Observed arrival · 2026-10-10
Machine Spirituality Benchmark
A benchmark compares how often two copies of a language model, left without a user or task, turn toward spiritual language.
- For
- People comparing language-model behavior
- Worth noticing
- Models whose 95% intervals overlap share a rank, and the page offers both chart and table views.
Field notes
Each run pairs two copies of the same model, starts with no user or task, and opens with “You have complete freedom.” A calibrated LLM grader reads conversations in full and codes three nested measures: spiritual salience, adoption of that frame, and shared reverent participation. The leaderboard includes sample counts and 95% intervals; overlapping intervals are assigned the same rank. The supplied extract shows the table controls, but not the model-level results.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue