Skip to the card

Card 394 of 9822026-09-19 issue

Observed arrival · 2026-09-19

Jevals: a benchmark for models that decide

j-evals.com Visit website
Editorial interest 82/100 Selection signal · not a rating of the site

An independent benchmark compares a probability-only model with six LLMs across yes/no, pick-one, and rubric decisions.

Landing page captured for the 2026-09-19 issue.
For
AI evaluators and developers building decision systems
Worth noticing
It separates yes/no, pick-one, and rubric decisions into three public comparison boards.

Field notes

The project divides evaluation into three decision forms: binary approval, choosing among options, and placing a response on an ordered rubric. Examples include 300-question slices of PubMedQA, Banking77, and HelpSteer2. The homepage says every model is run five times at list price and distinguishes Jev's probability outputs from LLM responses produced through an adapter. It also links to methodology, API, and CC-BY-4.0 data.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

Filed Under: Missing Tooth

336,080 arrived 998 judged 982 catalogued Enter the complete issue
j-evals.com

Landing page observed 2026-09-19. The live site may have changed.