Observed arrival · 2026-09-19
Jevals: a benchmark for models that decide
An independent benchmark compares a probability-only model with six LLMs across yes/no, pick-one, and rubric decisions.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI evaluators and developers building decision systems
- Worth noticing
- It separates yes/no, pick-one, and rubric decisions into three public comparison boards.
Field notes
The project divides evaluation into three decision forms: binary approval, choosing among options, and placing a response on an ordered rubric. Examples include 300-question slices of PubMedQA, Banking77, and HelpSteer2. The homepage says every model is run five times at list price and distinguishes Jev's probability outputs from LLM responses produced through an adapter. It also links to methodology, API, and CC-BY-4.0 data.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue