Observed arrival · 2026-09-18
Jevals: a benchmark for models that decide
An independent evaluation site comparing Jev and six LLMs on typed yes/no, pick-one, and rubric decisions.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Model evaluators and developers building decision systems
- Worth noticing
- Each model is run five times per question, with results compared across three decision formats.
Field notes
The benchmark separates three kinds of model decision: binary approval, selecting one option, and placing an answer on an ordered rubric. Its examples show how a returned probability can feed directly into an if-statement or a human-review fallback. The visible release compares Jev with six LLMs, repeats each question five times, and includes datasets such as PubMedQA and Banking77.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue