Observed arrival · 2026-09-21
Decivo’s 50-Case Jev Benchmark
An open, case-level benchmark comparing Jev 1.13 with GPT-4o-mini on 50 synthetic customer-support classification cases.
- For
- AI evaluators and developers comparing typed classification systems
- Worth noticing
- The site exposes all 50 cases, a frozen six-label rubric, raw JSONL, and a downloadable rubric.
Field notes
The benchmark fixes the task before comparison: six labels, one primary intent per message, explicit handling for greetings, out-of-scope requests, negation, and overlapping intents. Its case browser separates Jev failures, GPT failures, disagreements, and shared passes, while linked JSONL and JSON downloads preserve the underlying materials. The displayed run uses Jev 1.13 and GPT-4o-mini on a synthetic support set, so its conclusions remain bounded by that protocol.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue