Observed arrival · 2026-09-24
jeval: evaluator judgments in one API call
jeval presents an API that runs evaluators from tools such as RAGAS and promptfoo against one submitted item, returning scores, probabilities, and confidence.
- For
- Developers evaluating LLM and RAG outputs
- Worth noticing
- The page describes per-chunk grounding verdicts and per-step or per-call agent judgments alongside standard evaluator scores.
Field notes
The page describes a single request carrying an item’s input, output, expected answer, context, criteria, tool calls, or trajectory, with selected evaluator questions run against that state. Its examples extend beyond one overall score: retrieved chunks receive individual grounding verdicts, while agent trajectories can be judged step by step or call by call. The site reports 43 evaluators, with 33 answered by Jev and 10 in code; its latency and cost figures are homepage claims.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue