Skip to the card

Card 392 of 9722026-09-25 issue

Observed arrival · 2026-09-25

The Examination Room puts AI through a clinical viva

judgementlabs.app Visit website
Editorial interest 79/100 Selection signal · not a rating of the site

A clinician-written benchmark tests whether AI models gather evidence, order tests, and defend treatment plans under time and budget constraints.

Landing page captured for the 2026-09-25 issue.
For
AI evaluation researchers and clinical safety teams
Worth noticing
A private notebook stays hidden from three examiners while the model works through a ₹45,000 budget and 25-minute clock.

Field notes

The environment withholds case facts until a model requests them, then tracks tests against a ₹45,000 budget and a 25-minute clock. Three examiner agents can challenge a plan for up to 65 turns; an event ledger records actions, while two clinicians review runs blind, then compare marks and settle disagreements. The reported sample is 24 runs across six models, all using one clinician-written dental case.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Vandalism in Demo Mode

315,813 arrived 1,000 judged 972 catalogued Enter the complete issue
judgementlabs.app

Landing page observed 2026-09-25. The live site may have changed.