Observed arrival · 2026-09-09
Medical AI Benchmark and Study Review
A public review board that grades medical-AI benchmarks and clinical studies for methodological quality.
Field notes
The homepage records 14 benchmarks and studies, assigning each a letter grade alongside a 47–94 score range and a median of 77. Its rubric separates methodological concerns that are often collapsed into a single benchmark score, including model-version reporting, contamination controls, judge calibration, clinical comparators, and deployment framing. The page says reviewer disagreement of 1.5 points or more triggers discussion, while four named failures can impose grade ceilings.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue