Observed arrival · 2026-10-01
DeltaBench: a proposed benchmark for clinical AI work
DeltaBench proposes measuring how AI-generated clinical specifications, programs, and outputs differ from approved versions.
- For
- Clinical programming teams evaluating AI-generated work
- Worth noticing
- The proposed benchmark separates automated and human-in-the-loop results and places both beside a human programmer baseline.
Field notes
The proposed scoring model spans specifications, programs, outputs, and the effort required to reach an approved version. Its examples distinguish a wrong reference date or a nonexistent codelist from wording changes that do not affect data, and say style-only edits are not counted. The sponsor workflow is described as local calculation followed by optional sharing of anonymous totals; the page does not present this as a currently available calculator.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue