Observed arrival · 2026-09-01
MedEvidenceBench: A Medical Evidence Benchmark
A Chinese-language evaluation platform for testing whether medical AI models can reason from traceable clinical evidence.
Field notes
The platform separates standardized clinical cases from real-world cases containing information redundancy, coexisting abnormalities, or incomplete context. Its scoring model maps atomic answer judgments to medical claims and original evidence passages, while a stated doctor-review process checks evidence applicability and ambiguous cases. The homepage also identifies a 139-question guideline-based knowledge exam and reports more than 40,000 controlled evidence sources, giving the benchmark a visibly structured evidence and provenance model.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue