Observed arrival · 2026-08-28
Haocun Ye’s Benchmarks That Don’t Take Scores at Face Value
A researcher homepage documenting work on multimodal large models, reinforcement-learning post-training, image editing, and audits of whether benchmarks measure what they claim.
Field notes
The homepage presents research as a connected evaluation workflow rather than a list of model claims: papers, code, blog explanations, and citation records sit alongside each project. One clinical-text audit reports 0.944 macro-AUROC even after diagnosis redaction, while the proposed FlipTrack benchmark uses paired images whose answers must change, making image-blind answering easier to expose. The page also links to external scholarly profiles and a downloadable CV.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue