Skip to the card

Card 366 of 10002026-08-28 issue

Observed arrival · 2026-08-28

Haocun Ye’s Benchmarks That Don’t Take Scores at Face Value

haocunye.site Observed source
Editorial interest 82/100 Selection signal · not a rating of the site

A researcher homepage documenting work on multimodal large models, reinforcement-learning post-training, image editing, and audits of whether benchmarks measure what they claim.

Landing page captured for the 2026-08-28 issue.

Field notes

The homepage presents research as a connected evaluation workflow rather than a list of model claims: papers, code, blog explanations, and citation records sit alongside each project. One clinical-text audit reports 0.944 macro-AUROC even after diagnosis redaction, while the proposed FlipTrack benchmark uses paired images whose answers must change, making image-blind answering easier to expose. The page also links to external scholarly profiles and a downloadable CV.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use
HumanPersonal, local, civic, or handmade

One card from the complete issue

Porion Gets More Tolerance

315,344 arrived 1,000 judged 1000 catalogued Enter the complete issue
haocunye.site

Landing page observed 2026-08-28. The live site may have changed.