Observed arrival · 2026-08-25
Ebin Babu Thomas Measures What AI Safeguards Actually Stop
A personal AI-engineering portfolio documenting agent reliability, model-behaviour experiments, and open-source work with committed artifacts and failure reports.
Why it surfaced
Thomas publishes unusually concrete evidence: 434 of 434 simulated kill-point recoveries, 16 of 16 PDF conversions verified by render-back diffs, and a preregistered experiment whose reported effect survived after 65.8% of distress language was trained away. The portfolio treats limitations and failed tests as part of the work rather than burying them.
A personal AI-engineering portfolio documenting agent reliability, model-behaviour experiments, and open-source work with committed artifacts and failure reports.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue