Observed arrival · 2026-09-08
Covaric Wants to Make AI-Agent Failures Reproducible
Covaric tests AI agents by controlling their tool responses and model decisions, then reports which behaviors the tests exercised.
Field notes
The proposed workflow records a real agent run, replaces live tool behavior with selected conditions, and reruns the agent while controlling model decisions. The page names error, empty, partial, malformed, and delayed responses as test inputs, then measures coverage across steps, tools, response classes, and decisions. It reports evaluation on 35 open-source agents and cites AgentInspect at ISSTA 2026; the precision and recall figures remain claims made on the homepage.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue