Observed arrival · 2026-09-30
AssistantEval grades AI assistants against the record
AssistantEval describes a benchmark that puts personal AI assistants through simulated everyday tasks and grades them using saved service logs.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Teams evaluating personal AI assistants
- Worth noticing
- Tasks use per-run services over REST or MCP, saving request inputs, outputs, and before-and-after state for reproducible grading.
Field notes
Tasks give an assistant only the information a simulated user shares and require it to ask before actions that change records, affect other people, or spend money. The page describes code checking action order and value sources before judges assess meaning; people settle cases the judges cannot resolve. Example records are explicitly labeled illustrative, with made-up names and data.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue