Skip to the card

Card 015 of 9832026-09-30 issue

Observed arrival · 2026-09-30

AssistantEval grades AI assistants against the record

actually.help Visit website
Editorial interest 80/100 Selection signal · not a rating of the site

AssistantEval describes a benchmark that puts personal AI assistants through simulated everyday tasks and grades them using saved service logs.

Landing page captured for the 2026-09-30 issue.
For
Teams evaluating personal AI assistants
Worth noticing
Tasks use per-run services over REST or MCP, saving request inputs, outputs, and before-and-after state for reproducible grading.

Field notes

Tasks give an assistant only the information a simulated user shares and require it to ask before actions that change records, affect other people, or spend money. The page describes code checking action order and value sources before judges assess meaning; people settle cases the judges cannot resolve. Example records are explicitly labeled illustrative, with made-up names and data.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Ask the Descendants

390,826 arrived 1,000 judged 983 catalogued Enter the complete issue
actually.help

Landing page observed 2026-09-30. The live site may have changed.