Observed arrival · 2026-09-21
Personal Agent Bench: Four Agents, Taken Apart
A runtime-level comparison of Pine, Instinct, Muse, and Town, examining where their models run, how tools reach them, and where memory lives.
- For
- People choosing or building personal AI assistants
- Worth noticing
- The comparison records unknowns alongside findings and publishes no scores before task runs exist.
Field notes
The scorecard treats personal-agent evaluation as an inspection problem rather than a feature-counting exercise. Its table distinguishes server-side facts, read-only markdown, plain markdown in a home directory, and a server-side wiki across the four products. It also records tool-exposure counts and marks missing instruction text, schemas, hosts, and model information instead of converting those gaps into assumptions. Numeric scoring is reserved for a later stage involving recorded task runs.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue