Observed arrival · 2026-09-26
evalship reviews LLM evals on pull requests
evalship says it reads LLM evaluation results from GitHub Actions and comments on pull requests with regressions, likely causes, and prompt changes that lack test coverage.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Teams maintaining LLM prompt and evaluation suites
- Worth noticing
- Setup findings cite repository file-and-line locations, and the page says each citation is checked against the repo.
Field notes
The setup audit examples distinguish instructions with no test coverage from suites that rely entirely on an LLM judge. The page proposes concrete follow-up tests, including refund-limit cases, and explains that flaky evaluations are set aside when comparing results. Its sample also shows a pull-request comment updated in place, with merge blocking described as optional.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
⊠LoginAccess appeared gated
✦PrettyNotable craft visible
◎NicheUnusually specific use
One card from the complete issue