Observed arrival · 2026-08-24
BlueCollarBench Gives AI a Journeyman Exam
An open benchmark tests five AI models on 400 practical tasks across 10 skilled trades.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
Each model receives the same fixed bank of 400 questions, divided across 10 trades with roughly 40 tasks per trade. The grading scheme separates practical knowledge from code or specification lookups and from catching a bad premise or hidden hazard. Answers that hit a hard-fail safety item score zero, while the separate Arena reports blind human pairwise preferences as an Elo rating.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
ƒJavaScriptBrowser-side code central
One card from the complete issue