The Auditor's Checklist
This video presents the same text shown beside it, spoken and on screen. It adds nothing the text does not say.
Ten questions compress this course's evaluation machinery — each with a mechanism behind it, and a claim that survives all ten deserves provisional belief.
The questions, with their receipts. One: what is the task, as measured — named metric, named target, or no claim. Two: what distribution was tested, and how far from it will deployment sit. Three: what were the conditions — curation, retries, human assistance, stakes. Four: which room is the denominator — every rate is a share of something; find the something. Five: how does it score on the rare class, not the easy average — the screening count's lesson, since lopsided targets reward refusing to look. Six: how old and how public is the exam — contamination, aging, leaderboard-gaming. Seven: what bets does the method embed, and does this task reward them. Eight: what is the whole-chain reliability, not the per-step — multiply along. Nine: whose utilities are being optimized, and were the affected represented. Ten: can failures be explained, debugged, and monitored after deployment. Now run the drill on a vendor sentence: "Our agent autonomously handles 80 percent of customer tickets with 95 percent satisfaction." Question one falls first — handled, measured how? Question two: whose tickets, drawn from which month? Question three: what escalation and human cleanup hide inside "autonomously"? Question four: 95 percent of what — surveyed whom, counting whom? And question five: what happened inside the unmentioned 20 percent, where the hard cases live? Five conspicuous failures, one sentence — and the same drill, run on a dismissal, works identically. A claim that survives all ten has earned provisional belief — provisional because distributions drift, exams age, and monitoring is question ten for a reason.
A claim, or a critique, that cannot say which question it survives is mood — the checklist is deliberately symmetric.