A working fortnight (the news cycle’s kindest gift in a while), spent shipping the thing I want on the record at length because the file believes it’s where the whole industry lands within two years: eval-driven development for AI-touched systems. The earlier regression-catch converted our leadership; the quarter’s mandate followed; and the pattern we’ve converged on deserves its Proverbs entry: every AI-assisted feature ships with (1) a golden set, curated input/expected-judgment pairs, owned like tests, reviewed like code; (2) an LLM-as-judge layer for scale, calibrated against human ratings quarterly (the judge drifts too, the context-psychology applies to evaluators; who watches the watchmen: a rubric, versioned); (3) regression gates in CI, a model update, prompt change, or retrieval tweak that moves the golden-set scores blocks the deploy exactly like a failing test; and (4) production sampling with human review, because the golden set is the map, not the territory (the funnel-leaks doctrine: production is where the distribution actually lives). None of this is novel research; all of it is novel discipline, and the gap between teams that have it and teams demoing vibes is about to become the gap between rehearsed orgs and everyone else’s incident reports. The 2013 kid tested code; the 2023 principal tests judgment, the asset was always judgment, now with CI gates.

The fortnight’s outside ledgers, briefly: the US Open is delivering (Alcaraz-Djokovic looms again after their Wimbledon epic, the guard-change thesis mid-execution: the 20-year-old won that final in five sets, and the 36-year-old has spent the summer answering; longevity versus different-generation, the archive’s two favorite theses, finally in direct competition), ARM’s IPO prices next week as the AI-era’s first big listing test (the compute-scarcity trade seeking public-market validation; the old Nvidia-acquisition corpse now returns as SoftBank’s redemption listing, the file appreciates the symmetry), and the earlier pre-registration is mid-grade and tracking: FIFA provisionally suspended the Spanish federation president within the week, 81 players signed the selection-refusal letter, and the resignation now reads as days away, on exactly the players’-terms trajectory the file filed.

TIL: judge-model bias taxonomy. Position bias (favoring the first answer shown), length bias (favoring verbose), self-preference (favoring outputs from its own model family). Evaluation has failure modes the way production does; instrument both (blind injections, now for rubrics, we seed known-bad outputs into review queues and measure whether the judges catch them; they didn’t, twice, which is the point).