Patch Notes #273 — Three Claudes and a Logo Change
Anthropic shipped the Claude 3 family March 4th, Haiku/Sonnet/Opus, and the flagship posted benchmark wins over GPT-4 on the standard suites: the first time the same-day-launch rival has held the measured frontier, however temporarily, and the market-structure note matters more than the leaderboard (the dance is now a three-body problem, OpenAI/Microsoft, Google, Anthropic/Amazon-Google-money, with Meta’s open-weights flank as the fourth gravitational mass; frontier capability is officially plural, which changes procurement, eval discipline, and the governance math simultaneously: you cannot license a frontier that keeps electing new members). Our own quarterly bake-off (the tooling benchmarks) confirmed the delta on our tasks, which is the only leaderboard the file trusts (the golden-set doctrine): model choice is now a quarterly decision with a regression suite, exactly the commodity-with-switching-costs dynamics this archive filed for clouds a decade ago (the chokepoint homework, now with per-token pricing). The “sparks”-era question (which capabilities arrive at which scale) remains unanswered by anyone including the labs; the eval profession remains the only tractable response; the file remains on message because the message keeps grading correct. ...