Patch Notes #315 — Bake-Off Season and the Question of the Year
Claude Opus 4.5 shipped as scheduled (November 24th, agentic-coding benchmarks leading, pricing restructured downward: the cost-collapse now operating inside the frontier tier, not just beneath it), and our Q4 bake-off ran the year’s full portfolio, GPT-5.1-class, Gemini 3, Opus 4.5, plus the self-hosted distillation tier, through the golden sets with the year-end finding the file considers 2025’s actual technical summary: the frontier models are now functionally interchangeable on ~80% of our workload (the commodity tier arrived exactly as the portfolio thesis priced), meaningfully differentiated on the agentic 20% (sequence reliability, tool-use judgment, long-horizon coherence, the differentiation is the trusted-alone-longer axis, as pre-filed), and the procurement leverage this affords has inverted the vendor relationship entirely: the labs’ enterprise teams now ask to see our evals to understand why workloads move (the show-me-your-eval-suite doctrine, running in both directions, the instrument became the market). The boring layer’s thirteenth consecutive correct year closes its books. ...