The headline number is the benchmark score. The more useful number is the gap: for roughly two years, closed models held a consistent lead on multi-step reasoning tasks, and that lead was widening. This release closes most of it.
Benchmark parity is not production parity. Latency, tooling, and the operational cost of self-hosting all remain real. But for organisations that cannot send data to a third-party API — healthcare, defence, parts of finance — the question was never price. It was whether the capability existed at all on hardware they control.
Treat the benchmark with the usual suspicion. Contamination is endemic, and a score published by the party that trained the model is a claim, not a finding.



