Open-weight models close the reasoning gap

The headline number is the benchmark score. The more useful number is the gap: for roughly two years, closed models held a consistent lead on multi-step reasoning tasks, and that lead was widening. This release closes most of it.

Benchmark parity is not production parity. Latency, tooling, and the operational cost of self-hosting all remain real. But for organisations that cannot send data to a third-party API — healthcare, defence, parts of finance — the question was never price. It was whether the capability existed at all on hardware they control.

Treat the benchmark with the usual suspicion. Contamination is endemic, and a score published by the party that trained the model is a claim, not a finding.


Related