Microsoft and Hugging Face have published ThinkingBox, a benchmark that scores AI agents on whether a database actually changed rather than on whether the agent reported finishing. The headline finding is the gap between the two.
Across 507 stateful business workflows, each run 20 times against a range of language models, the authors report that of the runs that failed the executable state check, “67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error”. The agent looked successful by every available signal short of inspecting the data.
On pass@1, Claude Opus 5.5 leads overall at 67.16%. The consistency numbers are worth more than the leaderboard position: the authors note that Kimi-K3 solves 75 more tasks at least once than Opus 5, while Opus 5 solves 173 more tasks consistently than Kimi-K3. Breadth and reliability are not the same property, and a buyer piloting an agent on a handful of runs will mostly measure the first.
The tasks cover retail (98), auto insurance (100), travel (104), a neobank (104) and consulting (101). The harness is MIT-licensed, the benchmark data is released under CDLA-Permissive-2.0 and the OpenEnv environment under BSD-3-Clause, so the figures can be argued with rather than taken on trust.
Source: Microsoft and Hugging Face, The Agent Said It Was Done. The Database Disagreed.
