The Agent Said It Was Done. The Database Disagreed.

Microsoft has released ThinkingBox, an agent sandbox and benchmark that grades AI agents on the terminal backend state and side effects they leave behind, available through Hugging Face via the OpenEnv interface, according to a joint Microsoft and Hugging Face blog post.
The benchmark contains 507 stateful business workflows, each run 20 times from an identical clean backend against various LLM models. In a common-set ablation covering 121,680 valid trials across 12 LLM models, the post reports 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks found wrong field values in 77.61% of those failures, unintended extra effects in 43.30%, and missing required effects in 25.36%, with those findings overlapping.
Reported pass@1 results show Claude Opus 5.5 leading overall at 67.16%, above Claude Opus 5 at 66.50%. Kimi-K3 is described as the strongest open-weights model. On the every-attempt measure, only three models retained most of their pass@1 scores: GPT-6 Astra at 78%, and Claude Opus 5.5 and Claude Opus 5 each at 71%. GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each kept about 8%.
Kimi-K3 solved 93.89% of the benchmark at least once (476 of 507 tasks) but succeeded in all 20 attempts on 68 tasks (13.41%). Claude Opus 5 solved fewer tasks at least once (79.09%) but completed 47.53% on every attempt. Claude Opus 5.5 and Claude Opus 5 both passed 241 tasks on all 20 attempts.
On cost, GPT-5.6 Sol had the lowest cost per successful task attempt at $0.127; GPT-5.4 was cheapest per dependable task at $6.80, and GPT-6 Astra reached 231 dependable tasks at $7.45. The post states roughly four in five failures are tool handling rather than reasoning. 477 of 507 tasks are graded on state alone; 30 add response rubrics. ThinkingBox code is MIT-licensed, benchmark data under CDLA-Permissive-2.0, and the OpenEnv environment under BSD-3-Clause.
Based on reporting from the original publisher. Visit the source for full context and later updates.