ThinkingBox: Microsoft Grades AI Agents by the Database State They Leave Behind
An agent can make nine well-formed tool calls, close a ticket as resolved, and still leave the database in the wrong state — ThinkingBox measures exactly that gap across 507 stateful workflows.
ThinkingBox, from Microsoft and Hugging Face, is a sandbox and benchmark that grades agents on terminal backend state and side effects rather than tool-call validity or final responses. Each of 507 synthetic business workflows — retail, auto insurance, travel, neobank, consulting — runs 20 times from an identical clean backend against isolated MCP tool sessions. Executable judges compare the resulting database state against a required end state, rejecting wrong, missing, or extra effects. 477 tasks are graded on state alone; 30 add response rubrics. The benchmark is MIT-licensed for the framework, CDLA-Permissive-2.0 for the data, and runs through Hugging Face's OpenEnv interface.
The headline finding is that single-attempt accuracy hides consistency collapse. Claude Opus 5.5 leads pass@1 at 67.16%, but only 241 of 507 tasks (47.53%) pass all 20 attempts. Kimi-K3 solves the most tasks at least once — 476 of 507, or 93.89% — yet only 68 tasks (13.41%) pass 20/20. Claude Opus 5.5 and Claude Opus 5 pass exactly the same number of tasks consistently (241) despite the newer model's higher headline score. In a 121,680-trial ablation across 12 models, 67.24% of failed attempts still terminated cleanly with a state-changing tool call and no tool error; executable checks found wrong field values in 77.61% of those, unintended extra effects in 43.30%, and missing required effects in 25.36%.
Failure signatures are dominated by tool handling: 79.9% of failures across the ablation are tool usage errors, versus 10.3% wrong state updates, 7.0% incomplete user resolutions, and 2.9% no state-changing action. Cost analysis prices consistency separately from single successes. GPT-5.4 is cheapest per dependable task at $6.80 (128 tasks passing 20/20), GPT-6 Astra reaches 231 tasks at $7.45, and Claude Opus 5.5 reaches 241 at $7.80. GPT-5.6 Sol is cheapest per single success at $0.127 but costs $9.76 per dependable task.
The practical recommendation is to treat the 20/20 rate as a design input, check terminal state before committing rather than trusting the model's summary, classify tool errors so retries target recoverable ones, and require human approval on irreversible changes.