Key takeaways
- Microsoft and Hugging Face published the joint release at 22:56 UTC on October 3. ThinkingBox is now packaged behind the OpenEnv interface; the underlying v1.0 benchmark and paper were already public in August.
- ThinkingBox-Bench contains 507 synthetic, policy-conditioned workflows across retail, travel, auto insurance, neobank support and consulting IT/HR. Every attempt starts with isolated state and is graded on the records and side effects left behind.
- In the authors’ 121,680-trial common-set ablation, 79,853 attempts failed executable checks. Of those failures, 67.24% still ended cleanly after a state-changing tool call and no final tool error.
- The joint post reports Claude Opus 5.5 leading pass@1 at 67.16%. GPT-6 Astra retained 78% of its single-run score under the post’s observed 20/20 comparison, while Kimi-K3 solved 93.89% of tasks at least once but only 13.41% in all 20 recorded attempts.
- The dataset, framework and adapter are public, but inference is not included. Reproduction requires Linux or WSL, Python 3.11+, uv, Docker, Typesense, several local services and configured endpoints for agent, simulated-user and judge roles.
The October 3 release puts ThinkingBox behind OpenEnv
Microsoft and Hugging Face published their joint ThinkingBox article at 22:56 UTC on October 3, 2026. The practical release is an OpenEnv adapter for a benchmark that was introduced earlier: ThinkingBox-Bench v1.0 was tagged on August 17, and the paper first appeared on August 20 before an October 1 revision. The announcement therefore expands how teams can run the environment; it does not introduce 507 previously unavailable tasks.
Available now are an MIT-licensed ThinkingBox framework, CDLA-Permissive-2.0 benchmark data, a public Hugging Face dataset and a BSD-3-Clause OpenEnv integration. There is no hosted evaluation endpoint, bundled inference price or managed service tier in the reviewed release. Operators supply the infrastructure and model endpoints.
The benchmark grades outcomes, not plausible trajectories
Each task defines a starting backend state, user goal, policy, MCP-compatible tools and executable checks for the required final state. A simulated user holds details that the agent must request. Every attempt receives a clean isolated session; afterward, a side-effect extractor and deterministic checks compare what changed with the allowed result. Of 507 tasks, the authors say 477 use state-only grading and 30 add a narrow final-response rubric.
That design catches a failure a tool-call grader can miss: an agent may use valid tools, receive no final error and write to the database, yet still update the wrong field, omit a required effect or create collateral state. The public example follows a customer-service agent that correctly researches a delayed order but closes the ticket as resolved when it should remain on hold.
Clean execution concealed most failed outcomes in the ablation
The authors report a common-set ablation of 121,680 valid trials across 12 models. Executable checks rejected 79,853 attempts. Among those failures, 67.24% still terminated cleanly, called a state-changing tool and reported no final tool error. The checks found wrong field values in 77.61%, extra effects in 43.30% and missing required effects in 25.36%; those categories overlap.
A separate diagnostic assigns one signature to each failed trace. The post reports an unweighted cross-model average of 79.9% for tool usage, 10.3% for wrong state updates, 7.0% for incomplete user resolution and 2.9% for no state-changing action. These labels describe observed failure signatures, not proven root causes, and AccessAllGPT did not reproduce the classification.
Single-run capability and repeat reliability separate sharply
The full table covers 18 proprietary and open-weight models, 507 tasks and 20 independent attempts per task. In the authors’ results, Claude Opus 5.5 leads task-weighted pass@1 at 67.16%, followed by Claude Opus 5 at 66.50% and GPT-5.4 at 65.36%. Kimi-K3 leads the listed open-weight group at 57.37% and solves 476 tasks—93.89%—at least once.
That breadth does not imply dependable execution. Kimi-K3 passes only 68 tasks, or 13.41%, in all 20 recorded attempts. Claude Opus 5 and Opus 5.5 each pass 241 tasks, or 47.53%, every time. The post says GPT-6 Astra retains 78% of its pass@1 result under this comparison, while both Opus models retain 71%. These are benchmark-author results on synthetic workflows, not guarantees for production agents.
The cost table is a comparative estimate, not a serving quote
The authors priced recorded campaign tokens with undiscounted OpenRouter list rates captured on September 20, reversing promotional discounts and excluding endpoints that declared quantization. They divide an estimated 507-attempt run by successful attempts for cost per success, then divide the full 20-run campaign by tasks that passed all 20 attempts for cost per dependable task.
Under that method, GPT-5.4 is reported at $6.80 per dependable task for 128 all-20 successes, GPT-6 Astra at $7.45 for 231 and Claude Opus 5.5 at $7.80 for 241. Those figures are a normalized research index, not invoices or current vendor prices. They omit production integration, retries outside the benchmark, human review, networking and operational overhead, and the underlying rates can change.
Running it requires more than downloading one dataset
The documented path needs Linux or WSL, Python 3.11 or newer, uv and Docker. Users clone OpenEnv and the tagged benchmark data, install the ThinkingBox CLI, start Typesense 30.1, launch the session proxy and MCP servers, configure agent, simulated-user and judge endpoints, start the OpenEnv service and then run the packaged evaluator. Operational failures are written separately so they can be rerun rather than silently blended with model outcomes.
That is a useful reproducibility surface, but it leaves important work to the operator: endpoint compatibility, secrets, model snapshots, inference settings, service health, cost capture and error accounting. Teams should pin all three repositories, preserve the configuration and record every excluded or retried attempt before comparing a local result with the published table.
What agent teams should do next
Use ThinkingBox first as a measurement pattern. For a consequential workflow, define the permitted terminal state, forbidden side effects and response obligations independently of the model. Reset state between trials, repeat the same task enough times to expose variance, retain complete traces, and separate model failures from infrastructure failures. A fluent answer and successful HTTP status should never be the acceptance criterion for a database mutation.
Adopt the public harness when its five business domains and local service stack resemble the workflow you need to study. Constrain conclusions to the pinned configuration and report pass@1 beside an explicitly defined repeat metric. Wait before using its leaderboard as a procurement decision if your tools, policies or data differ materially. Reject autonomous production writes when the system cannot verify final state, prevent collateral effects and route uncertain outcomes to a human.
Copy-ready stateful-agent evaluation record
Complete this before accepting a benchmark run or allowing an agent to mutate production records.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
User goal, starting state, policies, required final state, forbidden side effects and response obligations.
Model IDs, provider endpoints, prompts, tools, framework and data commits, container images and inference settings.
Reset procedure, unique session state, cache controls, credential scope and evidence that attempts cannot contaminate one another.
Assertions for correct fields, missing effects, extra effects, response requirements and tests showing each judge can fail.
Trial count, pass@1 calculation, at-least-once and all-runs definitions, uncertainty and rationale for the production reliability target.
Model, tool, timeout, endpoint and evaluator failures; retry policy; excluded attempts; complete trace and error sidecars.
Token counts, rate snapshot and source, cache treatment, campaign cost, cost per success, human effort and infrastructure overhead.
Differences between synthetic tasks and production data, policies, tools, identity, scale, latency and irreversible consequences.
Adopt, constrain, wait or reject; approved writes, verification before commit, human escalation, rollback and revalidation date.
Primary sources
Browse the publication-wide evidence index →
- The Agent Said It Was Done. The Database Disagreed.Microsoft and Hugging Face · Reviewed: October 3, 2026 publication metadata; outcome and repeat-reliability definitions; 18-model results; cost methodology; failure signatures; sandbox design; OpenEnv setup; release links; licenses; synthetic-data disclaimer · Retrieved · Supports: The joint article announces ThinkingBox through Hugging Face and OpenEnv, describes 507 stateful workflows repeated 20 times, reports model, consistency, cost and failure results, and links the runnable harness, tagged benchmark data and paper. Its measurements are author-reported rather than independent AccessAllGPT results.
- One Success Isn’t Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows (v2)ThinkingBox authors on arXiv · Reviewed: August 20 submission and October 1 revision metadata; abstract; 507-workflow scope; isolated MCP-compatible sessions; terminal-state evaluation; repeat-reliability results; response-rubric boundary; released code and benchmark references · Retrieved · Supports: The author paper defines the benchmark and reports the reliability gap, including Claude Opus 5 moving from 66.50% pass@1 to 47.53% observed all-20 success and Kimi-K3 from 57.37% pass@1 to 17.60% under the paper’s pass^20 metric. These are paper results, not independent replication.
- ThinkingBox-Bench v1.0 release at commit fcaba4cMicrosoft on GitHub · Reviewed: Annotated thinkingbox-bench-v1.0 tag target; August 17 release notes; repository tree; task and server layout; setup documentation; canonical task-selection surface; CDLA-Permissive-2.0 benchmark-data license · Retrieved · Supports: The immutable tag target exposes the executable 507-task benchmark across five business domains. The release predates the October 3 Hugging Face announcement, so the news is broader packaging and access rather than creation of a brand-new dataset.
- ThinkingBox-Bench dataset at revision 5e49f89Microsoft on Hugging Face · Reviewed: Pinned public dataset revision; dataset card; 507-row task collection; five-domain scope; file inventory; task fields; CDLA-Permissive-2.0 license metadata; linked paper and source repository · Retrieved · Supports: The public, ungated registry revision establishes that the benchmark dataset is downloadable on Hugging Face and identifies its license and paper. Registry availability does not independently validate the authors’ model scores or cost estimates.
- ThinkingBox OpenEnv adapter at commit 1fe8090Hugging Face OpenEnv on GitHub · Reviewed: Pinned adapter tree; environment package; server and evaluator entry points; example usage; configuration and readiness surfaces; OpenEnv BSD-3-Clause repository license · Retrieved · Supports: The pinned OpenEnv artifact provides a public integration layer for running ThinkingBox episodes. It still requires local services, benchmark data and configured model endpoints; it is not a hosted benchmark API with included inference.
Limitations
AccessAllGPT did not install or run ThinkingBox, OpenEnv, Typesense, the MCP servers or the packaged evaluator. We did not download and inspect every benchmark row, audit executable assertions, call any of the 18 models, recreate 20 independent attempts, verify endpoint versions, examine private prompts, recompute token usage or September 20 prices, validate reported pass@1 or repeat scores, reproduce failure signatures, test contamination, measure confidence intervals, compare the synthetic workflows with confidential enterprise data, or conduct a security review. The blog, paper, repositories and dataset are controlled by the authors or release partners. Public code and pinned revisions improve inspectability but do not independently establish result correctness, domain realism, safe deployment or future model performance. The paper’s pass^20 terminology and the blog’s literal observed 20/20 count should not be substituted for one another without following each document’s definition.
Disclosures
AccessAllGPT did not receive Microsoft, Hugging Face, OpenEnv, ThinkingBox or model-provider access, credits, data, code, support, a briefing, review or compensation for this article. Microsoft, Hugging Face, Toloka, OpenRouter and the cited model vendors did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Microsoft, Hugging Face or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- ServiceNow Details AutoSynthData After a 2,000-Task Agent-Training Run Improved Pass@1 by 7.2 Points
- Design an Agent Benchmark That Predicts Production
- Where Human Approval Belongs in AI Automation
- AI Agents vs. Workflows: Choose the Right Automation Pattern
- Build or Buy an LLM Evaluation Platform?
- OpenAI Launches Dots, an Always-On GPT-6 Astra Agent
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.