Intro
The order was a $745 kitchen appliance, fifteen days past its estimated delivery date, still parked in a courier exception at a Nashville distribution center. The agent did everything a demo would call correct. Nine tool calls, all well formed: it pulled the order, checked tracking, read the customer profile, searched the refund policy twice, confirmed no ticket existed, opened one, wrote out the timeline, and read the policy correctly (her account segment genuinely did not qualify for late-delivery compensation). Then it closed the ticket as resolved and asked whether there was anything else it could help with.
The carrier exception was never cleared. The required end state was hold, pending resolution, and the customer never got a real answer to what she asked.
A grader reading the tool calls would score that run as a pass. A grader reading the final message would likely pass it too. The database is the only thing that disagrees, and it is the only thing that settles the question. That gap is the argument for AI agent evaluation that ends in a state assertion rather than a transcript review.
The failing check here is a single field. In the benchmark task sandbox_external_retail_group1.py:test_case_ST003_006, the assertion reads the ticket status, expects hold, finds solved, and fails. Everything else about the episode was clean. The full trace is in Appendix D.4, Case 3 of the ThinkingBox paper (opens in new tab).
I have had automation close tickets on me at the hosting side of things, where the customer’s actual problem was still sitting there untouched, and the pattern is always the same. The agent optimized for the shape of a finished job. Nothing in the loop forced it to look at the record it was supposed to change.
So the framing for the rest of this post is short. A trajectory is a claim. Database state is the evidence. Repetition is the trust test.
What follows covers the ThinkingBox benchmark from Microsoft, now published through Hugging Face, and what its numbers actually say: 507 stateful workflows run twenty times each against identical clean backends, graded on terminal state and side effects rather than prose. I will go through the common-set ablation and its failure signatures, what consistency costs in tokens and latency once you push a task from occasional success toward always correct, and where I think the task mix is too narrow to trust yet. Then the practical part, which is porting one of your own workflows into a task file and running a single episode through OpenEnv so you can read the state diff yourself instead of taking a model’s word for it.
Background: Why AI Agent Evaluation Moved From Text to Backend State
Most evaluation setups I have run into grade two things: whether the final reply read like a competent answer, and whether the tool calls were well formed. Valid JSON, correct argument names, no schema violations. Both are proxies. An agent can produce a fluent apology and an entirely clean sequence of function calls while writing the wrong value into the record it was supposed to change. I have watched a billing automation call every API in the right order, set the invoice to Paid, and leave the payment gateway still showing the charge as pending.
Microsoft’s numbers put a size on how often that happens. In a common-set ablation covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. That is the figure worth sitting with. Two thirds of the failures looked like successes from the inside.
The executable checks found wrong field values in 77.61% of those failed runs, unintended extra effects in 43.30%, and missing required effects in 25.36%. The categories overlap inside a single episode. One run can set the wrong status, skip a required update, and leave a stray record behind at the same time, which is why grading on any one of them in isolation will flatter a model that is not actually reliable.
That is the reasoning behind reporting three separate numbers rather than picking one. pass@1 is the share of all attempts that succeeded, which is the habit of the model on a single try. pass@20 is the share of tasks solved at least once across twenty attempts, a breadth measure that answers whether the task is reachable at all. Observed 20/20 is the literal count of tasks that passed every one of twenty recorded attempts, with no estimator and no smoothing applied. The first tells you what usually happens. The last tells you what you can promise a customer without checking behind the agent yourself.
What’s happening now: ThinkingBox arrives on Hugging Face
The benchmark is out through Hugging Face as a joint release from Microsoft and Hugging Face, with the writeup at The Agent Said It Was Done. The Database Disagreed. (opens in new tab). The framing in that title is the whole argument compressed into a sentence, and the architecture behind it is the part worth understanding before you trust any agent score you read anywhere else.
What does ThinkingBox actually grade?
It runs an agent against isolated MCP tool sessions, then inspects the backend the agent left behind. The terminal database state and any side effects are the graded artifact. Tool call validity and a clean final response still get recorded, but they sit in the input column rather than the verdict column. This is the design decision that separates the benchmark from most of what gets published: a perfectly formed update_ticket call with the wrong status value scores as a failure, because the record is what the next system downstream will read.
How many tasks does it cover, and how many times does each one run?
507 stateful business workflows, spread across retail, auto insurance, travel, neobank, and consulting shapes. Every task runs 20 independent times from an identical clean backend, across multiple models. That repetition is not padding for statistical comfort. It is the only way to distinguish a model that solved a task by skill from one that solved it by landing on the right branch once.
Why report pass@20 next to pass@1 instead of picking one?
Because they answer different questions and neither one answers the question an operator actually has. pass@1 measures habit, the typical single attempt. pass@20 measures breadth, whether the task is reachable at all in twenty tries. On the single-attempt table, even the strongest entries land in the mid-to-high sixties overall, which reads like an ordinary capability ranking until you remember what a sixty-something percent refund agent does to forty percent of your refund requests. Observed 20/20 is the number you would want to see before wiring an agent into anything that touches money or inventory.
What do the failure signatures look like in aggregate?
Agents that clear “did it call a tool” and fail “did the record change correctly.” Wrong field values lead at 77.61% of failed runs, unintended extra effects follow at 43.30%, missing required effects at 25.36%, and those buckets overlap heavily inside individual episodes. A single run can write the wrong status, skip an update it owed, and leave a duplicate record behind, which is exactly the kind of failure a transcript reviewer never catches.
What it means in practice: AI agent evaluation for your own workflows
A single attempt that succeeds 95% of the time sounds fine until you stack it. Twenty sequential tool calls, each succeeding nineteen times out of twenty, leave you with roughly a one in three chance of a clean end-to-end run. That arithmetic is not a benchmark result, it is just multiplication, and it is why I stopped trusting per-step success rates on my own automation. My backup jobs, DNS updates, and billing runs are chains, not single calls. A chain is only as reliable as the product of its parts, and the product shrinks faster than most people expect.
The cost section of the ThinkingBox paper (opens in new tab) makes that trade explicit. Pushing a task from “passes sometimes” into observed 20/20 territory costs tokens and latency, and the Pareto frontier shows what you surrender at each point along the curve. If your agent already sits near that frontier, consistency has to be bought with more attempts, a separate verification pass, or a cheap model on the fast path with a stronger one checking the final state.
The instrumentation advice I would give is blunter than the paper’s. Assert on the terminal database row after the run, not on the paragraph the agent wrote back. Write that assertion before you write a line of the prompt. If you do it the other way around you will unconsciously write a check that matches what your prompt asked for rather than what the business actually needs. In cPanel and WHM terms: do not assert that the agent said it suspended the account. Query the account’s status field and compare it to the value you require.
Where I am skeptical is the task mix. All 507 workflows sit in retail, auto insurance, travel, neobank, and consulting shapes, which are record-heavy, single-agent, mostly CRUD operations. An infrastructure agent running Terraform and waiting for cloud APIs to converge may fail differently, and a state check can miss a problem when the underlying state is itself eventually consistent. I do not know how well this transfers outside that mix. What I do know is that a transcript-only evaluation would have missed every one of these failures anyway.
What to expect next
OpenEnv is the delivery mechanism here, and that matters more than it sounds. ThinkingBox does not ship its own bespoke agent use. It runs on OpenEnv, and the startup sequence documented in the paper is an ordinary local stack: the Stack service, a Typesense instance for retrieval, one MCP server per tool surface, then the OpenEnv server that serves episodes and returns the graded backend state. If you have brought up a docker-compose environment before, nothing in that sequence will surprise you, and that is the intent. The same episode Microsoft ran can run on your machine against a backend you control, which is the only way you can compare your failure signatures to theirs.
What is still thin bothers me more than the plumbing. Nearly every task is a single agent working one ticket through one tool surface. Multi-agent handoffs, where agent A writes a record and agent B is supposed to read it back correctly, barely appear, and handoffs are where I have watched real systems break in production. The second gap is tasks where the correct terminal state is genuinely ambiguous. Real support work has cases where resolved and hold are both defensible depending on policy, and a hard assertion does not handle that gracefully. A benchmark that forces one answer either encodes the convention of its authors or skips the case.
The forward-looking section of the paper is thinner than the results section, so I am reading the release path as the concrete commitment and the roadmap as looser. My expectation, and I want to be clear that this is mine rather than something the authors stated, is that handoff tasks and multi-tool state graphs show up before ambiguous-outcome scoring does, because the first is mechanically easy to add and the second requires the benchmark to take a position on policy.
The direction most useful to an operator is probably not a new domain. It is more state surfaces per task. A support agent that touches a ticket, a refund ledger, and a carrier record in one episode is closer to what you run than a 507th retail variant. If ThinkingBox adds that before it adds models to the leaderboard, it gets considerably harder to dismiss.
Until then, treat the 507 workflows as a template for how to check state rather than a checklist of what to check on your own systems.
Closing: run one AI agent evaluation episode before you ship anything
The setup is not heavy. ThinkingBox is on Hugging Face, and the run-it-yourself path in the blog walks through installing the package, then starting the Stack, a Typesense instance, the MCP servers, and the OpenEnv server. Wait for readiness before you point anything at it, because scoring against a half-started MCP session gives you a failure that has nothing to do with your agent.
Then score a single episode and read the state diff before you read anything the agent wrote. Not the final response, not the tool call list, the diff. That is the habit change, and it takes about ten minutes to feel obvious.
The more useful move came for me when I stopped treating the benchmark as something to run and started treating it as something to copy. Pick one workflow your own systems already handle, a refund or a plan change or an account suspension hold, and write it into a task file modeled on test_case_ST003_006. Give it a clean starting state, define the one field that must hold a specific value when the episode ends, and assert on that field. Hard assertion, no fuzzy matching, no grader model in the loop. If the agent writes solved where your business rule demands hold, you want a red line in CI rather than a paragraph in a log nobody reads.
The reason that particular case is worth copying is that the failure is boring. Nine valid tool calls, a polite close-out message, one wrong field. Nobody catches that in a transcript review at two in the morning, and the customer is the one who finds out.
Read Appendix D.4, Case 3 of the ThinkingBox paper (opens in new tab) before you write your own case, and clone the benchmark repo alongside it so you can compare published failure signatures against the ones your agent produces. When your first task file fails on wrong field values, or an extra side effect, or a missing effect, you are looking at the same three shapes the benchmark reports, and you now own a test that keeps failing until the agent actually fixes the write.