Microsoft and Hugging Face published a joint post on October 3. ThinkingBox grades an agent on the records it leaves behind, not on the sentences it generates, and then asks whether it can do the same task 20 times in a row. The post says the benchmark is now available through Hugging Face. A well-formed tool call is not the same thing as a finished job.
[1]
The example is a $745 kitchen appliance stuck in a courier exception at a Nashville distribution center, fifteen days past the estimated delivery date. The agent makes nine tool calls: it pulls the order, checks tracking, looks up the customer, searches the refund policy twice, confirms that no ticket exists, opens one, documents the timeline, and reads the policy correctly. That account segment does not qualify for late-delivery compensation. It then closes the ticket as resolved. The post says two things are wrong: the carrier exception is still open, so the required end state is on hold, and the customer never gets a real answer. The executable check fails on one field. The ticket status is solved where the required end state is hold.
[1]The benchmark has 507 stateful business workflows. Each model runs every task 20 times, and the grade is the terminal backend state and its side effects. One common-set ablation covers 121,680 valid trials across 12 models. Of those, 79,853 attempts failed the executable checks. Among those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The checks nevertheless found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those three findings overlap. They are not three separate piles of failures.
[1]Claude Opus 5.5 leads the single-attempt score at 67.16% overall. Claude Opus 5 is at 66.50%. The two models pass the same number of tasks on all 20 attempts: 241. The post says that half a point of headline accuracy bought no additional dependability. Kimi-K3 has the broadest coverage of tasks solved at least once, but only 68 of 507 tasks, 13.41%, succeed on all 20 attempts. GPT-6 Astra retains 78% of its single-attempt rate. Opus 5.5 and Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro each keep about 8%.
[1]The prices are estimates at undiscounted list rates on OpenRouter. The post says they are not an invoice. Per successful attempt, GPT-5.6 Sol is the cheapest at $0.127. Per task that passes all 20 attempts, GPT-5.4 costs $6.80, but only 128 tasks meet that bar. GPT-6 Astra costs $7.45 across 231 tasks. Claude Opus 5.5 costs $7.80 across 241 tasks. Among failure labels, tool usage is 79.9%. The post says these are unweighted averages of per-model shares, not a single cause. It suggests checking the terminal state before committing a change, and it says the lift from that advice has not been measured on this benchmark.
[1]要点
- ThinkingBox grades the database end state and runs each task 20 times. In the $745 example, the ticket is marked solved when the required state is hold.
- Many failed trials among 121,680 still look clean. Opus 5.5 scores 67.16% on one attempt and, like Opus 5, passes only 241 tasks on all 20.
- The estimated cost of a dependable task is a different column from the cost of one success. Tool usage is 79.9% of failure labels, an unweighted average. The suggested fixes are not measured here.