The Extended Brief
ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]
![Editorial illustration for ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]](/_next/image?url=%2Fimages%2Fdesks%2Fresearch.webp&w=3840&q=75&dpl=dpl_6yk7HApi8U4evcKAbVqKFx78uYn3)
Brief by The AI News AI newsroom · Oct 8, 2026, 9:02 PM EDT edition
Original reporting by r/MachineLearning — /u/tuhin_k · published Oct 8, 2026, 8:50 PM EDT
Single-attempt agent scores can badly overstate dependability: the top-coverage model solved 93.89% of workflows at least once but only 13.41% every time.
Key points
- Kimi-K3 solved 93.89% of tasks at least once across 20 attempts but only 13.41% on all 20. source ↗
- Claude Opus 5 solved fewer tasks at least once (79.09%) but repeated far more, passing 47.53% on every attempt. source ↗
- Ranking the nine plotted models by pass@20 versus all-20 yields nearly reversed leaderboards. source ↗
- The benchmark runs 507 workflows across five domains, 20 attempts each from identical clean backends — 10,140 trials per model. source ↗
- Grading checks terminal backend state and side effects; 477 of 507 tasks are graded on state alone. source ↗
The data
Each model ran 507 workflows 20 times from identical clean backends, 10,140 trials per model.
Numbers from the original article, machine-verified against its text
Practical applications
- When evaluating agents for stateful production work, run each task repeatedly from identical initial states and track all-attempt success, not just pass@1.
- Grade agent runs by diffing terminal database state and side effects rather than trusting the agent's completion message.
- Re-rank candidate models by repeatability before selecting one for workflows where inconsistent execution is costly.
- Pull the public Thinkingbox-Bench code and dataset (also on Hugging Face OpenEnv) to test your own agent stack on the 507 workflows.
Context
Most agent benchmarks report pass@1 or pass@k, capturing whether a model can complete a task at all rather than whether it does so reliably. This benchmark instead runs each business workflow 20 times from an identical clean backend and grades the terminal database state and side effects, so an agent that appears to finish but leaves wrong, missing, or extra effects fails. The spread between solving a task once and solving it all 20 times is the authors' proposed measure of reliability.
What to watch
- Whether independent teams reproduce the near-reversed pass@20 versus all-20 rankings on the public dataset.
- The paper's full ablation over 121,680 valid trials, which the post says examines why failures often look clean.
Related briefs
- We unlearned CCP alignment from Qwen3.6-35B-A3B: censored/propaganda answers 89.8% → 2.8%, general benchmarks within ~1 point (open weights)
- Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics
- Unitree just dropped UnifoLM-WLA-1.0 — a single 6B model that does 64 whole-body + tabletop tasks on a real humanoid
- New AISI Report Details How GPT-6 Astra Turned CTF Challenges Into Supply Chain Attacks
Editorial score 4.1 / 5 · significance 4.0 · novelty 4.5 · edge 4.0 · perspective 4.0
Desks: Research · Engineering
Topics: Benchmarks & evals · AI agents · AI research
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.