The Extended Brief
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
Brief by The AI News AI newsroom · Jul 30, 2026, 12:02 PM EDT edition
Original reporting by Import AI (Jack Clark) — Jack Clark · published Jul 27, 2026, 9:30 AM EDT
New benchmark data shows frontier models can now autonomously complete multi-week human programming tasks in hours, redefining the economic viability of long-horizon agentic coding.
Key points
- Opus 4.7 completed a complex programming task in 14 hours for $251, replacing up to 17 human weeks.
- Epoch and METR released MirrorCode to evaluate AI performance on extended software reimplementation tasks without source code.
- AI models achieved perfect scores on 17 of 25 target programs, including large codebases up to 87000 lines.
- Leading models from a year ago scored only 30 percent on simpler programs like calendar utilities.
From the source
“The findings are already very striking; Opus 4.7 solved a task in 14 hours for $251 in inference cost which METR and Epoch believe would take a human 2-17 weeks to do.”
“In our results, 8/25 target programs were never solved to a 100% threshold, and 4/25 were never solved to a 99% threshold,” they write.”
“May 2026: Opus 4.7 acting autonomously completes all the tasks but one in 9 minutes (and 35 seconds).”
“The robots achieve a 99.1% success rate, performing 778 successful folds across 9 garment types.”
“The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub.”
Practical applications
- Benchmark your own long-horizon engineering tasks against MirrorCode-style reimplementation work to see whether agentic coding is now cost-effective for them.
- Compare the reported $251-for-14-hours cost profile against what equivalent multi-week human effort costs your team before scoping the next large refactor or port.
- Track MirrorCode results across model releases as a leading indicator of when agents can take on your longer engineering projects.
Who should care
Engineering leaders and researchers tracking agent capability, because the benchmark suggests multi-week programming projects are moving into the range of hours of autonomous model work.
Context
Long-horizon benchmarks measure whether AI systems can sustain coherent work over tasks that take humans days or weeks, not minutes — historically a weak point for language models. MirrorCode, from Epoch and METR, tests this by asking models to reimplement software without seeing the source code. The reported jump is steep: models from a year earlier scored around 30 percent on simple programs like calendar utilities, while Opus 4.7 reportedly completed a task estimated at up to 17 human weeks in 14 hours for $251.
What to watch
- Independent replications of the MirrorCode results and whether the perfect scores on 17 of 25 programs hold up across labs and task selections.
- How quickly the remaining 8 programs, and larger codebases beyond 87,000 lines, fall to newer models.
Editorial score 4.0 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Research · Engineering · Tags: agents, benchmark, coding
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.