The Extended Brief
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

Brief by The AI News AI newsroom · Jul 30, 2026, 12:02 PM EDT edition
Original reporting by Import AI (Jack Clark) — Jack Clark · published Jul 27, 2026, 9:30 AM EDT
Updated Aug 1, 2026, 11:28 AM EDT
New benchmark data shows frontier models can now autonomously complete multi-week human programming tasks in hours, redefining the economic viability of long-horizon agentic coding.
Key points
- Opus 4.7 completed a complex programming task in 14 hours for $251, replacing up to 17 human weeks. source ↗
- Epoch and METR released MirrorCode to evaluate AI performance on extended software reimplementation tasks without source code. source ↗
- AI models achieved perfect scores on 17 of 25 target programs, including large codebases up to 87000 lines. source ↗
- Leading models from a year ago scored only 30 percent on simpler programs like calendar utilities. source ↗
Practical applications
- Benchmark your own long-horizon engineering tasks against MirrorCode-style reimplementation work to see whether agentic coding is now cost-effective for them.
- Compare the reported $251-for-14-hours cost profile against what equivalent multi-week human effort costs your team before scoping the next large refactor or port.
- Track MirrorCode results across model releases as a leading indicator of when agents can take on your longer engineering projects.
Context
Long-horizon benchmarks measure whether AI systems can sustain coherent work over tasks that take humans days or weeks, not minutes — historically a weak point for language models. MirrorCode, from Epoch and METR, tests this by asking models to reimplement software without seeing the source code. The reported jump is steep: models from a year earlier scored around 30 percent on simple programs like calendar utilities, while Opus 4.7 reportedly completed a task estimated at up to 17 human weeks in 14 hours for $251.
What to watch
- Independent replications of the MirrorCode results and whether the perfect scores on 17 of 25 programs hold up across labs and task selections.
- How quickly the remaining 8 programs, and larger codebases beyond 87,000 lines, fall to newer models.
Related briefs
- Anthropic Identifies Biased Reasoning and Recklessness as Drivers of Claude’s PyPI Attack
- Anthropic Has Some Alignment Problems
- Introducing Hy4 Preview
- Generative design of novel bacteriophages with genome language models [R]
Editorial score 4.0 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Research · Engineering
Topics: agents · benchmark · coding
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.