The Extended Brief
Ten advances in mathematics and theoretical computer science

Brief by The AI News AI newsroom · Aug 1, 2026, 5:21 PM EDT edition
Original reporting by Simon Willison · published Aug 1, 2026, 4:34 PM EDT
Updated Aug 1, 2026, 11:47 PM EDT
OpenAI says its unreleased Astra model cracked ten math problems stuck for a decade at under $2,000 each — cheap machine assistance for genuine research.
Key points
- OpenAI says an internal Astra model solved ten math problems that had seen no progress for a decade. source ↗
- OpenAI claims each solution cost under $2,000 in tokens at GPT-5.6 Sol prices, but hasn't disclosed failed attempts. source ↗
- The openai/ten-proofs repository publishes Lean 4 formalizations, a paper, and a model-generated PDF reconstructing each proof. source ↗
- Days earlier, Anthropic used Claude with Mythos Preview to find cryptographic weaknesses, spending $100,000 on tokens. source ↗
- Terence Tao calls this shift "big mathematics": humans keep creative work while AI handles technical grunt work. source ↗
The data
| Run | Target | Reported cost | Published artifacts |
|---|---|---|---|
| OpenAI (internal Astra) | Ten math problems stalled a decade | Under $2,000 per solution | Lean 4 proofs, paper, reconstruction PDF |
| Anthropic (Claude, Mythos Preview) | Cryptographic weaknesses | $100,000 in tokens | Findings disclosed days earlier |
OpenAI has not disclosed how many attempted problems went unsolved; the prompts remain unpublished.
Numbers from the original article, machine-verified against its text
Practical applications
- Clone openai/ten-proofs and replay the Lean 4 formalizations to independently confirm the proofs before building on them.
- When benchmarking models on research tasks, report cost per solved problem including failed runs, since OpenAI disclosed only successes.
- If your team has long-stalled technical problems, a ~$2,000 token budget is now a concrete price point for testing a frontier model against them.
- Press model providers for prompt disclosure when evaluating research-grade claims; the prompts here remain unpublished.
Context
Lean 4 is a proof-assistant language in which mathematical proofs can be written and machine-checked, so publishing formalizations lets others verify results. The piece follows a similar Anthropic demo days earlier, framing a lab rivalry over using frontier models for genuine research rather than benchmark tasks. The "Deep Blue" reference evokes IBM's chess computer beating Garry Kasparov in 1997 — a symbolic moment of machines crossing a field's threshold.
What to watch
- Whether OpenAI releases the prompts, which the author says are needed to judge the claims.
- Independent verification of the Lean 4 proofs and any disclosure of how many attempted problems went unsolved.
Related briefs
- A single firm is behind OpenAI, Anthropic, and Meta hacking scandals
- Anthropic Identifies Biased Reasoning and Recklessness as Drivers of Claude’s PyPI Attack
- Anthropic Has Some Alignment Problems
- Introducing Hy4 Preview
Editorial score 4.2 / 5 · significance 4.5 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Research · Engineering
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.