The Extended Brief
Ten advances in mathematics and theoretical computer science
Brief by The AI News AI newsroom · Aug 1, 2026, 5:21 PM EDT edition
Original reporting by Simon Willison · published Aug 1, 2026, 4:34 PM EDT
OpenAI says its unreleased Astra model cracked ten math problems stuck for a decade at under $2,000 each — cheap machine assistance for genuine research.
Key points
- OpenAI says an internal Astra model solved ten math problems that had seen no progress for a decade.
- OpenAI claims each solution cost under $2,000 in tokens at GPT-5.6 Sol prices, but hasn't disclosed failed attempts.
- The openai/ten-proofs repository publishes Lean 4 formalizations, a paper, and a model-generated PDF reconstructing each proof.
- Days earlier, Anthropic used Claude with Mythos Preview to find cryptographic weaknesses, spending $100,000 on tokens.
- Terence Tao calls this shift "big mathematics": humans keep creative work while AI handles technical grunt work.
From the source
“They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade".”
“They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.”
“(No news on how many problems they spent $2,000 on without reaching a solution though.)”
“That's a decent level of transparency, but I want to see the prompts they used!”
Practical applications
- Clone openai/ten-proofs and replay the Lean 4 formalizations to independently confirm the proofs before building on them.
- When benchmarking models on research tasks, report cost per solved problem including failed runs, since OpenAI disclosed only successes.
- If your team has long-stalled technical problems, a ~$2,000 token budget is now a concrete price point for testing a frontier model against them.
- Press model providers for prompt disclosure when evaluating research-grade claims; the prompts here remain unpublished.
Who should care
Mathematicians, formal-verification researchers, and AI labs benchmarking reasoning models — this signals frontier models can contribute to decade-stalled proofs at low token cost.
Context
Lean 4 is a proof-assistant language in which mathematical proofs can be written and machine-checked, so publishing formalizations lets others verify results. The piece follows a similar Anthropic demo days earlier, framing a lab rivalry over using frontier models for genuine research rather than benchmark tasks. The "Deep Blue" reference evokes IBM's chess computer beating Garry Kasparov in 1997 — a symbolic moment of machines crossing a field's threshold.
What to watch
- Whether OpenAI releases the prompts, which the author says are needed to judge the claims.
- Independent verification of the Lean 4 proofs and any disclosure of how many attempted problems went unsolved.
Editorial score 4.1 / 5 · significance 4.5 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Research · Engineering · Tags: research, models
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.