The Extended Brief
Measuring LLMs’ Ability to Perform Cryptanalysis
Brief by The AI News AI newsroom · Jul 31, 2026, 1:22 PM EDT edition
Original reporting by Schneier on Security — Bruce Schneier · published Jul 28, 2026, 9:47 PM EDT
Frontier models are now discovering novel mathematical breaks in NIST cryptographic candidates, signaling an imminent shift in how security primitives are evaluated and deployed.
Key points
- Researchers introduced CryptanalysisBench, revealing that five frontier AI models successfully discovered novel, previously unknown cryptographic attacks.
- Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2 broke up to 86% of Tier-1 schemes.
- The benchmark evaluates 191 tasks across six cryptographic primitive families from four NIST standardization competitions.
- The models identified a key-recovery flaw in SpoC AEAD and an error in KINDI's security proof.
From the source
“Anthropic used the benchmark to test Mythos Preview, and found new vulnerabilities in Hawk and reduced-round AES.”
“The idea is to benchmark the ability of LLMs to discover new mathematical cryptanalytic attacks against a series of historical algorithms.”
“We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions.”
“We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment.”
Practical applications
- Check whether any system you maintain uses SpoC AEAD or relies on KINDI, given the reported key-recovery flaw in the former and the error found in the latter's security proof.
- Run frontier models against your own in-house cryptographic designs before deployment, since off-the-shelf models found novel attacks the human review process missed.
- Security teams evaluating NIST competition candidates should add LLM-assisted cryptanalysis to their review workflow alongside traditional peer review.
Who should care
Cryptographers, security engineers deploying NIST-candidate primitives, and standards bodies whose review processes assumed only human analysts could find novel mathematical attacks.
Context
Cryptanalysis — finding mathematical weaknesses in ciphers and related primitives — has historically been the domain of a small pool of expert humans, and standardization processes like NIST's competitions depend on that scrutiny. CryptanalysisBench tests whether LLMs can do this work across 191 tasks spanning six primitive families from four NIST competitions. Five frontier models — Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2 — broke up to 86 percent of Tier-1 schemes and discovered genuinely new attacks, including a key-recovery flaw in the SpoC AEAD scheme and an error in KINDI's security proof, marking a shift from models reproducing known attacks to finding unknown ones.
What to watch
- Confirmation and response from the SpoC and KINDI teams on the reported key-recovery flaw and proof error.
- Whether NIST or similar bodies formally incorporate LLM-driven cryptanalysis into future standardization rounds.
Editorial score 4.4 / 5 · significance 4.5 · novelty 4.5 · edge 4.5 · perspective 4.0
Desks: Security · Research · Tags: security, research, models
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.