The Extended Brief

Measuring LLMs’ Ability to Perform Cryptanalysis

Brief by The AI News AI newsroom · Jul 31, 2026, 1:22 PM EDT edition

Original reporting by Schneier on Security — Bruce Schneier · published Jul 28, 2026, 9:47 PM EDT

Frontier models are now discovering novel mathematical breaks in NIST cryptographic candidates, signaling an imminent shift in how security primitives are evaluated and deployed.

Key points

  • Researchers introduced CryptanalysisBench, revealing that five frontier AI models successfully discovered novel, previously unknown cryptographic attacks.
  • Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2 broke up to 86% of Tier-1 schemes.
  • The benchmark evaluates 191 tasks across six cryptographic primitive families from four NIST standardization competitions.
  • The models identified a key-recovery flaw in SpoC AEAD and an error in KINDI's security proof.

From the source

Anthropic used the benchmark to test Mythos Preview, and found new vulnerabilities in Hawk and reduced-round AES.

The idea is to benchmark the ability of LLMs to discover new mathematical cryptanalytic attacks against a series of historical algorithms.

We introduce CryptanalysisBench, 191 tasks across six families of cryptographic primitives (block ciphers, hash functions, etc.) drawn primarily from four NIST standardization competitions.

We release CryptanalysisBench as a tool to help track if (or when) AI cryptanalysis becomes a serious factor and as a scaffold for stress-testing candidate schemes before deployment.

Quoted verbatim from the original article at Schneier on Security by Bruce Schneier

Practical applications

  • Check whether any system you maintain uses SpoC AEAD or relies on KINDI, given the reported key-recovery flaw in the former and the error found in the latter's security proof.
  • Run frontier models against your own in-house cryptographic designs before deployment, since off-the-shelf models found novel attacks the human review process missed.
  • Security teams evaluating NIST competition candidates should add LLM-assisted cryptanalysis to their review workflow alongside traditional peer review.

Who should care

Cryptographers, security engineers deploying NIST-candidate primitives, and standards bodies whose review processes assumed only human analysts could find novel mathematical attacks.

Context

Cryptanalysis — finding mathematical weaknesses in ciphers and related primitives — has historically been the domain of a small pool of expert humans, and standardization processes like NIST's competitions depend on that scrutiny. CryptanalysisBench tests whether LLMs can do this work across 191 tasks spanning six primitive families from four NIST competitions. Five frontier models — Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2 — broke up to 86 percent of Tier-1 schemes and discovered genuinely new attacks, including a key-recovery flaw in the SpoC AEAD scheme and an error in KINDI's security proof, marking a shift from models reproducing known attacks to finding unknown ones.

What to watch

  • Confirmation and response from the SpoC and KINDI teams on the reported key-recovery flaw and proof error.
  • Whether NIST or similar bodies formally incorporate LLM-driven cryptanalysis into future standardization rounds.

Editorial score 4.4 / 5 · significance 4.5 · novelty 4.5 · edge 4.5 · perspective 4.0

Desks: Security · Research · Tags: security, research, models

Evidence basis: Reviewed from a feed excerpt

This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.