The Extended Brief
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
Brief by The AI News AI newsroom · Jul 31, 2026, 4:44 PM EDT edition
Original reporting by The Decoder — Thomas Joos · published Jul 31, 2026, 12:39 PM EDT
DeepSeek's latest Flash update delivers near-frontier performance at a fraction of the cost of comparable OpenAI models, shifting the cost-performance baseline for high-volume API routing.
Key points
- Deepseek V4 Flash matches OpenAI GPT-5.6 Luna performance at roughly sixty percent lower cost.
- The 0731 update increased the model's Artificial Analysis Intelligence Index score by ten points.
- Deepseek V4 Flash reached a final intelligence index score of fifty following this update.
- The updated model trails the OpenAI system by just one point on the intelligence index.
From the source
“According to the Artificial Analysis Intelligence Index , the new version scores 50 points, ten more than the previous V4 Flash that launched in April 2026.”
“That puts it just one point behind OpenAI's budget model GPT-5.6 Luna, but it costs about 60 percent less per task, even after OpenAI's 80 percent price cut .”
“A big reason for the gap is Deepseek's 98 percent cache discount , well above the industry-standard 90 percent.”
“On GDPval, a benchmark designed to test models on complex real-world office work , it climbs from 1,189 to 1,559 Elo points.”
“The architecture stays the same: 284 billion total parameters, 13 billion active, with a one-million-token context window.”
Practical applications
- Teams with high-volume API workloads can benchmark DeepSeek V4 Flash 0731 against GPT-5.6 Luna on their own tasks to validate the cost-performance claim.
- Engineers running model routers can add the updated Flash as a cheap tier for tasks that previously required a frontier model.
- Finance and procurement owners can re-run vendor cost models given a roughly 60 percent lower cost per task at near-parity performance.
Who should care
Engineers and engineering managers routing high-volume LLM traffic, and business leads whose API spend depends on where the cost-performance frontier sits.
Context
Aggregate benchmarks like the Artificial Analysis Intelligence Index are widely used to compare models on a single capability score, and cost per task at a given score is a key routing input. DeepSeek's budget V4 Flash model jumped ten points to 50 on the index with its 0731 update, landing one point behind OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost per task. DeepSeek has repeatedly pressured pricing from the low end, and near-parity at a fraction of the cost shifts the default calculus for high-volume API deployments.
What to watch
- Whether OpenAI responds with price cuts or a cheaper tier as budget models close the intelligence-index gap.
- Task-level evaluations beyond the aggregate index confirming parity holds on real workloads such as coding, extraction, or reasoning.
Editorial score 3.8 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 3.0
Desks: Engineering · Business · Tags: models, business
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.