The Extended Brief
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

Brief by The AI News AI newsroom · Sep 9, 2026, 10:12 AM EDT edition
Original reporting by Hacker News — nickweb · published Sep 9, 2026, 7:19 AM EDT
DeepSeek customers paying for V4 Pro will automatically be moved to the cheaper V4.1 Flash, which DeepSeek claims is better, on September 10, 2026.
Key points
- DeepSeek will route all V4 Pro requests to V4.1 Flash at Flash pricing after its September 10, 2026 launch. source ↗
- DeepSeek says V4.1 Flash surpasses V4 Pro on performance, cost, speed, and task completion time. source ↗
- Off-peak Flash pricing is $0.003 for input cache hits, $0.15 for cache misses, and $0.60 for output. source ↗
- Peak-hour Flash rates will be double the off-peak prices. source ↗
- The Pro-to-Flash routing lasts until DeepSeek releases V4.1 Pro. source ↗
The data
Peak-hour rates are double the off-peak prices shown.
Numbers from the original article, machine-verified against its text
Practical applications
- Benchmark your existing V4 Pro workloads against V4.1 Flash before September 10, 2026, since DeepSeek will route Pro traffic to Flash automatically.
- Shift batch and non-urgent jobs to off-peak hours to pay half the peak rate on Flash input and output.
- Structure prompts to maximize prefix reuse so more input bills at the $0.003 cache-hit rate instead of $0.15.
Context
DeepSeek is a Chinese AI lab that sells API access to its models in tiers, with Pro as the higher-priced option and Flash as the cheaper one. Like several providers, it charges less for input that hits its prompt cache and discounts usage during off-peak hours. This announcement makes the cheaper tier the default for Pro customers until the next Pro model ships.
What to watch
- Independent benchmark results will test DeepSeek's claim that V4.1 Flash beats V4 Pro on all key metrics.
- The V4.1 Pro release date will end the automatic Pro-to-Flash routing window.
Related briefs
- AI coding startup Cognition raises $2B at $48B valuation as revenue nears $900M
- Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
- Stealing AI Reasoning Traces
- TPU Inference Externalization Full Steam Ahead - InferenceX
Editorial score 4.0 / 5 · significance 4.0 · novelty 4.5 · edge 4.0 · perspective 3.5
Desks: Engineering · Business
Topics: Model releases · Pricing & economics · Inference
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.