The Extended Brief
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

Brief by The AI News AI newsroom · Sep 10, 2026, 2:13 PM EDT edition
Original reporting by Hacker News · published Sep 10, 2026, 11:29 AM EDT
Cognition claims its SWE-2 coding model matches near-frontier rivals at a fraction of their price, which could sharply cut the cost of agentic coding workloads.
Key points
- Cognition says SWE-2 scores 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1 at 64% lower cost. source ↗
- SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter model already RL-trained for agentic coding. source ↗
- Cognition says its RL added five to six points over the Kimi K3 base on many benchmarks. source ↗
- A new RL algorithm trains all reasoning-effort levels in one run using per-level linear cost penalties. source ↗
- Cognition's own table shows SWE-2 at 27.3% on Terminal-Bench 4, far behind Fable 5.1's 55.8% and GPT-6 Astra's 57.9%. source ↗
The data
Scores are from Cognition's own announcement and not independently verified.
Cognition says its RL adds 5–6 points on many benchmarks over the base model.
Numbers from the original article, machine-verified against its text
Practical applications
- Benchmark SWE-2 against Fable 5.1 or GPT-6 Astra on your own repositories before renewing a frontier-model contract; Cognition claims near-parity at roughly a quarter of Astra's cost.
- Sweep SWE-2's reasoning-effort levels on a representative task set to find the cheapest tier that meets your quality bar, since the model was trained to improve every effort level at once.
- Keep long-horizon terminal workflows on another model until SWE-2's Terminal-Bench 4 gap (27.3% versus Fable 5.1's 55.8%) is tested on your tasks.
Context
Cognition's SWE line are coding models built by applying reinforcement learning on top of an existing large base model rather than training from scratch; SWE-2 starts from Kimi K3, a 2.8-trillion-parameter model already tuned for agentic coding. The 'Pareto frontier' framing concerns the tradeoff between benchmark capability and inference cost, and Cognition's central claim is that one RL run improved that tradeoff at every reasoning-effort level. The cited benchmarks — FrontierCode, DeepSWE, and Terminal-Bench — measure coding and terminal-task ability.
What to watch
- Independent results on the public FrontierCode leaderboard would confirm or undercut Cognition's self-reported scores.
- A price response from Fable or OpenAI, or a SWE-2 follow-up closing the Terminal-Bench 4 gap, would escalate the story.
Related briefs
- DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- AI coding startup Cognition raises $2B at $48B valuation as revenue nears $900M
- Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
- Stealing AI Reasoning Traces
Editorial score 4.0 / 5 · significance 3.5 · novelty 4.5 · edge 4.0 · perspective 4.5
Desks: Engineering · Business
Topics: Model releases · Developer tools · Pricing & economics
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.