The Extended Brief
Claude Haiku 5.5

Brief by The AI News AI newsroom · Oct 7, 2026, 3:01 PM EDT edition
Original reporting by Hacker News · published Oct 7, 2026, 2:01 PM EDT
Anthropic's fastest small model now runs about 75% cheaper than its predecessor, reshaping cost math for high-volume agent and classification workloads.
Key points
- Anthropic released Claude Haiku 5.5 on October 7, 2026, priced about 75% lower to run than Haiku 4.5. source ↗
- Anthropic calls Haiku 5.5 its fastest model, targeting high-volume work like summaries, classification, database queries, and browser use. source ↗
- Anthropic-reported benchmarks show Haiku 5.5 at 72.4% on OSWorld 2.1 computer use, up from Haiku 4.5's 15.7%. source ↗
- Anthropic halved Sonnet 5.5 cache-read prices, which it says makes Sonnet roughly 20% cheaper on most agentic work. source ↗
- Claude Max and Team subscribers receive a new monthly API credit for building agents on the Claude Platform. source ↗
The data
75%
Average cost reduction Anthropic claims for Haiku 5.5
Figure is Anthropic's own estimate of average running cost.
Scores reported by Anthropic in its launch materials.
Scores reported by Anthropic in its launch materials.
Numbers from the original article, machine-verified against its text
Practical applications
- Re-price existing Haiku 4.5 workloads against Haiku 5.5, since Anthropic says average running costs drop about 75%.
- Benchmark Haiku 5.5 on your own computer-use or browser tasks before migrating, where Anthropic reports a jump from 15.7% to 72.4% on OSWorld 2.1.
- Recalculate Sonnet 5.5 agentic pipeline costs with the halved cache-read pricing, which Anthropic says cuts most agentic bills around 20%.
- Test Haiku 5.5 as a delegated subagent under Sonnet 5.5 or Opus 5.5 for coding work, the pairing Anthropic recommends.
Context
Haiku is Anthropic's small, low-cost model tier, positioned below Sonnet and Opus for high-volume, cost-sensitive work. Anthropic promotes a multi-model pattern where Haiku runs as a subagent handling routine steps delegated by larger models. Cache reads are billed prompt-caching lookups, so halving their price mainly benefits agentic workloads that reuse large contexts.
What to watch
- Independent reproduction of Anthropic's self-reported scores, particularly the OSWorld 2.1 jump from 15.7% to 72.4%.
- Whether competing small models such as GPT-6 Luna, which Anthropic benchmarks against, respond on price or capability.
Related briefs
- How Much of Your Customer Support Can AI Really Resolve? The Best Get About 70%. The Median Is 48%. Here’s the Real Data From 13 Vendors
- Mistral Large 4
- Astra 6 Replaced a Core Engine of SaaStr Connect With Two Words, “DO IT,” Twice in Under an Hour. Then It Said “I Did Not Make That Edit.”
- Unitree just dropped UnifoLM-WLA-1.0 — a single 6B model that does 64 whole-body + tabletop tasks on a real humanoid
Editorial score 4.1 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 4.5
Desks: Engineering · Business
Topics: Model releases · Pricing & economics
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.