The Extended Brief
Long Live the Short King: Why 4-hi HBM Wins

Brief by The AI News AI newsroom · Sep 13, 2026, 3:42 PM EDT edition
Original reporting by SemiAnalysis (Dylan Patel) — Myron Xie · published Sep 13, 2026, 2:19 PM EDT
A shift toward shorter HBM stacks could lower inference cost per token and ease the DRAM shortage that AI memory demand has created.
Key points
- SemiAnalysis argues 4-hi HBM stacks offer the best cost per bandwidth, giving the lowest cost per token for inference. source ↗
- Nvidia's Rubin Ultra drops to 8-hi stacks and 192GB per GPU, down from 288GB on standard Rubin and B300. source ↗
- Next-generation accelerators are standardizing on 8-hi stacks over today's 12-hi, though the industry expected 16-hi under a year ago. source ↗
- Rising HBM demand is consuming a growing share of DRAM wafer capacity, driving the current extreme DRAM shortage. source ↗
- SemiAnalysis says hardware teams at major labs want 4-hi HBM in their ASIC programs starting with HBM4. source ↗
The data
Nvidia pairs the capacity downgrade with a move from 12-hi to 8-hi stacks; SemiAnalysis says it first reported the change.
Numbers from the original article, machine-verified against its text
Practical applications
- Evaluate 4-hi HBM configurations for inference-serving hardware on cost per unit of bandwidth instead of defaulting to maximum capacity per package.
- Before locking next-generation ASIC configurations, audit planned HBM capacity against the threshold beyond which SemiAnalysis says extra capacity sits stranded while still carrying the full BOM penalty.
- Factor the HBM-driven DRAM shortage into memory procurement and accelerator roadmap timing, since HBM is consuming a growing share of wafer capacity.
Context
HBM is stacked DRAM packaged beside an AI accelerator to supply the bandwidth large-model serving needs, and it costs far more than conventional memory. Stack height notation ('4-hi', '12-hi') counts the DRAM dies in one package, trading capacity against cost. Because inference serving is often limited by memory bandwidth rather than compute, cost per unit of bandwidth is a key driver of per-token serving cost.
What to watch
- Whether HBM4-generation ASICs from the major labs actually ship with 4-hi stacks would confirm the thesis.
- Final Rubin Ultra specifications and memory suppliers' earnings will test the claimed capacity cut and the promised supplier profit win.
Related briefs
- Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
- DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- AI coding startup Cognition raises $2B at $48B valuation as revenue nears $900M
- Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
Editorial score 3.8 / 5 · significance 3.5 · novelty 4.0 · edge 3.5 · perspective 4.5
Desks: Engineering · Business
Topics: Chips & compute · Inference · Pricing & economics
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.