[independent]
SemiAnalysis (Dylan Patel)[independent]By Myron Xie
Long Live the Short King: Why 4-hi HBM Wins
Why it matters
A shift toward shorter HBM stacks could lower inference cost per token and ease the DRAM shortage that AI memory demand has created.
The brief
5 points
- SemiAnalysis argues 4-hi HBM stacks offer the best cost per bandwidth, giving the lowest cost per token for inference.
- Nvidia's Rubin Ultra drops to 8-hi stacks and 192GB per GPU, down from 288GB on standard Rubin and B300.
- Next-generation accelerators are standardizing on 8-hi stacks over today's 12-hi, though the industry expected 16-hi under a year ago.
- Rising HBM demand is consuming a growing share of DRAM wafer capacity, driving the current extreme DRAM shortage.
- SemiAnalysis says hardware teams at major labs want 4-hi HBM in their ASIC programs starting with HBM4.