Topic Archive
Inference news and analysis
Every published brief tagged Inference, newest first. Each story cleared the same two-reviewer editorial gate and links to its evidence.
SemiAnalysis (Dylan Patel) · Sep 13, 2026, 3:42 PM EDT
Long Live the Short King: Why 4-hi HBM Wins
A shift toward shorter HBM stacks could lower inference cost per token and ease the DRAM shortage that AI memory demand has created.
Hacker News · Sep 9, 2026, 10:12 AM EDT
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
DeepSeek customers paying for V4 Pro will automatically be moved to the cheaper V4.1 Flash, which DeepSeek claims is better, on September 10, 2026.
Constellation Research · Sep 8, 2026, 11:21 AM EDT
Qualcomm lands AWS deal for custom chips, interconnects
AWS committing to co-develop custom inference silicon with Qualcomm gives the mobile-chip maker a hyperscaler foothold against Nvidia and AMD in AI data centers.
SemiAnalysis (Dylan Patel) · Sep 7, 2026, 5:12 PM EDT
TPU Inference Externalization Full Steam Ahead - InferenceX
Third-party benchmarks showing Google's TPUv7 beating NVIDIA's flagship chips on inference cost give AI teams a credible second supplier for serving models.
The Decoder · Sep 4, 2026, 11:12 AM EDT
Deepseek plans the largest known Huawei chip cluster with 160,000 processors in Inner Mongolia
A 160,000-chip DeepSeek inference cluster would be the largest known Huawei deployment, but supply bottlenecks could delay it past a year.
SemiAnalysis (Dylan Patel) · Aug 25, 2026, 11:43 AM EDT
OpenAI Jalapeño: Better Than Nvidia Blackwell
OpenAI's first custom chip beat every Nvidia, AMD, and Google chip in the authors' benchmarks, a result that could erode Nvidia's grip on inference economics.
SemiAnalysis (Dylan Patel) · Aug 18, 2026, 10:12 PM EDT
Cerebras's Next Generation CS-4: Fast Just Got Faster
Cerebras customers get roughly double the inference throughput per wafer at about the same hardware cost, doubling potential token revenue without new silicon.
PYMNTS — AI · Aug 13, 2026, 12:22 PM EDT
Anthropic Pursues $6 Billion Decart Deal to Cut AI Costs
Anthropic's reported $6 billion bid for Decart would give the AI lab in-house technology to cut training and operating costs as demand for its software surges.
Vercel Blog · Aug 11, 2026, 8:11 PM EDT
DeepSeek overtakes Google on volume, cost per token falls 13.6%
Enterprise token routing is flipping to cheap open-weight models: DeepSeek now carries a quarter of this gateway's traffic, more than double Google, squeezing incumbents' pricing power.
Hugging Face · Jul 29, 2026, 10:05 PM EDT
LFM2.5-Encoders for Fast Long-Context Inference on CPU
Enables cost-effective long-context inference on CPUs, significantly reducing deployment costs for edge and high-throughput applications.