Topic Archive
Chips & compute news and analysis
Every published brief tagged Chips & compute, newest first. Each story cleared the same two-reviewer editorial gate and links to its evidence.
TechNode · Sep 18, 2026, 10:32 AM EDT
Huawei sets commercial launch dates for Ascend 950 AI cluster cloud service
Teams training large models in China can rent a 1,024-card Huawei Ascend cluster from Sept. 30, with global access following Nov. 30.
SemiAnalysis (Dylan Patel) · Sep 18, 2026, 2:02 AM EDT
Everyone Says Datacenter Moratoriums Are Killing the US Buildout. We disagree
SemiAnalysis's satellite-tracked model says US datacenter capacity will more than double in 2027 despite four states and 300-plus localities moving to block new builds.
PYMNTS — AI · Sep 14, 2026, 3:14 PM EDT
Anthropic Makes $13.7 Compute Deal With Trump-Linked Rum Group
Anthropic is reportedly paying $13.7 billion to lease compute from a Trump-linked neocloud whose Georgia data center isn't built — or financed — yet.
SemiAnalysis (Dylan Patel) · Sep 13, 2026, 3:42 PM EDT
Long Live the Short King: Why 4-hi HBM Wins
A shift toward shorter HBM stacks could lower inference cost per token and ease the DRAM shortage that AI memory demand has created.
SemiAnalysis (Dylan Patel) · Sep 11, 2026, 3:11 PM EDT
Nvidia’s Backstop Universe – Heads I Win, Tails Who Loses?
Nvidia now backstops $530B of the AI buildout off its balance sheet, so a demand shortfall would land partly on its own books.
Constellation Research · Sep 8, 2026, 11:21 AM EDT
Qualcomm lands AWS deal for custom chips, interconnects
AWS committing to co-develop custom inference silicon with Qualcomm gives the mobile-chip maker a hyperscaler foothold against Nvidia and AMD in AI data centers.
Tech Funding News · Sep 8, 2026, 5:11 AM EDT
Samsung leads Mistral’s €3B round, pushing Europe’s AI champion to a €21B valuation
Europe's open-weight champion now has €3 billion and Samsung as an owner, giving enterprises a better-funded alternative to US frontier labs.
SemiAnalysis (Dylan Patel) · Sep 7, 2026, 5:12 PM EDT
TPU Inference Externalization Full Steam Ahead - InferenceX
Third-party benchmarks showing Google's TPUv7 beating NVIDIA's flagship chips on inference cost give AI teams a credible second supplier for serving models.
Tech Funding News · Sep 7, 2026, 6:11 AM EDT
Nscale seeks up to $3.5B in pre-IPO financing as AI infrastructure race intensifies
Nscale's planned $3.5 billion raise shows AI cloud providers are leaning on enormous contracted backlogs — and illustrative projections — to justify IPO valuations far above actual revenue.
PYMNTS — AI · Sep 6, 2026, 8:11 PM EDT
ByteDance Lands $29.6 Billion Loan to Fuel AI Advances
A reported unsecured $29.6 billion loan gives ByteDance fresh capital to build AI data centers outside China, sharpening competition with hyperscalers.
The Decoder · Sep 4, 2026, 11:12 AM EDT
Deepseek plans the largest known Huawei chip cluster with 160,000 processors in Inner Mongolia
A 160,000-chip DeepSeek inference cluster would be the largest known Huawei deployment, but supply bottlenecks could delay it past a year.
CIO · Sep 3, 2026, 5:12 PM EDT
What Nvidia’s $13B acquisition of Hugging Face means for AI model choice
Teams that build on Hugging Face have no full open alternative if its new owner, Nvidia, ever tilts the platform toward its own hardware.
Ars Technica AI · Aug 27, 2026, 3:20 PM EDT
AI industry says Trump plans to tax chips in the “single dumbest way imaginable”
Proposed US tariffs could tax not just chips but servers and other products built with them, raising hardware costs across the AI industry within weeks or months.
PYMNTS — AI · Aug 26, 2026, 8:11 PM EDT
Anthropic Escalates Cloud Spending With $45 Billion Nscale Agreement
Anthropic has now reportedly committed roughly $150 billion to rented compute, including $45 billion to Nscale for a site that won't operate until late 2027.
SemiAnalysis (Dylan Patel) · Aug 25, 2026, 11:43 AM EDT
OpenAI Jalapeño: Better Than Nvidia Blackwell
OpenAI's first custom chip beat every Nvidia, AMD, and Google chip in the authors' benchmarks, a result that could erode Nvidia's grip on inference economics.
CIO · Aug 24, 2026, 9:32 PM EDT
Nvidia to hike prices by 15%, on top of an even larger increase in July
Budgets for AI infrastructure face another hit: a reported 15% Nvidia server price increase for early-2027 deliveries, on top of July's 30% hikes.
Tech Funding News · Aug 24, 2026, 6:13 AM EDT
Nvidia reportedly eyes Perplexity at a $30B+ valuation as AI search becomes an agent business
Nvidia doubling down on Perplexity at a reported $30B-plus valuation would price the AI search startup near 40x revenue, setting a benchmark for agent businesses.
Latent Space · Aug 21, 2026, 2:43 AM EDT
Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud
Poolside's talent and technology moving into NVIDIA shows frontier-model startups can no longer raise compute capital fast enough to stay independent.
SemiAnalysis (Dylan Patel) · Aug 18, 2026, 10:12 PM EDT
Cerebras's Next Generation CS-4: Fast Just Got Faster
Cerebras customers get roughly double the inference throughput per wafer at about the same hardware cost, doubling potential token revenue without new silicon.
PYMNTS — AI · Aug 13, 2026, 12:22 PM EDT
Anthropic Pursues $6 Billion Decart Deal to Cut AI Costs
Anthropic's reported $6 billion bid for Decart would give the AI lab in-house technology to cut training and operating costs as demand for its software surges.
SiliconANGLE — AI · Aug 12, 2026, 6:21 AM EDT
SK hynix approves $38B+ investment in two new memory fabs
A major data-center memory supplier is spending $38.3 billion on two new fabs, adding future manufacturing capacity that data-center memory buyers depend on.
CIO · Aug 12, 2026, 12:12 AM EDT
Nvidia’s half-trillion-dollar AI investment fund could impact enterprise chip pricing, availability
Analysts warn Nvidia's $500B-plus financing fund could push enterprise AI infrastructure costs higher and deepen the data-center chip shortage.
SemiAnalysis (Dylan Patel) · Aug 7, 2026, 5:21 PM EDT
SpaceX 10GW in 2027 – Why It’s Real, Will Drive $300B ARR for SpaceX, and Why Microsoft Will Be the Largest Offtaker
SpaceX aims to build about 10GW of AI datacenter capacity by end-2027, putting its spending on par with AWS and Google.
Tech Funding News · Aug 4, 2026, 3:41 PM EDT
A high school dropout’s nuclear startup just landed $1B from Sequoia at a $6B valuation
Soaring AI power demand just made a two-year-old nuclear startup worth $6 billion, signaling investors now treat reactors as core AI infrastructure.
PYMNTS — AI · Aug 4, 2026, 1:11 PM EDT
Volta Exits Stealth at $2.4 Billion Valuation to Build AI Infrastructure
A startup just raised at a $2.4 billion valuation solely to finance and run dedicated AI data centers, with Anthropic reportedly signed as a $10 billion anchor customer.
Hacker News · Aug 3, 2026, 2:53 AM EDT
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Wafer's benchmark claims AMD's MI355X serves the 2.8-trillion-parameter Kimi K3 at better performance per dollar than NVIDIA's B300, challenging NVIDIA's grip on frontier-model inference.
CIO · Aug 3, 2026, 2:23 AM EDT
Data center backlash could slow CIOs’ AI plans
State moratoriums and new power tariffs are raising US data center costs, forcing CIOs to recalculate the economics of planned AI deployments.
Hacker News · Aug 1, 2026, 12:02 AM EDT
Everyone is building LLM routers, we deprecated ours
A gateway vendor killed its own LLM router after real-world use, undercutting the cost-saving promise driving the current model-routing hype.
Vercel Blog · Jul 30, 2026, 5:42 PM EDT
WebSocket support for OpenAI Responses API live on AI Gateway
Vercel's WebSocket support for the OpenAI Responses API cuts latency and token costs by up to 40% for complex, multi-step agentic workflows.
NVIDIA Developer Blog · Jul 30, 2026, 5:12 PM EDT
NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure
Misconfiguring kernel settings on new Blackwell or H100 clusters can silently waste up to 12% of your compute throughput, costing millions at scale.
NVIDIA Developer Blog · Jul 30, 2026, 5:12 PM EDT
ModelExpress: Distributing Model Artifacts at the Speed of Light
Distributing terabyte-scale model weights for RL post-training and autoscaling introduces massive I/O bottlenecks that new distribution techniques can now bypass.
r/LocalLLaMA · Jul 30, 2026, 5:10 PM EDT
4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s
Prosumer hardware can now run 122B parameter MoE models at interactive speeds by strategically spilling layers to system RAM, redefining local inference economics.
Simon Willison · Jul 30, 2026, 12:02 PM EDT
An Inside Look at the Relay Market Powering Token Resellers and Fraud
Exposed LLM endpoints are actively targeted by sophisticated relay networks for token arbitrage and model distillation, making strict API spend caps mandatory for public deployments.