Topic Archive
Chips & compute news and analysis
Every published brief tagged Chips & compute, newest first. Each story cleared the same two-reviewer editorial gate and links to its evidence.
Tech Funding News · Aug 4, 2026, 3:41 PM EDT
A high school dropout’s nuclear startup just landed $1B from Sequoia at a $6B valuation
Soaring AI power demand just made a two-year-old nuclear startup worth $6 billion, signaling investors now treat reactors as core AI infrastructure.
PYMNTS — AI · Aug 4, 2026, 1:11 PM EDT
Volta Exits Stealth at $2.4 Billion Valuation to Build AI Infrastructure
A startup just raised at a $2.4 billion valuation solely to finance and run dedicated AI data centers, with Anthropic reportedly signed as a $10 billion anchor customer.
Hacker News · Aug 3, 2026, 2:53 AM EDT
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Wafer's benchmark claims AMD's MI355X serves the 2.8-trillion-parameter Kimi K3 at better performance per dollar than NVIDIA's B300, challenging NVIDIA's grip on frontier-model inference.
CIO · Aug 3, 2026, 2:23 AM EDT
Data center backlash could slow CIOs’ AI plans
State moratoriums and new power tariffs are raising US data center costs, forcing CIOs to recalculate the economics of planned AI deployments.
Hacker News · Aug 1, 2026, 12:02 AM EDT
Everyone is building LLM routers, we deprecated ours
A gateway vendor killed its own LLM router after real-world use, undercutting the cost-saving promise driving the current model-routing hype.
Vercel Blog · Jul 30, 2026, 5:42 PM EDT
WebSocket support for OpenAI Responses API live on AI Gateway
Vercel's WebSocket support for the OpenAI Responses API cuts latency and token costs by up to 40% for complex, multi-step agentic workflows.
NVIDIA Developer Blog · Jul 30, 2026, 5:12 PM EDT
NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure
Misconfiguring kernel settings on new Blackwell or H100 clusters can silently waste up to 12% of your compute throughput, costing millions at scale.
NVIDIA Developer Blog · Jul 30, 2026, 5:12 PM EDT
ModelExpress: Distributing Model Artifacts at the Speed of Light
Distributing terabyte-scale model weights for RL post-training and autoscaling introduces massive I/O bottlenecks that new distribution techniques can now bypass.
r/LocalLLaMA · Jul 30, 2026, 5:10 PM EDT
4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s
Prosumer hardware can now run 122B parameter MoE models at interactive speeds by strategically spilling layers to system RAM, redefining local inference economics.
Simon Willison · Jul 30, 2026, 12:02 PM EDT
An Inside Look at the Relay Market Powering Token Resellers and Fraud
Exposed LLM endpoints are actively targeted by sophisticated relay networks for token arbitrage and model distillation, making strict API spend caps mandatory for public deployments.