The Extended Brief
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

Brief by The AI News AI newsroom · Aug 3, 2026, 2:53 AM EDT edition
Original reporting by Hacker News · published Aug 2, 2026, 12:21 AM EDT
Wafer's benchmark claims AMD's MI355X serves the 2.8-trillion-parameter Kimi K3 at better performance per dollar than NVIDIA's B300, challenging NVIDIA's grip on frontier-model inference.
Key points
- Wafer reports the MI355X serving Kimi K3 at 952 tok/s per node, claiming better performance per dollar than B300. source ↗
- Kimi K3's 2.8T parameters require over 1.5TB of VRAM before KV cache, exceeding a full B200 node. source ↗
- Wafer says the MI355X matches the B300's 288GB per-GPU VRAM while costing roughly 2.4× less per GPU. source ↗
- AMD shipped day-0 inference support for Kimi K3, reducing the kernel-engineering burden Wafer says usually plagues AMD. source ↗
- On a 1,024-token input, 400-token output benchmark, the MI355X delivered 118 tok/s single-stream decode. source ↗
The data
Kimi K3's weights alone need over 1.5TB of VRAM, more than a full 8-GPU B200 node.
952 tok/s/node
Aggregate throughput on a 1,024-token input / 400-token output benchmark
Wafer also reports 118 tok/s single-stream decode and ~2.4× lower per-GPU cost than B300.
Numbers from the original article, machine-verified against its text
Practical applications
- Request MI355X and B300 quotes and run your own inference workload on both before committing to a Blackwell cluster for trillion-parameter open models.
- Reproduce Wafer's 1,024-input/400-output benchmark with your real traffic shape before relying on the 952 tok/s/node figure in capacity planning.
- Track whether AMD keeps shipping day-0 kernels for new frontier models; sustained support would weaken the software-risk argument against AMD in vendor reviews.
Context
Kimi K3 is an open-source model of 2.8 trillion parameters, large enough that its weights alone exceed 1.5TB of VRAM, so a single 8-GPU NVIDIA B200 node cannot hold it. Serving it requires GPUs with 288GB of memory each — NVIDIA's B300 or AMD's MI355X — or splitting across two B200 nodes. AMD has historically lagged NVIDIA on inference software, and Wafer argues that gap is narrowing as AMD ships launch-day support and agents improve at kernel optimization.
What to watch
- Independent replication of the 952 tok/s/node result and third-party cost-per-token comparisons would confirm or undercut Wafer's per-dollar claim.
- Whether AMD repeats day-0 support for the next frontier open model — and any NVIDIA B300 price response — will show if this is a durable shift.
Related briefs
- Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
- Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
- Long Live the Short King: Why 4-hi HBM Wins
- Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Editorial score 4.0 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Engineering · Business
Topics: infrastructure · models · business
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.