The Extended Brief
Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Brief by The AI News AI newsroom · Aug 3, 2026, 2:53 AM EDT edition
Original reporting by Hacker News · published Aug 2, 2026, 12:21 AM EDT
Wafer's benchmark claims AMD's MI355X serves the 2.8-trillion-parameter Kimi K3 at better performance per dollar than NVIDIA's B300, challenging NVIDIA's grip on frontier-model inference.
Key points
- Wafer reports the MI355X serving Kimi K3 at 952 tok/s per node, claiming better performance per dollar than B300.
- Kimi K3's 2.8T parameters require over 1.5TB of VRAM before KV cache, exceeding a full B200 node.
- Wafer says the MI355X matches the B300's 288GB per-GPU VRAM while costing roughly 2.4× less per GPU.
- AMD shipped day-0 inference support for Kimi K3, reducing the kernel-engineering burden Wafer says usually plagues AMD.
- On a 1,024-token input, 400-token output benchmark, the MI355X delivered 118 tok/s single-stream decode.
The data
Kimi K3's weights alone need over 1.5TB of VRAM, more than a full 8-GPU B200 node.
952 tok/s/node
Aggregate throughput on a 1,024-token input / 400-token output benchmark
Wafer also reports 118 tok/s single-stream decode and ~2.4× lower per-GPU cost than B300.
Numbers from the original article, machine-verified against its text
From the source
“GLM5.2 has 753B parameters, DeepSeek V4-Pro 1.6T, and Kimi K3 weighs in at 2.8T (!!) parameters.”
“That's over 1.5TB of VRAM before allocating a KV cache for 1M tokens of context.”
“At around ~2.4× cheaper per GPU on average versus a B300 and ~1.7× cheaper than a B200, the MI355X is a cost-efficient alternative to Blackwells with comparable hardware specs.”
“But with AMD shipping day-0 support for Kimi K3, most of the work was already done for us.”
Practical applications
- Request MI355X and B300 quotes and run your own inference workload on both before committing to a Blackwell cluster for trillion-parameter open models.
- Reproduce Wafer's 1,024-input/400-output benchmark with your real traffic shape before relying on the 952 tok/s/node figure in capacity planning.
- Track whether AMD keeps shipping day-0 kernels for new frontier models; sustained support would weaken the software-risk argument against AMD in vendor reviews.
Who should care
Inference infrastructure and procurement teams weighing AMD MI355X against NVIDIA B300/B200 for serving multi-trillion-parameter open models, and cost-sensitive labs betting on open-source frontier models.
Context
Kimi K3 is an open-source model of 2.8 trillion parameters, large enough that its weights alone exceed 1.5TB of VRAM, so a single 8-GPU NVIDIA B200 node cannot hold it. Serving it requires GPUs with 288GB of memory each — NVIDIA's B300 or AMD's MI355X — or splitting across two B200 nodes. AMD has historically lagged NVIDIA on inference software, and Wafer argues that gap is narrowing as AMD ships launch-day support and agents improve at kernel optimization.
What to watch
- Independent replication of the 952 tok/s/node result and third-party cost-per-token comparisons would confirm or undercut Wafer's per-dollar claim.
- Whether AMD repeats day-0 support for the next frontier open model — and any NVIDIA B300 price response — will show if this is a durable shift.
Editorial score 4.0 / 5 · significance 4.0 · novelty 4.0 · edge 4.0 · perspective 4.0
Desks: Engineering · Business · Tags: infrastructure, models, business
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.