AMD and Kimi AI

AMD and Kimi AI Just Made Coding Models Faster Than Nvidia’s B200

AMD's latest GPUs are now outrunning Nvidia on real AI coding workloads — and the numbers come from a fully public benchmark. AMD has published new results showing its Instinct…

July 23, 2026
3 min read

AMD’s latest GPUs are now outrunning Nvidia on real AI coding workloads — and the numbers come from a fully public benchmark.

AMD has published new results showing its Instinct MI355X GPUs now deliver lower latency and higher throughput than Nvidia’s single-node B200 when serving Moonshot AI’s Kimi K2.5, K2.6, and K2.7-Code models — the same coding-focused AI models developers are increasingly using for real software engineering tasks. The improvements come from AMD’s own ATOM inference engine paired with a set of kernel-level optimizations built with the open-source community.

This is a meaningful data point in the broader AI infrastructure race TechnoSports has been following closely, where the real competition increasingly plays out not just in raw chip specs, but in how efficiently that hardware serves the massive coding and reasoning models developers actually use.

The Headline Numbers

MetricResult
Peak throughput5,369.6 tokens/sec per GPU (at concurrency 128)
Sustained speed at peak load19.2 tokens/sec per user
Interactive-mode speed116.4 tokens/sec per user (at concurrency 4)
ComparisonOutperforms single-node Nvidia B200 across multiple concurrency levels
Hardware usedSingle AMD Instinct MI355X node, 4-way tensor parallelism

How AMD Got There

The gains didn’t come from one single trick — they’re the result of stacking several targeted optimizations across AMD’s serving stack. The core of the work rebuilds how the Kimi models’ “Mixture of Experts” architecture runs on AMD hardware, fusing several processing steps together so data moves through the GPU with less overhead. AMD also tuned how GPUs communicate with each other during each response step, and adjusted how the system schedules incoming requests so heavy traffic doesn’t slow down users already mid-conversation.

Notably, AMD says a chunk of this work was developed together with outside contributors through an open collaboration program, with the resulting code merged into AMD’s public AITER and ATOM repositories rather than kept proprietary.

Why This Matters Beyond Benchmarks

For any company or developer running Kimi’s coding models at scale, faster inference directly translates to lower cost per response and snappier AI coding assistants. Kimi K2.6 already scores competitively against top closed models on real-world coding benchmarks, so cheaper, faster serving on AMD hardware could make it a more attractive option for companies weighing GPU choices for their AI infrastructure — an increasingly important decision as AI compute costs climb across the industry.

Bottom Line

This isn’t a one-off marketing claim — the results were measured on InferenceX, a continuously running public benchmark, meaning the numbers get re-verified over time rather than reported once and forgotten. For AMD, it’s another sign that its GPUs are becoming a genuine alternative to Nvidia for serving today’s largest AI coding models, not just training them.


Based on AMD’s official technical blog post, published July 21, 2026. Results reflect specific test configurations and may vary by workload and system setup.

Follow us on Google News Get real-time updates & exclusive tech coverage
Follow

Leave a Reply

Your email address will not be published. Required fields are marked *

wp_enqueue_script('jquery', false, [], false, true); // load in footer