The leak was right, and the reality turned out even bigger than expected. Moonshot AI has officially released Kimi K3, a 2.8-trillion-parameter model the company calls the world’s largest open-weight AI system, and early benchmarks suggest it’s landed remarkably close to the very top of the field.
Kimi K3 Is Officially Here: What Moonshot Actually Confirmed
K3 arrives with a 1-million-token context window, native visual understanding, and an always-on “thinking mode.” Under the hood, Moonshot built it around two architectural innovations developed in-house: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, a replacement for standard residual connections designed to keep delivering gains as the model scales. Both techniques were previously published as open research, so this wasn’t an invention Moonshot kept entirely under wraps.

| Detail | Confirmed Specification |
|---|---|
| Parameters | 2.8 trillion (largest open-weight model to date) |
| Context window | 1 million tokens |
| Architecture | Kimi Delta Attention + Attention Residuals |
| Reasoning mode | Always-on “thinking mode” |
| Vision | Native visual understanding |
| API compatibility | OpenAI SDK compatible |
| Full weights release | July 27, 2026 |
| Third-party ranking | 1st on web interface building, 2nd overall (Arena.ai) |
| Beats on benchmarks | Claude Opus 4.8, GPT-5.5 |
| Still trails | Claude Fable 5, GPT-5.6 Sol |
An Honest Launch Line, for Once
What’s notable is how Moonshot framed its own model. Rather than overselling, the company’s own launch messaging stated plainly that overall performance still trails Fable 5 and GPT-5.6 Sol, while consistently outperforming every other tested model. That’s an unusually candid admission in an industry known for cherry-picked benchmark charts, and outside evaluators have largely backed it up: Arena.ai ranked K3 first on web interface construction and second overall, ahead of GPT-5.6 Sol and just behind Fable 5.

On raw compute efficiency, Moonshot says K3 performed competitively against Fable 5 and clearly outperformed both Opus 4.8 and GPT-5.5 on GPU kernel optimisation, the technical work of squeezing more throughput out of the same hardware.
Why the Timing and the Reaction Matter
The release lands just ahead of the 2026 World Artificial Intelligence Conference in Shanghai, and it’s already moved markets: shares in rival Chinese labs Zhipu and MiniMax fell sharply in Hong Kong following the announcement, as investors reassessed where the real competitive edge sits. Bank of America analysts noted that K3 shows pre-training scale paired with architectural innovation can still deliver meaningful gains for Chinese labs, even under persistent hardware and compute constraints.
This continues a pattern we flagged in our earlier coverage of the Kimi K3 launch leak and its rumoured specs, and fits the broader trend we’ve tracked around Chinese AI labs closing the performance gap with US frontier models faster than most expected this year.
The Takeaway
Kimi K3 is real, it’s officially the largest open-weight model in the world, and it lands closer to Fable 5 and GPT-5.6 Sol than any Chinese open model has managed before. Full weights land July 27, which is when independent developers will get to properly stress-test whether the benchmarks hold up outside Moonshot’s own evaluation suite.





