Cloud AI just got a serious competitor from your living room. AMD’s official guide demonstrates running the Kimi K2.5 — a one trillion open-source parameter AI model — across just four desktop PCs, completely offline, with no cloud subscription and no per-token charge.
Table of Contents

The Setup: 4 PCs Acting as One Giant AI Brain
| Component | Detail |
|---|---|
| Hardware | 4x Framework Desktop, AMD Ryzen AI Max+ 395, 128GB each |
| Total Memory Available | 480GB unified GPU memory (120GB per node) |
| AI Framework | AMD ROCm 7 |
| Inference Engine | llama.cpp RPC |
| OS | Ubuntu 24.04 LTS |
| Model | Kimi K2.5 Q2 quantised (375GB) |
| Network | 5Gbps Ethernet between nodes |
| Text Generation Speed | 9.45 tokens/sec with Flash Attention |
How It Works — Explained Simply
Think of it like splitting a massive jigsaw puzzle across four tables. Each PC handles a portion of the AI model’s layers — and a special protocol called RPC (Remote Procedure Call) stitches all four together so they behave like a single, very powerful computer. One machine acts as the “brain” that manages everything; the other three are workers that lend their memory and processing power.
The trick that makes this possible is AMD’s unified memory architecture on the Ryzen AI Max+ 395. Unlike traditional setups where CPU RAM and GPU VRAM are separate pools, this chip treats all 128GB as one shared pool. That means the GPU can access the full 128GB — and with a Linux kernel tweak (TTM modification), each node can unlock 120GB for GPU use instead of the BIOS-limited 96GB. Across four nodes, that’s 480GB of total AI memory.

AMD found that enabling Flash Attention via the rocWMMA library is the single biggest performance lever — boosting text generation from 3.46 to 8.30 tokens/second at long context lengths, and reducing time-to-first-token at a 4096-token prompt from 53.7 seconds down to 39.7 seconds. Larger batch sizes compound this further, delivering up to a 2x improvement in prompt processing vs baseline.
The result runs Kimi K2.5 — a model that rivals GPT-4 class performance on coding and reasoning — privately, on your own hardware. As TechnoSports has noted, this is a landmark moment for on-device AI: Kimi K2.5 is Moonshot AI’s most advanced open reasoning model, and it now runs at home.
FAQs
Can I run a trillion parameter AI model on a single PC?
Not yet at this scale — this guide uses four Framework Desktop PCs with Ryzen AI Max+ 395, each with 128GB unified memory, totalling 480GB across the cluster.
Is this setup suitable for everyday users or just developers?
AMD provides two setup paths — an easy one-click Lemonade SDK route and a manual build for developers, making it accessible to both audiences.





