Growing Pains: How Distributed AI Training Changes the Network Between Datacenters

The network linking distributed AI datacenters has become the toughest part of training a frontier model. A distributed AI training cluster combines GPU systems in separate buildings, cities, or cloud…

October 10, 2026
5 min read

The network linking distributed AI datacenters has become the toughest part of training a frontier model. A distributed AI training cluster combines GPU systems in separate buildings, cities, or cloud regions and makes them work like one machine. Keeping those systems in sync is the real headache.

Why Can’t a Single Datacenter Hold a Frontier Model Anymore?

Power and floor space now set the limits. Industry estimates reportedly suggest that training a modern model can require clusters with tens of thousands of GPUs, while analysts reportedly project that the largest individual frontier training runs could draw between 4 GW and 16 GW of power by 2030.

No single campus can reportedly take on that load within a realistic timeline, so operators spread the work across sites and connect them. Dividing the calculations is easy. Making that division invisible to the training job is much harder — and that’s where the network takes centre stage.

What Actually Changes in the Network Between Distributed AI Datacenters?

Interconnects become far more demanding once model parameters reach the hundreds of billions. Training needs high-bandwidth, low-latency links between sites because each gradient-synchronisation step makes thousands of GPUs exchange data before they can continue. Modern interconnect designs increasingly use custom optical switching and remote direct memory access over converged Ethernet, known as RoCEv2.

RoCEv2 carries remote direct memory access traffic over converged Ethernet, allowing one accelerator to read another’s memory directly without waking the host processor for every transfer.

Long-haul fibre brings packet loss and jitter that a single building never encounters. Specialised transport layers help reduce those issues, while training across cloud regions requires congestion controls such as Explicit Congestion Notification (ECN) and Priority Flow Control (PFC). The same congestion logic used by a home router in our Wi-Fi Setup Network guide has become a hyperscale engineering discipline.

Who Is Already Running Training Across Sites?

Most of the largest labs have made the move. Google has said Gemini trained synchronously across clusters in multiple locations. Microsoft connected AI datacenters in Wisconsin and Georgia into what it calls one distributed AI supercomputer — a project running alongside its Microsoft Windows Changes push into AI-ready hardware. AWS connected compute clusters across wide areas so Anthropic could build Claude models, while Meta built high-capacity datacenter interconnects for training.

OperatorDistributed training approach
GoogleGemini trained synchronously across clusters in multiple locations
MicrosoftWisconsin and Georgia AI datacenters joined as one supercomputer
AWSCompute clusters linked across wide areas for Anthropic’s Claude models
MetaHigh-capacity datacenter interconnects built for model training
CoreWeave + Google CloudCross-cloud training conducted over a private interconnect

CoreWeave and Google Cloud announced cross-cloud training over a private interconnect, with Azure expected to follow later this year. A sponsored feature at Theregister provided the details.

What Happens Next for the Network?

The transport layer will likely become a product category of its own. Optical switching vendors, Ethernet silicon makers, and cloud providers are all pursuing the same goal: making a 100-kilometre link behave like a local network.


FAQs

What is the significance of Microsoft’s distributed AI supercomputer?

Microsoft’s distributed AI supercomputer links datacenters in Wisconsin and Georgia, marking a significant step forward for AI infrastructure. By combining those sites, the approach provides more computing power and efficiency, allowing teams to train complex AI models across multiple locations. It also fits Microsoft’s wider plan to develop AI-ready hardware and support new artificial intelligence advances.

How are AWS and Meta contributing to distributed AI training?

AWS has connected compute clusters across wide areas to help Anthropic build its Claude models, showing how companies are working together on AI development. Meta, meanwhile, has built high-capacity datacenter interconnects specifically for model training. That work highlights how important capable infrastructure has become as distributed AI training grows.

Was this article helpful?

Your feedback directly improves future articles on this site.

Follow us on Google News Get real-time updates & exclusive tech coverage
Follow

Leave a Reply

Your email address will not be published. Required fields are marked *

wp_enqueue_script('jquery', false, [], false, true); // load in footer