Why GPU clusters need NVLink and InfiniBand
Training a frontier model means thousands of GPUs taking the same step at the same time. Ethernet wasn't built for that, and PCIe gave up a long time ago.
On this page
The picture version
Five pictures for a reader who has never seen a training cluster, following one forced sync: 10,000 participants, nobody moves until everyone has the same copy.
1 · The problem
It isn’t a lot of traffic. It’s a barrier.
2 · Inside the box
PCIe is about an order of magnitude short.
3 · Between boxes
Dropping a packet and recovering is a web habit.
4 · The seam most posts skip
Doubling the cluster doesn’t double the bytes.
5 · Keep this card
The whole thing on one index card.
Why it exists
Imagine 10,000 students all editing the same Google Doc, and every 30 seconds the system forces them to pause, share their changes with each other, and only resume once everyone has the same copy. With normal home Wi-Fi this would be hopeless — by the time the slowest student finishes syncing, the rest are sitting idle. Training a frontier AI model is that problem at thousands-of-GPUs scale. Every step, every GPU has to compare notes with every other GPU, and the “comparison” is hundreds of gigabytes of numbers. Ordinary network cables choke on this. NVLink and InfiniBand are what you build instead — fatter, but above all steadier — so the comparison step doesn’t dominate the whole training run. The steadiness turns out to matter more than the fatness, for reasons that take the rest of the post to unpack. That forced sync — 10,000 participants, nobody moves until everyone has the same copy — is the example this post follows.
The analogy breaks in one place worth naming: students editing a doc can merge changes lazily and mostly ignore each other. Gradient averaging is the opposite — it’s a barrier, and no GPU may take its next step until every GPU’s contribution is in the sum.
Open the spec sheet of a top-end AI training server — a DGX or HGX H100 SXM box, or one of the OEM clones — and you will find two networks you did not expect to find on a server.
Inside the box, eight GPUs are wired to each other through something called NVLink, not the PCIe slots you would normally expect to carry a GPU’s traffic. Between boxes, the cluster runs over InfiniBand or a carefully tuned variant of Ethernet — never the same boring 10/25/100G fabric the rest of the data center uses.
The obvious question is: why? PCIe is the universal interconnect. Ethernet runs the entire internet. They are both fast, both standardized, both cheap. Why does training one neural network require building two extra networks on top?
The honest answer is that training a frontier model is one of the strangest networking workloads ever invented. Every training step, often on sub-second timescales, every GPU in the cluster has to stop, sum its results with every other GPU, and only then can it take the next step. There is no “fast path” and “slow path” — the slowest link in the cluster sets the training speed for everybody. PCIe and ordinary Ethernet were designed for traffic that averages out. AI training is traffic that has to synchronize. Those are different problems, and they need different hardware.
Why it matters now
If you write software, you probably never touch NVLink directly. But almost everything you’d care about downstream is shaped by it.
- Cluster cost. A surprisingly large slice of a frontier-training bill is interconnect, not silicon. Networking gear, optics, switches, and cables make up a non-trivial fraction of an AI data center’s capex, though company-by-company breakdowns generally aren’t published.
- Why GPUs ship in groups of 8. A “node” — DGX, HGX, and the OEM clones — is built around the unit that fits inside one non-blocking NVLink fabric. Eight isn’t a marketing number; it’s the size of the domain where every GPU can talk to every other GPU at full bandwidth.
- Why training scales the way it does. A large training run only pays off if you can keep all the GPUs in lockstep. One reason 100k-GPU-class runs are possible at all is that the network improved alongside the chips, not just the chips themselves.
The short answer
GPU cluster network = PCIe replacement (NVLink) + Ethernet replacement (InfiniBand or RoCE) + collectives that match how training actually communicates
Picture to keep: two nested rings — eight GPUs inside a box passing gradient chunks hand-to-hand over short fat wires, and the boxes themselves passing partial sums around a much longer ring, with nobody allowed to start the next step until the sum has gone all the way around.
Inside a server, NVLink replaces PCIe as the GPU-to-GPU link because PCIe is roughly an order of magnitude too slow and not optimised for sustained all-to-all GPU-to-GPU traffic. Between servers, InfiniBand (or carefully tuned Ethernet) replaces ordinary networking because training is bottlenecked on the tail of latency, not the average — and on collective operations like all-reduce, which have to finish before the next step can begin.
How it works
The thing to hold in your head is what training actually does on the network.
In standard data-parallel training, every GPU has a copy of the model. Each step, every GPU does a forward and backward pass on a different slice of the batch and produces its own gradients. Before the optimizer can take a step, those gradients have to be averaged across every GPU in the world that holds a copy of that parameter. The collective operation that does this is called all-reduce.
All-reduce has two properties that wreck normal networking:
- Everybody waits for the slowest GPU. A 1% tail-latency spike on one link stalls the entire cluster for that step. There is no “retry later” — the next step literally cannot start.
- The volume is roughly the size of the model, every step. For a 100B-parameter model in bfloat16, that’s ~200 GB of gradients to reconcile per optimizer step. If your compute step is fast, the network has milliseconds to move that data, or it becomes the bottleneck.
Naive attempt: use what’s already in the server. Plug the GPUs into PCIe slots, run the cluster over the same Ethernet as everything else, let TCP handle the losses. This is exactly how a normal distributed system is built, and for a normal distributed system it’s the right call. Here it fails twice over — once inside the box and once between boxes.
Inside the box: NVLink vs. PCIe. PCIe Gen5 ×16 gives you 128 GB/s
total bidirectional (64 GB/s each way). NVLink 4 on an H100 SXM gives
each GPU 900 GB/s of bidirectional bandwidth — NVIDIA’s number, and the
one most third-party write-ups repeat. (The PCIe-form-factor H100 NVL
is lower; this post is about the SXM parts that go into HGX/DGX nodes.)
That’s roughly 7× per GPU. NVLink is also topologically better:
GPU-to-GPU traffic over PCIe has to traverse switches and sometimes a
shared root complex, and the bandwidth is shared with everything else on
that PCIe tree — network cards, SSDs, the other GPU you’re trying to talk to.
NVLink is a dedicated mesh of GPU-to-GPU links with its own switch fabric.
In an HGX H100 8-GPU board, every GPU is wired to four NVSwitch chips, and the NVSwitches are wired to each other in a way that gives you a non-blocking all-to-all fabric: every GPU can simultaneously talk to every other GPU at full NVLink speed. NVIDIA quotes 3.6 TB/s of bisection bandwidth for that 8-GPU domain. The point isn’t the headline number; it’s that there are no contention surprises within a node.
flowchart LR
subgraph Node A
G1[GPU] <-->|NVLink ~900 GB/s| SW1[NVSwitch]
G2[GPU] <-->|NVLink| SW1
end
subgraph Node B
G3[GPU] <-->|NVLink| SW2[NVSwitch]
G4[GPU] <-->|NVLink| SW2
end
SW1 <-->|InfiniBand / RoCE via NICs<br/>much thinner, much longer| SW2
Between boxes: InfiniBand vs. Ethernet. Ordinary Ethernet has two problems for this workload. The first is latency: training cares about microsecond tail latencies, and a switch chain that occasionally drops packets and recovers is fine for HTTP and disastrous for all-reduce. The second is that ordinary Ethernet is best-effort — under congestion it drops packets and expects upper layers to recover. InfiniBand was designed lossless from day one, with credit-based flow control: you never send a packet the receiver isn’t ready for.
In practice, InfiniBand is usually a little lower-latency and more predictable out of the box, while RoCE (RDMA over Converged Ethernet) can match it on throughput with careful fabric tuning — flow control, ECN, congestion control, switch buffer sizing. Specific microsecond numbers floating around (e.g. ~1 µs vs 1.5–2.5 µs) are vendor- and tuning-dependent and I’d be cautious about treating any single pair of numbers as canonical.
InfiniBand also ships with a feature called SHARP that does part of the all-reduce sum inside the network switches, so more of the reduction happens in the fabric before the result reaches the GPUs. That’s the kind of thing you can’t bolt onto general-purpose Ethernet without reinventing it — and that reinvention, under the name “Ultra Ethernet” plus AI-specific silicon like NVIDIA’s Spectrum-X and Broadcom’s Tomahawk 6, is exactly what’s been happening for the last couple of years. In mid-2024 Meta publicly described training Llama 3 across two 24K-GPU clusters — one on InfiniBand, one on RoCE — tuned to equivalent performance, with the RoCE cluster used for the largest model. So the “InfiniBand mandatory” answer is becoming “either, with care.”
The seam most posts skip. The naive story is that a bigger cluster
mainly means more bytes per GPU. Ring all-reduce breaks that intuition, and
all of this only matters because of how all-reduce decomposes. The classical implementation, ring all-reduce,
splits the gradient buffer into chunks and pipelines them around a
ring of GPUs in two passes (reduce-scatter, then all-gather). The
clever part: per-GPU bytes-on-the-wire is 2(N-1)/N × K for a buffer
of size K and N GPUs — which approaches 2K from below as N grows, and
is essentially independent of cluster size. So the cost of doubling
the cluster is dominated by latency (more hops around the ring), not
bandwidth — which is why microsecond tail behavior matters so much,
and why you build the cheap fast network inside the box (NVLink, where
the ring is short) and the expensive predictable network between boxes
(InfiniBand or tuned RoCE, where the ring is long). Ring isn’t the
only collective shape in use anymore — at very large scale, recursive
halving/doubling and hierarchical schemes are common — but the ring
analysis is the cleanest way to see why the bandwidth/latency split
shows up in the hardware.
You started with GPU cluster network = NVLink + InfiniBand/RoCE + collectives. What did the 10,000 synchronized students add? — + the workload is a barrier, so the cluster runs at the speed of its worst link.
Averages don’t help you when every step waits for the last arrival, and that
one property is why the money goes into lossless fabrics and tail latency
rather than into more aggregate bandwidth.
Check yourself
Before you go — you double a cluster from 1,000 to 2,000 GPUs and keep the per-GPU batch size the same. Using the ring all-reduce result, does each GPU now have to push roughly twice as many bytes per step?
Answer
No. Per-GPU bytes on the wire are 2(N-1)/N × K — at N=1000 that’s already
essentially 2K, and at N=2000 it’s still essentially 2K. What actually
grows is the number of hops the data makes around the ring, so the cost of
doubling shows up as latency, not volume. That’s why per-hop predictability
matters more than raw bandwidth at large N, and why hierarchical collectives
get used instead of one giant ring.
And one more — a cluster shows 99th-percentile link latency 20× the median, but median throughput looks great on every dashboard. How much should you worry?
Answer
A lot. Because all-reduce is a barrier, every GPU in the collective waits for the slowest participant in that step, so a tail that fires occasionally is sampled by every step across every GPU. This is the specific way in which training traffic differs from web traffic, where a slow request affects one user and averages out. Median dashboards are close to useless here; the distribution’s right tail is the number that sets your training throughput.
Famous related terms
- NVLink —
NVLink = direct GPU-to-GPU link + non-blocking switch fabric (NVSwitch)— replaces PCIe for in-box GPU traffic; ~900 GB/s bidirectional per H100. - NVSwitch —
NVSwitch ≈ Ethernet switch, but for NVLink— what makes the 8-GPU all-to-all fabric inside a DGX/HGX node non-blocking. - InfiniBand —
InfiniBand = lossless fabric + RDMA + microsecond latency— the historical default between AI servers; designed for HPC long before “AI cluster” was a phrase. - RoCE —
RoCE = RDMA semantics + lossless-tuned Ethernet— Ethernet’s answer to InfiniBand for AI fabrics. - All-reduce —
all-reduce = sum-across-everyone + result-to-everyone— the collective op that gates every training step. - Ring all-reduce —
ring all-reduce ≈ reduce-scatter + all-gather around a ring— the pattern that makes per-GPU traffic approach 2× the buffer regardless of cluster size, at the price of being latency-sensitive (more hops as N grows).
Going deeper
- Patarasuk & Yuan, Bandwidth Optimal All-reduce Algorithms for Clusters of Workstations (J. Parallel Distrib. Comput., 2009) — the primary source for the question “why is per-GPU traffic independent of cluster size?”, with the proof rather than the summary.
- Andrew Gibiansky, Bringing HPC techniques to deep learning (Baidu Research, 2017) — the explainer for “what does a ring all-reduce actually do, step by step,” and the piece that put the idea in front of ML practitioners.
- Meta Engineering, RoCE networks for distributed AI training at scale (2024) — the rabbit hole for “is InfiniBand actually required?”, written by people who ran a 24K-GPU Llama 3 cluster on Ethernet instead.
- NVIDIA’s HGX H100 material — where the 8-GPU topology, NVSwitch count, and 3.6 TB/s bisection figure come from; first-party and self-interested, so read the numbers as vendor specs rather than measurements.