Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why GPU clusters need NVLink and InfiniBand

Training a frontier model means thousands of GPUs taking the same step at the same time. Ethernet wasn't built for that, and PCIe gave up a long time ago.

Networking intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Five pictures for a reader who has never seen a training cluster, following one forced sync: 10,000 participants, nobody moves until everyone has the same copy.

1 · The problem

It isn’t a lot of traffic. It’s a barrier.

every GPU must finish the sum before any GPU may take the next step GPUGPUGPUGPUGPUGPUGPUGPU one slow link nobody crosses until every contribution is in the sum so the slowest link in the cluster sets the training speed for everybody and the volume is roughly the size of the model, every step a 100B-parameter model in bfloat16 is about 200 GB of gradients to reconcile per optimizer step if the compute step is fast, the network has milliseconds to move that, or it becomes the bottleneck Traffic that has to synchronise, not traffic that averages out.
There is no fast path and slow path here, and no “retry later” — the next step literally cannot start. A 1% tail-latency spike on one link is sampled by every step on every GPU, which is exactly the property ordinary networking was never designed around.

2 · Inside the box

PCIe is about an order of magnitude short.

inside one server: NVLink instead of the PCIe slots PCIe Gen5 ×16 128 GB/s total (64 each way) NVLink 4, H100 SXM 900 GB/s bidirectional per GPU roughly 7× per GPU — and that is only the bandwidth half of the argument the other half is topology over PCIe, GPU-to-GPU traffic crosses switches and shares the tree with network cards and SSDs over NVLink it is a dedicated mesh with its own switch chips, every GPU to every other at full speed Eight is not a marketing number. It is the size of that non-blocking domain.
Vendor figures, not measurements. The topology half matters as much as the bandwidth half: on PCIe the link you want is sharing a tree with everything else in the machine, so what you lose is not just speed but predictability — and predictability is what a barrier is made of.

3 · Between boxes

Dropping a packet and recovering is a web habit.

between servers, the boring fabric doesn’t work either ordinary Ethernet best-effort: under congestion it drops and expects recovery fine for a web request, disastrous for a barrier InfiniBand or tuned RoCE lossless by design: credit-based flow control you never send a packet the receiver isn’t ready for training cares about microsecond tail latency, not average throughput so an occasional drop-and-recover is invisible on a web request and fatal to a step everyone is waiting on Meta trained Llama 3 on both, so this is “either, with care.”
In mid-2024 Meta publicly described training Llama 3 across two 24K-GPU clusters — one on InfiniBand, one on RoCE — tuned to equivalent performance, using the RoCE cluster for the largest model. “InfiniBand mandatory” has become “either, with enough tuning”, and the tuning is the real cost.

4 · The seam most posts skip

Doubling the cluster doesn’t double the bytes.

chunks pass hand to hand reduce-scatter, then all-gather per-GPU bytes on the wire 2(N−1)/N × K for a buffer of size K across N GPUs → approaches 2K from below as N grows so doubling the cluster costs latency, not volume more hops around the ring, the same bytes each Which is why the money goes into tail latency, not aggregate bandwidth.
The naive story is that a bigger cluster mainly means more bytes per GPU; the ring result breaks it. That is also why you build the short fast network inside the box and the long predictable one between boxes. Ring isn’t the only collective shape in use at very large scale — hierarchical and recursive schemes are common — but it is the cleanest way to see why the hardware splits the way it does.

5 · Keep this card

The whole thing on one index card.

GPU cluster network = NVLink inside the box, replacing PCIe + InfiniBand or tuned RoCE between boxes + collectives shaped like the sync training needs ∴ a barrier — the cluster runs at its worst link averages don’t help you when every step waits for the last arrival
Picture to keep: two nested rings — eight GPUs inside a box passing gradient chunks hand-to-hand over short fat wires, and the boxes themselves passing partial sums around a much longer ring, with nobody allowed to start the next step until the sum has gone all the way around.

Why it exists

Imagine 10,000 students all editing the same Google Doc, and every 30 seconds the system forces them to pause, share their changes with each other, and only resume once everyone has the same copy. With normal home Wi-Fi this would be hopeless — by the time the slowest student finishes syncing, the rest are sitting idle. Training a frontier AI model is that problem at thousands-of-GPUs scale. Every step, every GPU has to compare notes with every other GPU, and the “comparison” is hundreds of gigabytes of numbers. Ordinary network cables choke on this. NVLink and InfiniBand are what you build instead — fatter, but above all steadier — so the comparison step doesn’t dominate the whole training run. The steadiness turns out to matter more than the fatness, for reasons that take the rest of the post to unpack. That forced sync — 10,000 participants, nobody moves until everyone has the same copy — is the example this post follows.

The analogy breaks in one place worth naming: students editing a doc can merge changes lazily and mostly ignore each other. Gradient averaging is the opposite — it’s a barrier, and no GPU may take its next step until every GPU’s contribution is in the sum.

Open the spec sheet of a top-end AI training server — a DGX or HGX H100 SXM box, or one of the OEM clones — and you will find two networks you did not expect to find on a server.

Inside the box, eight GPUs are wired to each other through something called NVLink, not the PCIe slots you would normally expect to carry a GPU’s traffic. Between boxes, the cluster runs over InfiniBand or a carefully tuned variant of Ethernet — never the same boring 10/25/100G fabric the rest of the data center uses.

The obvious question is: why? PCIe is the universal interconnect. Ethernet runs the entire internet. They are both fast, both standardized, both cheap. Why does training one neural network require building two extra networks on top?

The honest answer is that training a frontier model is one of the strangest networking workloads ever invented. Every training step, often on sub-second timescales, every GPU in the cluster has to stop, sum its results with every other GPU, and only then can it take the next step. There is no “fast path” and “slow path” — the slowest link in the cluster sets the training speed for everybody. PCIe and ordinary Ethernet were designed for traffic that averages out. AI training is traffic that has to synchronize. Those are different problems, and they need different hardware.

Why it matters now

If you write software, you probably never touch NVLink directly. But almost everything you’d care about downstream is shaped by it.

The short answer

GPU cluster network = PCIe replacement (NVLink) + Ethernet replacement (InfiniBand or RoCE) + collectives that match how training actually communicates

Picture to keep: two nested rings — eight GPUs inside a box passing gradient chunks hand-to-hand over short fat wires, and the boxes themselves passing partial sums around a much longer ring, with nobody allowed to start the next step until the sum has gone all the way around.

Inside a server, NVLink replaces PCIe as the GPU-to-GPU link because PCIe is roughly an order of magnitude too slow and not optimised for sustained all-to-all GPU-to-GPU traffic. Between servers, InfiniBand (or carefully tuned Ethernet) replaces ordinary networking because training is bottlenecked on the tail of latency, not the average — and on collective operations like all-reduce, which have to finish before the next step can begin.

How it works

The thing to hold in your head is what training actually does on the network.

In standard data-parallel training, every GPU has a copy of the model. Each step, every GPU does a forward and backward pass on a different slice of the batch and produces its own gradients. Before the optimizer can take a step, those gradients have to be averaged across every GPU in the world that holds a copy of that parameter. The collective operation that does this is called all-reduce.

All-reduce has two properties that wreck normal networking:

  1. Everybody waits for the slowest GPU. A 1% tail-latency spike on one link stalls the entire cluster for that step. There is no “retry later” — the next step literally cannot start.
  2. The volume is roughly the size of the model, every step. For a 100B-parameter model in bfloat16, that’s ~200 GB of gradients to reconcile per optimizer step. If your compute step is fast, the network has milliseconds to move that data, or it becomes the bottleneck.

Naive attempt: use what’s already in the server. Plug the GPUs into PCIe slots, run the cluster over the same Ethernet as everything else, let TCP handle the losses. This is exactly how a normal distributed system is built, and for a normal distributed system it’s the right call. Here it fails twice over — once inside the box and once between boxes.

Inside the box: NVLink vs. PCIe. PCIe Gen5 ×16 gives you 128 GB/s total bidirectional (64 GB/s each way). NVLink 4 on an H100 SXM gives each GPU 900 GB/s of bidirectional bandwidth — NVIDIA’s number, and the one most third-party write-ups repeat. (The PCIe-form-factor H100 NVL is lower; this post is about the SXM parts that go into HGX/DGX nodes.) That’s roughly 7× per GPU. NVLink is also topologically better: GPU-to-GPU traffic over PCIe has to traverse switches and sometimes a shared root complex, and the bandwidth is shared with everything else on that PCIe tree — network cards, SSDs, the other GPU you’re trying to talk to. NVLink is a dedicated mesh of GPU-to-GPU links with its own switch fabric.

In an HGX H100 8-GPU board, every GPU is wired to four NVSwitch chips, and the NVSwitches are wired to each other in a way that gives you a non-blocking all-to-all fabric: every GPU can simultaneously talk to every other GPU at full NVLink speed. NVIDIA quotes 3.6 TB/s of bisection bandwidth for that 8-GPU domain. The point isn’t the headline number; it’s that there are no contention surprises within a node.

flowchart LR
    subgraph Node A
      G1[GPU] <-->|NVLink ~900 GB/s| SW1[NVSwitch]
      G2[GPU] <-->|NVLink| SW1
    end
    subgraph Node B
      G3[GPU] <-->|NVLink| SW2[NVSwitch]
      G4[GPU] <-->|NVLink| SW2
    end
    SW1 <-->|InfiniBand / RoCE via NICs<br/>much thinner, much longer| SW2

Between boxes: InfiniBand vs. Ethernet. Ordinary Ethernet has two problems for this workload. The first is latency: training cares about microsecond tail latencies, and a switch chain that occasionally drops packets and recovers is fine for HTTP and disastrous for all-reduce. The second is that ordinary Ethernet is best-effort — under congestion it drops packets and expects upper layers to recover. InfiniBand was designed lossless from day one, with credit-based flow control: you never send a packet the receiver isn’t ready for.

In practice, InfiniBand is usually a little lower-latency and more predictable out of the box, while RoCE (RDMA over Converged Ethernet) can match it on throughput with careful fabric tuning — flow control, ECN, congestion control, switch buffer sizing. Specific microsecond numbers floating around (e.g. ~1 µs vs 1.5–2.5 µs) are vendor- and tuning-dependent and I’d be cautious about treating any single pair of numbers as canonical.

InfiniBand also ships with a feature called SHARP that does part of the all-reduce sum inside the network switches, so more of the reduction happens in the fabric before the result reaches the GPUs. That’s the kind of thing you can’t bolt onto general-purpose Ethernet without reinventing it — and that reinvention, under the name “Ultra Ethernet” plus AI-specific silicon like NVIDIA’s Spectrum-X and Broadcom’s Tomahawk 6, is exactly what’s been happening for the last couple of years. In mid-2024 Meta publicly described training Llama 3 across two 24K-GPU clusters — one on InfiniBand, one on RoCE — tuned to equivalent performance, with the RoCE cluster used for the largest model. So the “InfiniBand mandatory” answer is becoming “either, with care.”

The seam most posts skip. The naive story is that a bigger cluster mainly means more bytes per GPU. Ring all-reduce breaks that intuition, and all of this only matters because of how all-reduce decomposes. The classical implementation, ring all-reduce, splits the gradient buffer into chunks and pipelines them around a ring of GPUs in two passes (reduce-scatter, then all-gather). The clever part: per-GPU bytes-on-the-wire is 2(N-1)/N × K for a buffer of size K and N GPUs — which approaches 2K from below as N grows, and is essentially independent of cluster size. So the cost of doubling the cluster is dominated by latency (more hops around the ring), not bandwidth — which is why microsecond tail behavior matters so much, and why you build the cheap fast network inside the box (NVLink, where the ring is short) and the expensive predictable network between boxes (InfiniBand or tuned RoCE, where the ring is long). Ring isn’t the only collective shape in use anymore — at very large scale, recursive halving/doubling and hierarchical schemes are common — but the ring analysis is the cleanest way to see why the bandwidth/latency split shows up in the hardware.

You started with GPU cluster network = NVLink + InfiniBand/RoCE + collectives. What did the 10,000 synchronized students add? — + the workload is a barrier, so the cluster runs at the speed of its worst link. Averages don’t help you when every step waits for the last arrival, and that one property is why the money goes into lossless fabrics and tail latency rather than into more aggregate bandwidth.

Check yourself

Before you go — you double a cluster from 1,000 to 2,000 GPUs and keep the per-GPU batch size the same. Using the ring all-reduce result, does each GPU now have to push roughly twice as many bytes per step?

Answer

No. Per-GPU bytes on the wire are 2(N-1)/N × K — at N=1000 that’s already essentially 2K, and at N=2000 it’s still essentially 2K. What actually grows is the number of hops the data makes around the ring, so the cost of doubling shows up as latency, not volume. That’s why per-hop predictability matters more than raw bandwidth at large N, and why hierarchical collectives get used instead of one giant ring.

And one more — a cluster shows 99th-percentile link latency 20× the median, but median throughput looks great on every dashboard. How much should you worry?

Answer

A lot. Because all-reduce is a barrier, every GPU in the collective waits for the slowest participant in that step, so a tail that fires occasionally is sampled by every step across every GPU. This is the specific way in which training traffic differs from web traffic, where a slow request affects one user and averages out. Median dashboards are close to useless here; the distribution’s right tail is the number that sets your training throughput.

Going deeper