Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why networks are big-endian but your CPU is little-endian

Two halves of the same machine disagree on which end of a number comes first. The split is older than you, and it's never going away.

Computer Science intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who has never opened a memory viewer. The prose below fills in the seams the pictures skip.

1 · The problem

You stored the number four. Memory shows it backwards.

length = 4 04000000 address +0+1+2+3 And the debugger reports this as completely normal. nothing in “a 32-bit integer” says which of its four bytes should come first
Four bytes have to be laid out in some order, and the arithmetic of a 32-bit number says nothing about which. That single 32-bit field holding the value 4 is the running example for everything below.

2 · Two answers, both defensible

Start from the important end, or start from the unimportant end.

big end first 00000004 reads the way you would write it little end first 04000000 the byte that matters is at the front Same number. Same four bytes. Opposite direction of travel. Your laptop and phone use the right-hand one. Every packet you send over the network uses the left.
Neither order is wrong; they are two conventions for walking the same four boxes. The trouble is that both are in daily use, on opposite sides of a boundary you cross constantly.

3 · Where the naive assumption breaks

Just copy the bytes across, and nobody reports an error.

sender writes 04 00 00 00 copied faithfully receiver reads 04 00 00 00 67,108,864 It now waits for a 64-megabyte message that was never sent. nothing crashed · nothing logged · it just hangs The bytes arrived perfectly. The disagreement is silent by construction — nothing in the four bytes records which convention wrote them
The wire delivers the bytes exactly as sent; the two machines simply read them in opposite directions. Because the bytes carry no note saying which convention produced them, the mismatch cannot announce itself — it shows up as a hang, not an error.

4 · The fix

Pick one order for the wire and make everybody convert.

whatever this machine uses inside 04 00 00 00 swap the one agreed order, on the wire 00 00 00 04 swap whatever that machine uses inside 04 00 00 00 On a machine that already agrees, the swap costs nothing. On one that doesn’t, it is a real byte-reverse — and it happens every time. which is why the conversion functions have lived in every network stack since the early 1980s
The agreement is out-of-band: the protocol says which order the wire uses, and every machine converts on the way in and out. Nothing detects a mismatch — the convention is simply obeyed or the field is garbage.

5 · Why each side chose what it chose

One was picked for human eyes. The other for a hardware convenience.

big end first, on the wire a hex dump reads like the number it represents which mattered a great deal to people reading packet traces by hand little end first, in the chip read the first byte and you have the small version of the number no shifting, no moving — a real convenience for early hardware both of these are the standard accounts rather than documented decisions — the what is solid, the why is reconstruction
The network order was chosen so a packet dump reads the way a person writes a number; the chip order makes a narrow read of a wide value free. Both explanations are the standard accounts rather than confirmed design records — no contemporary document settles the chip side.

6 · Keep this card

The whole thing on one index card.

endianness = which end of a multi-byte number goes first + two answers, both still in daily use ∴ somebody has to convert, and nobody is told The bytes never say which convention wrote them. Which is why getting it wrong is quiet, and why it is still worth knowing.
Picture to keep: four numbered mailboxes in a row and a four-digit number to post — one convention drops the most important digit in box 0, the other drops the least important there. The number is the same; only the direction you walk the boxes differs. Where it breaks: mailboxes are labelled and bytes aren’t, which is why the reader has to be told out of band.

Why it exists

You’ve probably had this moment: you set a variable to 4, then open the memory view in your debugger — or dump the file with xxd — and the four bytes read 04 00 00 00. Not 00 00 00 04. The number you just wrote looks like it’s stored backwards, and the tooling reports this as completely normal.

That’s the running example for the rest of this post: a single 32-bit field holding the value 4 — say, a length prefix in front of a four-byte message. Four bytes have to be laid out in some particular order in memory and on the wire, and nothing in the math of “a 32-bit integer” tells you which byte goes first.

That ambiguity — which end of a multi-byte number is “first”? — is endianness. And the answer turns out to be: it depends on who you ask, and the two answers we landed on disagree.

So the moment a number crosses from your CPU’s registers to a network packet, somebody has to swap bytes. We’ve been swapping since the early 1980s. Why?

Why it matters now

Endianness is one of those things that’s invisible until you cross a boundary. The boundaries it lives on:

A bug in this layer doesn’t crash loudly — it produces wrong numbers. A length field of 4 read as 67_108_864 is the kind of failure that shows up as “the packet parser hangs forever” rather than “segfault on line 47.” That’s why the convention exists at all: pick one, write it down, swap if you have to.

The short answer

endianness = which byte of a multi-byte number lives at the lowest memory address

Picture to keep: four numbered mailboxes in a row, and a four-digit number to post. Big-endian drops the most important digit in box 0; little-endian drops the least important digit in box 0. The number is the same; only the direction you walk the boxes differs. (Where the picture breaks: mailboxes are labelled, so you could always look. Bytes aren’t — nothing in the four bytes says which convention wrote them, which is why the reader has to be told out-of-band and why disagreement is silent rather than loud.)

Both work. Both have defensible arguments. The split between “network byte order is big” and “x86 is little” is a historical accident that calcified into a standard.

How it works

Start with the naive assumption: a 32-bit number is a 32-bit number, so just copy the bytes. Take our length field, the value 4. In a register, it’s just 32 bits — no “order” exists. The order only appears when you ask: what byte is at address N, N+1, N+2, N+3?

                addr  +0  +1  +2  +3
big-endian          00  00  00  04     (matches written order)
little-endian       04  00  00  00     (low byte first)

The CPU’s load and store instructions pick one convention and bake it into the silicon. On x86-64, MOV of a 32-bit word writes those four bytes in little-endian order, full stop. On a big-endian machine like a classic PowerPC or a SPARC, the same instruction would write them in the opposite order. Same number in the register, different bytes in RAM.

Why the naive assumption breaks. Send those four bytes across a network. The wire is a stream of bytes, no concept of “register.” The receiver picks them up in the order they arrived. A little-endian sender writes 04 00 00 00; a receiver that reads most-significant-byte-first sees 0x04000000 — 67,108,864. It now waits for a 64-megabyte message that will never arrive. Nothing crashed. Nothing logged an error. The parser just hangs.

The fix the early internet adopted: pick one order for the wire and make everyone convert. That order, defined in the early TCP/IP RFCs, is big-endian — what we now call “network byte order.” Every host, regardless of native endianness, swaps to big-endian on the way out and back on the way in. Our length field goes onto the wire as 00 00 00 04 no matter who sent it. On a big-endian host, the swap is a no-op. On a little-endian host (most of them today), it’s a real byte-reverse.

Why big for the network?

The standard account: big-endian is what humans write. When you write 12,345, the most significant digit is leftmost — it’s big-endian positional notation. Putting bytes on the wire most-significant-first means a hex dump of a packet looks like the number it represents. For protocol designers reading network traces by hand in 1981, that mattered a lot.

There’s also a small algorithmic argument: when comparing numbers lexicographically as byte strings, big-endian sorts the same way as the numbers themselves. Useful for some routing tricks, less so today.

Why little for x86?

The standard account here is more mechanical and arguably more interesting. Little-endian has a property that mattered to early hardware designers: reading a smaller integer from the start of a larger one Just Works.

Imagine you stored a 32-bit value 0x0000_00FF in little-endian: bytes FF 00 00 00. If you load just one byte from that address, you get 0xFF. Load two, you get 0x00FF. Load all four, you get 0x000000FF. The address of the value is the same regardless of how wide a load you do — because the low byte sits at the bottom. Big-endian would put the FF at the end, so a 1-byte load at the same address gives you 0x00, not 0xFF. You’d have to adjust the address by the difference in widths.

This made arithmetic carry propagation, pointer truncation between widths, and certain kinds of mixed-width arithmetic a hair simpler in hardware. Intel adopted little-endian for the 8086 in 1978, AMD inherited it for x86-64, and ARM made little-endian the default in practice. RISC-V is a partial exception worth stating precisely: instruction fetch is fixed little-endian, but the spec allows little-endian, big-endian, and bi-endian data accesses — real implementations are overwhelmingly little-endian anyway. The momentum, not the spec, is what’s overwhelming.

No clean primary source explains why exactly Intel picked little-endian for the 8086. The often-repeated story is the mixed-width-load argument above, but no contemporary Intel design document confirms that was the deciding factor versus, say, compatibility with the earlier 8080 and its accumulator conventions. Take the “why” with a grain of salt; the “what” is solid.

Bi-endian and the awkward middle

Some architectures — older ARM, older MIPS, IA-64, PowerPC — are bi-endian: a configuration bit selects which mode the CPU runs in. In practice almost everyone configures these as little-endian today, because that’s where the software ecosystem is. The mode bit is a relic.

Then there’s mixed-endian (“middle-endian”), which used to exist on the PDP-11 for 32-bit words and occasionally shows up in formats where one field is little-endian and another is big-endian in the same struct — typically because a format grew up on little-endian machines but embedded fields inherited from a big-endian protocol. It’s universally regarded as a footgun, for the obvious reason: correctness now depends on remembering the direction field by field. (SMB/CIFS is often named as the canonical example; the base spec actually says multi-byte fields are little-endian unless otherwise noted, so treat “SMB is mixed-endian” as lore with no clean source behind it.)

The seams

The deeper observation: endianness is a coordination problem the industry solved twice — once for hosts (little wins), once for the wire (big wins) — and never reconciled. It’s cheap enough to swap that nobody has to.

You started with endianness = which byte lives at the lowest address. What did the length-field story add? — + it is only ever a question at a boundary. Inside one machine the convention is invisible and self-consistent; the bug surface is entirely where bytes cross from one convention’s world into another’s, which is exactly why the fix is a conversion function and not a better number format.

Check yourself

Before you go — a colleague argues that since almost every machine today is little-endian, network byte order is a pointless legacy tax and new protocols should just use little-endian and skip the swap. What’s the strongest counter?

Answer

Two things. First, the tax is nearly free: htonl compiles to a single BSWAP/REV instruction, so you’re arguing about a cycle. Second — and this is the real point — the value of a convention is that it’s fixed, not that it’s optimal. A new protocol choosing little-endian doesn’t remove the conversion; it adds a second convention that every parser and every hex-dump reader now has to disambiguate per-field. That’s the SMB/CIFS failure mode. Note the counter isn’t “big-endian is better” — it’s that switching costs more than it saves.

And: your program stores a uint32_t and then reads the first byte of it via a uint8_t*. Same code, x86 laptop and a big-endian machine. Same answer?

Answer

No — and this is the little-endian design argument in miniature. On little-endian, that first byte is the low byte: for our length field 4 you read 0x04. On big-endian, it’s the high byte: you read 0x00. This is why the pattern is a portability bug, and also why little-endian made narrowing loads convenient for early hardware — the address of a value doesn’t change with the width you read it at.

Going deeper