Why UTF-8 won
Unicode could have been a fixed 4-byte-per-character encoding. Instead, the web runs on a variable-width hack — and that hack is why everything still works.
On this page
The picture version
Five pictures for a reader who has only ever seen the mangled name. The prose below fills in the seams the pictures skip.
1 · The problem
One character in. Two nonsense characters out.
2 · The obvious answer, and why it lost
Give every character four bytes. Watch everything break.
3 · The trick
Put the length in the high bits, and mark every follower.
4 · Why old code survived
Three ranges of byte, and none of them overlap.
5 · Keep this card
The whole thing on one index card.
Why it exists
You’ve seen a name come back wrong. You type José into a form, or open a CSV
in the wrong program, and out comes José. One character became two, and the
two are nonsense. That’s not corruption — every byte arrived intact. It’s two
programs disagreeing about what the bytes mean.
Here’s what actually happened, and it’s the running example for this whole
post: the single character é. In UTF-8 it’s two bytes, C3 A9. A program
that assumes one byte per character reads them as two separate old-style
characters — Ã and © — and prints both. The encoding is the contract, and
someone broke it.
To see why that contract is shaped so strangely, rewind. Imagine it’s 1992.
You’re writing C code. A string is a char* — a pointer to bytes that ends with a 0. Every library, every kernel call, every config file, every protocol assumes this. ASCII fits in 7 bits, so one byte per character, and the world holds together.
Then you want to support Japanese. And Arabic. And Cyrillic. And you discover there are roughly a million possible characters in the world’s writing systems once you count CJK ideographs and historic scripts. A byte isn’t enough. So what do you do?
The obvious answer — make every character 4 bytes — is what UTF-32 does. It’s clean. Index i of the string is at offset i * 4. No ambiguity. The problem: every existing program that reads bytes, every filesystem, every network protocol, every C library breaks. A file that used to be 1KB is now 4KB. And 90% of the bytes in an English document are zeros, because ASCII characters only need 7 bits but you’re padding them out to 32. The disk groans. The wire groans. And worst of all, a \0 byte appears inside every English character — so strlen, which stops at the first zero byte, reports a length of 0 or 1 and every string in your program is silently truncated at its first letter.
UTF-8 is the answer to “how do we add a million characters without breaking anything that already exists.” It exists because backwards compatibility with ASCII and with byte-oriented C code wasn’t a nice-to-have — it was the only way the new encoding could possibly win adoption.
Why it matters now
Nearly every web page, JSON payload, source file and commit message you’ve touched this week is UTF-8. The WHATWG HTML Standard requires it for conforming documents. JSON exchanged between systems must use it (RFC 8259 §8.1). Linux filenames are byte sequences that are almost always interpreted as UTF-8. Go and Rust string literals are UTF-8 by definition. Even Windows, which spent two decades on UTF-16 internally, has been quietly migrating APIs toward UTF-8.
This means a software engineer who doesn’t understand UTF-8 will eventually hit a bug they can’t explain — a string that reverses wrong, a length that’s off, a regex that matches the middle of a character. AI-era engineers hit this constantly because LLM tokenizers operate on byte-pair-encoded UTF-8, not on “characters” in any human sense. (See tokenization.)
The short answer
UTF-8 = variable-width encoding (1–4 bytes) + ASCII as a literal subset + self-synchronizing byte pattern
Picture to keep: a train of carriages where the first carriage has a flag on the roof saying how many carriages this train has, and every following carriage is painted “I am not a first carriage.” Walk up to any carriage in the yard and you can tell instantly whether you’re at the front of a train or in the middle of one — and walking backwards a step or two always finds the front.
UTF-8 encodes each Unicode code point in 1, 2, 3, or 4 bytes. The first 128 code points (ASCII) are encoded as a single byte identical to ASCII — so every ASCII file is already valid UTF-8. Non-ASCII characters use multi-byte sequences whose bytes are all in the range 0x80–0xFF, which means they never collide with ASCII control characters or with \0.
How it works
The encoding rule is almost suspiciously simple. Look at the high bits of each byte:
0xxxxxxx → 1 byte, 7 bits of payload (ASCII)
110xxxxx 10xxxxxx → 2 bytes, 11 bits
1110xxxx 10xxxxxx 10xxxxxx → 3 bytes, 16 bits
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx → 4 bytes, 21 bits
A leading byte tells you how long the sequence is. A continuation byte always starts with 10. The high-bit pattern of the leading byte (0, 110, 1110, 11110) is also the count of leading 1s, which is the byte length. From any byte in the stream, you can find the start of the current character by scanning backward at most 3 bytes for one that doesn’t start with 10. That property — self-synchronizing — means a corrupted byte loses one character, not the whole rest of the file.
Run our é through it. Its code point is U+00E9 = decimal 233 = binary 11101001 — eight significant bits, one more than the seven a single UTF-8 byte can carry, so it takes the 2-byte form: 110xxxxx 10xxxxxx. Right-align the value in the eleven x slots (00011 101001) and you get 11000011 10101001 = C3 A9. Both bytes are ≥ 0x80, so neither can be mistaken for an ASCII character; the second starts with 10, so it can never be mistaken for the start of anything. The mojibake at the top of this post is what you get when a reader ignores all of that and treats each byte as its own character.
The clever bit: continuation bytes are 10xxxxxx, which is 0x80–0xBF. A leading byte for a multi-byte sequence is 0xC2–0xF4 — RFC 3629 says outright that C0, C1 and F5–FF never appear in well-formed UTF-8. ASCII bytes are 0x00–0x7F. These ranges don’t overlap. So you can search a UTF-8 string for the ASCII byte '/' (0x2F) using plain memchr and you will never get a false hit inside a multi-byte character. This is the magic that let UTF-8 slot into existing C code without rewriting it. Byte-oblivious code mostly keeps working.
The trade you make: indexing is no longer O(1). s[5] doesn’t mean “the 6th character” anymore — to find the 6th character you have to walk from the start, decoding one code point at a time. In practice this matters less than people fear, because most string operations (search, concatenate, split on a delimiter, send over a socket) don’t need character-indexed access. They just need bytes that round-trip cleanly.
The seam: “characters” is a lie
Here’s the honest part. Even after you decode UTF-8 into code points, “character” still doesn’t mean what you think. The user-perceived character é can be one code point (U+00E9) or two (U+0065 U+0301 — e plus a combining acute accent). Family emoji like 👨👩👧 are sequences of multiple code points joined by zero-width joiners. What humans call a “character” is a grapheme cluster, and finding grapheme boundaries requires a Unicode-aware library and a giant lookup table that ships with every new emoji release.
So “string length” has at least four reasonable answers: bytes (UTF-8), code units (UTF-16), code points (Unicode scalars), and grapheme clusters (what users count). Our é can be 2 bytes and 1 code point, or 3 bytes and 2 code points, and it looks identical either way. Bugs love this gap. UTF-8 didn’t cause it — Unicode itself did — but UTF-8 made the gap visible at the byte layer where most engineers live.
You started with UTF-8 = variable-width + ASCII subset + self-synchronizing. What did é add to that line? — + the design goal was never elegance, it was not breaking existing code. Every awkward property here — the wasted bits in the prefix, the loss of O(1) indexing, the four different string lengths — is the price of the one property UTF-32 couldn’t offer: a program written before Unicode existed keeps working, byte for byte, on English text.
Famous related terms
- UTF-16 —
UTF-16 = variable-width 2-or-4-byte encoding + endianness baggage— the model behind JavaScript strings, the Windows API, and Java’s string semantics (though modern JVMs store the bytes compactly under the hood). Lost the web because it isn’t ASCII-compatible. - ASCII —
ASCII = 7-bit fixed encoding for English + control characters— UTF-8’s superset relationship with ASCII is the entire reason it won. - Code point —
code point = a Unicode integer (0 to 0x10FFFF)— the abstract identity of a character, separate from how it’s encoded as bytes. - Grapheme cluster —
grapheme ≈ what a human would call "one character"— usually a sequence of code points. Where most “string length” bugs come from.
Going deeper
- RFC 3629 (Yergeau, 2003) — the IETF specification, and the place to check the exact bit patterns and the rules about which byte sequences are invalid (a detail this post skipped, and a common source of security bugs).
- UAX #29, Unicode Text Segmentation — the grapheme-cluster machinery behind the seam above: the actual boundary rules for “what counts as one character.” (The Unicode Standard’s chapter 3 defines conformance and grapheme clusters; UAX #29 is where the segmentation algorithm lives.)
- Rob Pike’s UTF-8 turned 20 years old (in 2012) — the rabbit hole, and a first-hand account: the diner story is real, and Pike names the placemat it was sketched on.