Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why UTF-8 won

Unicode could have been a fixed 4-byte-per-character encoding. Instead, the web runs on a variable-width hack — and that hack is why everything still works.

Computer Science intro Apr 29, 2026 · updated Aug 25, 2026 · 9 min read

On this page

The picture version

Five pictures for a reader who has only ever seen the mangled name. The prose below fills in the seams the pictures skip.

1 · The problem

One character in. Two nonsense characters out.

José what you typed stored 4A 6F 73 C3 A9 every byte arrives intact read back José by a program expecting one byte each Nothing was corrupted. Two programs disagreed about what the bytes mean. the single character é is two bytes, C3 A9 — and it is the running example for this whole post The encoding is the contract. Someone broke it.
To see why that contract is shaped so strangely, you have to rewind to a world where a string was a pointer to bytes ending in a zero, and one byte meant one character. UTF-8’s every awkward property is a concession to that world.

2 · The obvious answer, and why it lost

Give every character four bytes. Watch everything break.

the letter H, in a fixed 4-byte encoding 00 00 00 48 three padding bytes, one real one — for every English letter clean to index character i sits at byte i × 4 and four times the size the disk groans, the wire groans and fatal in C those 00 bytes are end-of-string A C string ends at its first zero byte, so every string truncates at its first letter. Every existing program would have to be rewritten. That was the whole problem.
UTF-8’s design brief wasn’t elegance. It was: add a million characters without breaking anything that already exists — every filesystem, every network protocol, every C library that walks bytes until it hits a zero.

3 · The trick

Put the length in the high bits, and mark every follower.

0xxxxxxx 1 byte — and identical to ASCII 110xxxxx 10xxxxxx 2 bytes 1110xxxx 10xxxxxx 10xxxxxx 3 bytes 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx 4 bytes count the leading 1s and that is the length every follower starts 10 so it can never look like a start So from any byte in the stream you can find where the character began. scan backwards at most three bytes for one that doesn’t start 10. that property is called self-synchronising. A corrupted byte costs you one character, not the rest of the file.
Our é is code point 233 — eight significant bits, one more than a single byte can carry — so it takes the two-byte form and comes out C3 A9. Both bytes are above 0x7F, so neither can be mistaken for an ASCII character.

4 · Why old code survived

Three ranges of byte, and none of them overlap.

00–7F — ASCII, one byte each 80–BF — followers C2–F4 — leaders C0, C1 and F5–FF never appear at all So searching for the ASCII byte '/' can never hit the middle of a character. which is the magic that let UTF-8 slot into existing C code without rewriting it — byte-oblivious code mostly keeps working the trade: s[5] is no longer “the 6th character”. to find that you walk from the start.
In practice the lost indexing matters less than people fear, because most string work — search, concatenate, split on a delimiter, send over a socket — never needs character-indexed access. It just needs bytes that round-trip cleanly.

5 · Keep this card

The whole thing on one index card.

UTF-8 = variable width, 1 to 4 bytes + ASCII as a literal subset + a self-synchronising byte pattern ∴ designed for compatibility, not for elegance
Picture to keep: a train of carriages where the first carriage has a flag on the roof saying how many carriages this train has, and every following carriage is painted “I am not a first carriage.” Walk up to any carriage in the yard and you can tell instantly whether you are at the front or in the middle — and walking backwards a step or two always finds the front.

Why it exists

You’ve seen a name come back wrong. You type José into a form, or open a CSV in the wrong program, and out comes José. One character became two, and the two are nonsense. That’s not corruption — every byte arrived intact. It’s two programs disagreeing about what the bytes mean.

Here’s what actually happened, and it’s the running example for this whole post: the single character é. In UTF-8 it’s two bytes, C3 A9. A program that assumes one byte per character reads them as two separate old-style characters — Ã and © — and prints both. The encoding is the contract, and someone broke it.

To see why that contract is shaped so strangely, rewind. Imagine it’s 1992. You’re writing C code. A string is a char* — a pointer to bytes that ends with a 0. Every library, every kernel call, every config file, every protocol assumes this. ASCII fits in 7 bits, so one byte per character, and the world holds together.

Then you want to support Japanese. And Arabic. And Cyrillic. And you discover there are roughly a million possible characters in the world’s writing systems once you count CJK ideographs and historic scripts. A byte isn’t enough. So what do you do?

The obvious answer — make every character 4 bytes — is what UTF-32 does. It’s clean. Index i of the string is at offset i * 4. No ambiguity. The problem: every existing program that reads bytes, every filesystem, every network protocol, every C library breaks. A file that used to be 1KB is now 4KB. And 90% of the bytes in an English document are zeros, because ASCII characters only need 7 bits but you’re padding them out to 32. The disk groans. The wire groans. And worst of all, a \0 byte appears inside every English character — so strlen, which stops at the first zero byte, reports a length of 0 or 1 and every string in your program is silently truncated at its first letter.

UTF-8 is the answer to “how do we add a million characters without breaking anything that already exists.” It exists because backwards compatibility with ASCII and with byte-oriented C code wasn’t a nice-to-have — it was the only way the new encoding could possibly win adoption.

Why it matters now

Nearly every web page, JSON payload, source file and commit message you’ve touched this week is UTF-8. The WHATWG HTML Standard requires it for conforming documents. JSON exchanged between systems must use it (RFC 8259 §8.1). Linux filenames are byte sequences that are almost always interpreted as UTF-8. Go and Rust string literals are UTF-8 by definition. Even Windows, which spent two decades on UTF-16 internally, has been quietly migrating APIs toward UTF-8.

This means a software engineer who doesn’t understand UTF-8 will eventually hit a bug they can’t explain — a string that reverses wrong, a length that’s off, a regex that matches the middle of a character. AI-era engineers hit this constantly because LLM tokenizers operate on byte-pair-encoded UTF-8, not on “characters” in any human sense. (See tokenization.)

The short answer

UTF-8 = variable-width encoding (1–4 bytes) + ASCII as a literal subset + self-synchronizing byte pattern

Picture to keep: a train of carriages where the first carriage has a flag on the roof saying how many carriages this train has, and every following carriage is painted “I am not a first carriage.” Walk up to any carriage in the yard and you can tell instantly whether you’re at the front of a train or in the middle of one — and walking backwards a step or two always finds the front.

UTF-8 encodes each Unicode code point in 1, 2, 3, or 4 bytes. The first 128 code points (ASCII) are encoded as a single byte identical to ASCII — so every ASCII file is already valid UTF-8. Non-ASCII characters use multi-byte sequences whose bytes are all in the range 0x80–0xFF, which means they never collide with ASCII control characters or with \0.

How it works

The encoding rule is almost suspiciously simple. Look at the high bits of each byte:

0xxxxxxx                             → 1 byte,  7 bits of payload   (ASCII)
110xxxxx 10xxxxxx                    → 2 bytes, 11 bits
1110xxxx 10xxxxxx 10xxxxxx           → 3 bytes, 16 bits
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx  → 4 bytes, 21 bits

A leading byte tells you how long the sequence is. A continuation byte always starts with 10. The high-bit pattern of the leading byte (0, 110, 1110, 11110) is also the count of leading 1s, which is the byte length. From any byte in the stream, you can find the start of the current character by scanning backward at most 3 bytes for one that doesn’t start with 10. That property — self-synchronizing — means a corrupted byte loses one character, not the whole rest of the file.

Run our é through it. Its code point is U+00E9 = decimal 233 = binary 11101001 — eight significant bits, one more than the seven a single UTF-8 byte can carry, so it takes the 2-byte form: 110xxxxx 10xxxxxx. Right-align the value in the eleven x slots (00011 101001) and you get 11000011 10101001 = C3 A9. Both bytes are ≥ 0x80, so neither can be mistaken for an ASCII character; the second starts with 10, so it can never be mistaken for the start of anything. The mojibake at the top of this post is what you get when a reader ignores all of that and treats each byte as its own character.

The clever bit: continuation bytes are 10xxxxxx, which is 0x80–0xBF. A leading byte for a multi-byte sequence is 0xC2–0xF4 — RFC 3629 says outright that C0, C1 and F5–FF never appear in well-formed UTF-8. ASCII bytes are 0x00–0x7F. These ranges don’t overlap. So you can search a UTF-8 string for the ASCII byte '/' (0x2F) using plain memchr and you will never get a false hit inside a multi-byte character. This is the magic that let UTF-8 slot into existing C code without rewriting it. Byte-oblivious code mostly keeps working.

The trade you make: indexing is no longer O(1). s[5] doesn’t mean “the 6th character” anymore — to find the 6th character you have to walk from the start, decoding one code point at a time. In practice this matters less than people fear, because most string operations (search, concatenate, split on a delimiter, send over a socket) don’t need character-indexed access. They just need bytes that round-trip cleanly.

The seam: “characters” is a lie

Here’s the honest part. Even after you decode UTF-8 into code points, “character” still doesn’t mean what you think. The user-perceived character é can be one code point (U+00E9) or two (U+0065 U+0301 — e plus a combining acute accent). Family emoji like 👨‍👩‍👧 are sequences of multiple code points joined by zero-width joiners. What humans call a “character” is a grapheme cluster, and finding grapheme boundaries requires a Unicode-aware library and a giant lookup table that ships with every new emoji release.

So “string length” has at least four reasonable answers: bytes (UTF-8), code units (UTF-16), code points (Unicode scalars), and grapheme clusters (what users count). Our é can be 2 bytes and 1 code point, or 3 bytes and 2 code points, and it looks identical either way. Bugs love this gap. UTF-8 didn’t cause it — Unicode itself did — but UTF-8 made the gap visible at the byte layer where most engineers live.

You started with UTF-8 = variable-width + ASCII subset + self-synchronizing. What did é add to that line? — + the design goal was never elegance, it was not breaking existing code. Every awkward property here — the wasted bits in the prefix, the loss of O(1) indexing, the four different string lengths — is the price of the one property UTF-32 couldn’t offer: a program written before Unicode existed keeps working, byte for byte, on English text.

Going deeper