Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does tokenization exist?

Computers can already read bytes. So why do language models insist on chopping text into these weird half-words first?

AI & ML intro Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has never wondered what a model actually reads. The prose below fills in the seams the pictures skip.

1 · The problem

It can write a compiler. It cannot count three letters.

“how many r’s in strawberry?” “There are two.” a model that can write you a working compiler s t r a w b e r r y three of them, plainly there — to you the model never saw these letters something ran before it and took them away why? Not a reasoning failure. A perception failure. and it happens one layer below anything you would call thinking
The model’s input is not text. Before anything that resembles thinking, your string is chopped into units and turned into integer IDs — and the choice of unit is what decides which questions the model can even see.

2 · Two obvious answers

Whole words, or single bytes. Both fail — in opposite directions.

one row per word run running ran runs Running five rows for one idea — and millions of rows in total, most seen only a handful of times in training xyzzy42 — no row exists at all one row per byte a tiny vocabulary — only 256 rows, nothing ever missing s t r a w b e r r y one word → ten symbols the model must reassemble first sequences get long — and standard attention costs grow quadratically with sequence length context window evaporates Too many units, or too many of them per sentence.
Word units break on every typo, plural and unseen name — the out-of-vocabulary problem. Byte units never break, but blow up sequence length, which is the expensive axis. (Byte units would, however, have no trouble counting the rs.)

3 · The compromise

Pieces bigger than letters, smaller than words.

 the   JavaScript   tokenization whole common words — one piece each ing   tion   er   th short common fragments s  t  r  a  w  …  255 every single byte — the floor more common → bigger piece nothing is ever out-of-vocabulary worst case: spell it out one byte at a time Vocabulary stays manageable. Sequences stay short. roughly 30,000 to 200,000 entries in current models
Common things are cheap (one token) and rare things are merely expensive (several tokens) rather than impossible. The byte floor is what kills the out-of-vocabulary problem — there is always some way to spell a string, even one nobody has ever typed before.

4 · How the pieces get chosen

Nobody picked them. A counting loop did.

start: vocabulary = all 256 bytes count every adjacent pair in the corpus the most frequent pair becomes a new symbol apply it everywhere, then repeat — until the vocabulary is the size you wanted what that produces, watched on one word: t · o · k · e · n · i · z · a · t · i · o · n to · ke · n · iz · at · ion token · ization bytes, before any merge after the common pairs merge after the frequent ones merge again illustrative intermediate steps — the exact merges depend on the tokenizer and its training corpus
ization is one piece because a great many English words end that way — not because a suffix means anything. A piece boundary tells you what was common in the training corpus, not what carries meaning.

5 · The bill

And then the letters stop existing.

strawberry what you typed split str aw berry three pieces look up 3 integers what the model receives “how many r’s are in this integer?” a question about a representation the model does not have It can often answer anyway — by reconstructing, not reading. models memorise spellings from text that discusses them, which is why the failure is unreliable rather than total the split shown here is illustrative — check your own strings in a tokenizer playground
This is the price of the compromise, and it is not a bug anyone forgot to fix: short sequences and a small vocabulary are bought by going permanently blind to anything below the token. Same family of bug: arithmetic on long numbers, exact string surgery, counting characters.

6 · Keep this card

The whole thing on one index card.

tokenization = a learned vocabulary of pieces + a deterministic way to split any string into them ∴ after it runs, the letters are gone
Picture to keep: a set of pre-cut jigsaw pieces — common shapes like  the and ing are single big pieces, and anything unusual gets assembled from smaller ones, down to individual bytes if it has to be. Where the jigsaw analogy breaks: these pieces were cut to follow frequency, not the picture.

Why it exists

Ask a chatbot how many rs are in “strawberry” and there’s a decent chance it tells you two. A model that can write you a working compiler cannot reliably count three letters in a nine-letter word. That looks like a reasoning failure. It isn’t — it’s a perception failure, and the reason lives one layer below anything you’d call thinking. The model never saw the letters. It’s the running example for this whole post: hold “strawberry” in your head and we’ll come back to it.

Here’s why it never saw them. A language model’s input layer is a lookup table: one row per token, each row a vector. Before any “thinking” happens, your text has to be turned into integer IDs that index into that table. The question is: what should those units be?

Two obvious answers fail in opposite directions.

One unit per word. Sounds clean. Falls apart immediately. Every typo, every plural, every hyphenation, every language other than English needs its own row. “run”, “running”, “ran”, “runs”, “Running” — five rows for one idea. The lookup table balloons to millions of rows, most of them seen only a handful of times during training, which is thin evidence to learn a good vector from. And the first time a user types a word the table has never seen, the model is blind: there’s no row to look up. This is the out-of-vocabulary problem, and it haunted NLP for decades.

One unit per byte. Also clean. Also fails, but in the other direction. Now the vocabulary is tiny — 256 rows — but every sentence is enormous. “strawberry” is one concept to a human and ten bytes to a model, and the model has to re-discover that those ten symbols form a word before it can do anything interesting with it. (It would, however, have no trouble counting the rs.) Sequence length blows up, context window budget evaporates, and standard attention — whose compute cost grows quadratically with sequence length — gets brutally expensive.

Tokenization is the compromise. Cut text into pieces that are bigger than characters but smaller than words, chosen so that common stuff (the, ing, tion, JavaScript) becomes a single token, and rare or unseen stuff (xyzzy42, a new product name, an emoji nobody’s seen before) gracefully falls back to several smaller tokens. Vocabulary stays manageable (~30k–200k entries), sequences stay short, and — in the byte-level tokenizers the GPT-family models use — nothing is ever truly out-of-vocabulary, because in the worst case you can always spell it out one byte at a time.

Why it matters now

Once you start building on top of LLMs, tokenization stops being an academic detail and starts showing up in your bills, your bugs, and your benchmarks.

If you ship anything LLM-shaped to production, tokenization is in the critical path of cost, latency, and correctness.

The short answer

tokenization = a learned vocabulary + a deterministic algorithm that splits any string into pieces from that vocabulary

Picture to keep: a set of pre-cut jigsaw pieces — common shapes like the and ing are single big pieces, and anything unusual gets assembled from smaller ones, down to individual bytes if it has to be. Where the jigsaw analogy breaks: real jigsaw pieces are cut to follow the picture, and these are cut to follow frequency in the training corpus. A piece boundary tells you what was common, not what means something.

The vocabulary is built once, by scanning a giant corpus and greedily merging the byte pairs that co-occur most. At inference time, a fixed algorithm replays those learned merges in priority order until no more apply, starting from raw bytes. The model only ever sees the resulting sequence of integer IDs.

How it works

The dominant algorithm in modern LLMs is BPE. (The neighbouring names are easy to confuse: WordPiece is a sibling algorithm with a different merge criterion, SentencePiece is a library that can train either BPE or unigram models, and tiktoken is OpenAI’s implementation of byte-level BPE.) The training procedure is almost embarrassingly simple:

1. Start with the vocabulary = all individual bytes (or characters).
2. Count every adjacent pair of symbols in your training corpus.
3. The most frequent pair becomes a new symbol; add it to the vocabulary.
4. Apply that merge everywhere in the corpus.
5. Repeat until the vocabulary is the size you want (e.g. 100k).

You end up with a vocabulary that has the bytes at the bottom (so nothing is ever unrepresentable), short common sequences in the middle (th, ing, er), and whole common words or sub-words at the top (tokenization, JavaScript, the). Note that leading space — most GPT-style tokenizers treat the and the as different tokens, because spacing has to survive the round trip back to text.

At inference, splitting a string is the merge process replayed: start from bytes, keep applying the highest-priority merges until no more apply. It’s deterministic, it’s fast, and it produces the same tokens for the same input every time.

One detail that recipe glosses over: GPT-style tokenizers don’t run BPE loose across the whole string. They first cut the text into chunks with a fixed regular expression — roughly, at word, number, punctuation and whitespace boundaries — and then run byte-level BPE inside each chunk, so a merge can never span two chunks. That’s the reason a leading space rides along with the word it precedes, and the reason the cat can never become one token no matter how often the pair occurs.

A worked example, GPT-style. The exact pieces depend on which tokenizer and which version you use — what’s stable is the shape, so check your own strings in a tokenizer playground rather than trusting these splits verbatim:

"tokenization is fun"  →  ["token", "ization", " is", " fun"]   (4 tokens)
"tokenizashun is fun"  →  ["token", "iz", "ash", "un", " is", " fun"] (6)
"strawberry"           →  ["str", "aw", "berry"]                (3 tokens)
"日本語"                →  ["日", "本", "語"] or several bytes each, depending on tokenizer

There’s the answer to the hook. Asked how many rs are in “strawberry,” the model isn’t looking at ten letters. It’s looking at a handful of integer IDs, and “how many rs are in this ID?” is a question about a representation it doesn’t have. It can often get there anyway — models memorize spellings from text that discusses them — but it’s reconstructing, not reading. (There’s a whole post on that failure.)

A few things that surprise people the first time:

The deep reason tokenization is good enough to stay is that it pushes a hard problem (segmenting text) out of the model and into a cheap, fixed preprocessing step. The model gets to spend its capacity on the things it’s uniquely good at — composing meaning across the sequence — and not on re-deriving “these ten bytes are the word strawberry” on every forward pass.

You started with tokenization = a learned vocabulary + a splitting algorithm. What did this post add? — + the letters stop existing after it runs. That’s the price of the compromise: the model gets short sequences and a small vocabulary, and in exchange it goes permanently blind to anything below the token, which is why it can write your compiler but not count your rs.

Going deeper

A note on what I’m sure of: the algorithmic shape (BPE-style merges, byte-level fallback, deterministic encoding) and the practical consequences (cost, context, non-English overhead, the strawberry-r family of bugs) are well-established. The relative quality and adoption of specific tokenizers shifts model-by-model and year-by-year — verify against the current model card rather than memorize.