Why does tokenization exist?
Computers can already read bytes. So why do language models insist on chopping text into these weird half-words first?
On this page
The picture version
Six pictures for a reader who has never wondered what a model actually reads. The prose below fills in the seams the pictures skip.
1 · The problem
It can write a compiler. It cannot count three letters.
2 · Two obvious answers
Whole words, or single bytes. Both fail — in opposite directions.
3 · The compromise
Pieces bigger than letters, smaller than words.
4 · How the pieces get chosen
Nobody picked them. A counting loop did.
5 · The bill
And then the letters stop existing.
6 · Keep this card
The whole thing on one index card.
Why it exists
Ask a chatbot how many rs are in “strawberry” and there’s a decent chance
it tells you two. A model that can write you a working compiler cannot
reliably count three letters in a nine-letter word. That looks like a
reasoning failure. It isn’t — it’s a perception failure, and the reason
lives one layer below anything you’d call thinking. The model never saw the
letters. It’s the running example for this whole post: hold “strawberry” in
your head and we’ll come back to it.
Here’s why it never saw them. A language model’s input layer is a lookup table: one row per token, each row a vector. Before any “thinking” happens, your text has to be turned into integer IDs that index into that table. The question is: what should those units be?
Two obvious answers fail in opposite directions.
One unit per word. Sounds clean. Falls apart immediately. Every typo, every plural, every hyphenation, every language other than English needs its own row. “run”, “running”, “ran”, “runs”, “Running” — five rows for one idea. The lookup table balloons to millions of rows, most of them seen only a handful of times during training, which is thin evidence to learn a good vector from. And the first time a user types a word the table has never seen, the model is blind: there’s no row to look up. This is the out-of-vocabulary problem, and it haunted NLP for decades.
One unit per byte. Also clean. Also fails, but in the
other direction. Now the vocabulary is tiny — 256 rows — but every
sentence is enormous. “strawberry” is one concept to a human and ten
bytes to a model, and the model has to re-discover that those ten symbols
form a word before it can do anything interesting with it. (It would,
however, have no trouble counting the rs.) Sequence length blows up,
context window
budget evaporates, and standard attention — whose compute cost grows
quadratically with sequence length — gets brutally expensive.
Tokenization is the compromise. Cut text into pieces that are bigger than
characters but smaller than words, chosen so that common stuff (the,
ing, tion, JavaScript) becomes a single token, and rare or unseen stuff
(xyzzy42, a new product name, an emoji nobody’s seen before) gracefully
falls back to several smaller tokens. Vocabulary stays manageable
(~30k–200k entries), sequences stay short, and — in the byte-level tokenizers
the GPT-family models use — nothing is ever truly out-of-vocabulary, because
in the worst case you can always spell it out one byte at a time.
Why it matters now
Once you start building on top of LLMs, tokenization stops being an academic detail and starts showing up in your bills, your bugs, and your benchmarks.
- Pricing is per-token. Every major LLM API charges by tokens in and tokens out. “How long is this prompt?” is not a question about characters or words — it’s a question about that model’s tokenizer. The same string costs different amounts on different models.
- Context windows are measured in tokens. “200k context” is 200,000 tokens, not characters. How much of your codebase actually fits depends on how token-efficient the tokenizer is for your content. Code, non-English languages, and structured data tokenize very differently from English prose.
- Weird model failures trace back here. The classic “how many
rs in strawberry?” miss happens partly because the model never sees the letters individually — it sees a couple of tokens and is being asked a question about a representation it doesn’t have. Same family of bug: arithmetic on long numbers, exact string manipulation, counting characters. - Odd inputs land in odd corners of the vocabulary. Unicode lookalikes, unusual whitespace, and rare “glitch” tokens tokenize into pieces the model saw very little of during training, and behavior there is less predictable. There is no good public measurement of how much real-world jailbreaking depends on this — treat it as a known rough edge, not a quantified attack surface.
If you ship anything LLM-shaped to production, tokenization is in the critical path of cost, latency, and correctness.
The short answer
tokenization = a learned vocabulary + a deterministic algorithm that splits any string into pieces from that vocabulary
Picture to keep: a set of pre-cut jigsaw pieces — common shapes like the
and ing are single big pieces, and anything unusual gets assembled from
smaller ones, down to individual bytes if it has to be. Where the jigsaw
analogy breaks: real jigsaw pieces are cut to follow the picture, and these
are cut to follow frequency in the training corpus. A piece boundary tells
you what was common, not what means something.
The vocabulary is built once, by scanning a giant corpus and greedily merging the byte pairs that co-occur most. At inference time, a fixed algorithm replays those learned merges in priority order until no more apply, starting from raw bytes. The model only ever sees the resulting sequence of integer IDs.
How it works
The dominant algorithm in modern LLMs is BPE. (The neighbouring names are easy to confuse: WordPiece is a sibling algorithm with a different merge criterion, SentencePiece is a library that can train either BPE or unigram models, and tiktoken is OpenAI’s implementation of byte-level BPE.) The training procedure is almost embarrassingly simple:
1. Start with the vocabulary = all individual bytes (or characters).
2. Count every adjacent pair of symbols in your training corpus.
3. The most frequent pair becomes a new symbol; add it to the vocabulary.
4. Apply that merge everywhere in the corpus.
5. Repeat until the vocabulary is the size you want (e.g. 100k).
You end up with a vocabulary that has the bytes at the bottom (so nothing is
ever unrepresentable), short common sequences in the middle (th, ing,
er), and whole common words or sub-words at the top (tokenization,
JavaScript, the). Note that leading space — most GPT-style tokenizers
treat the and the as different tokens, because spacing has to survive
the round trip back to text.
At inference, splitting a string is the merge process replayed: start from bytes, keep applying the highest-priority merges until no more apply. It’s deterministic, it’s fast, and it produces the same tokens for the same input every time.
One detail that recipe glosses over: GPT-style tokenizers don’t run BPE
loose across the whole string. They first cut the text into chunks with a
fixed regular expression — roughly, at word, number, punctuation and
whitespace boundaries — and then run byte-level BPE inside each chunk, so
a merge can never span two chunks. That’s the reason a leading space rides
along with the word it precedes, and the reason the cat can never become
one token no matter how often the pair occurs.
A worked example, GPT-style. The exact pieces depend on which tokenizer and which version you use — what’s stable is the shape, so check your own strings in a tokenizer playground rather than trusting these splits verbatim:
"tokenization is fun" → ["token", "ization", " is", " fun"] (4 tokens)
"tokenizashun is fun" → ["token", "iz", "ash", "un", " is", " fun"] (6)
"strawberry" → ["str", "aw", "berry"] (3 tokens)
"日本語" → ["日", "本", "語"] or several bytes each, depending on tokenizer
There’s the answer to the hook. Asked how many rs are in “strawberry,” the
model isn’t looking at ten letters. It’s looking at a handful of integer IDs,
and “how many rs are in this ID?” is a question about a representation it
doesn’t have. It can often get there anyway — models memorize spellings from
text that discusses them — but it’s reconstructing, not reading. (There’s
a whole post on that failure.)
A few things that surprise people the first time:
- Tokens are not morphemes. They’re whatever pairs happened to co-occur
often in the training corpus.
izationis one token because lots of English words end that way; the tokenizer doesn’t know it’s a suffix, it just knows those bytes show up together. - The same word tokenizes differently in different positions.
Helloat the start of a string andHellomid-sentence are usually different tokens. This is correct behavior, but it makes “count the tokens of this word” a slightly ill-posed question. - Non-English text usually costs more tokens per unit of meaning. The merges were learned from whatever the trainer’s corpus was heaviest in; scripts and languages that were underrepresented fall back to shorter, less efficient pieces. A Japanese sentence and its English translation can have very different token counts even carrying the same meaning — which means a different price and a different share of the context window for the same content. The size of the gap depends entirely on the tokenizer and the language pair, so measure yours rather than trusting a multiplier you read somewhere (including this sentence).
- You can’t casually swap a tokenizer. A trained model is welded to the vocabulary it was trained on: the embedding rows are indexed by token ID, so feeding it IDs from a different tokenizer means feeding it the wrong rows. Getting a model onto a new vocabulary means remapping embeddings and retraining, not editing a config line.
- There’s an active research thread on getting rid of tokenization entirely. Byte-level and “tokenizer-free” architectures (ByT5, Charformer, more recently work on byte-level transformers) try to operate directly on bytes, paying the longer-sequence cost in exchange for removing a brittle preprocessing step. As of writing, the production frontier is still tokenized; the case for going byte-native is real but hasn’t won, and which way it settles is an open question in the field rather than a decided one.
The deep reason tokenization is good enough to stay is that it pushes a
hard problem (segmenting text) out of the model and into a cheap, fixed
preprocessing step. The model gets to spend its capacity on the things it’s
uniquely good at — composing meaning across the sequence — and not on
re-deriving “these ten bytes are the word strawberry” on every forward
pass.
You started with tokenization = a learned vocabulary + a splitting algorithm. What did this post add? — + the letters stop existing after it runs. That’s the price of the compromise: the model gets short sequences and
a small vocabulary, and in exchange it goes permanently blind to anything
below the token, which is why it can write your compiler but not count your
rs.
Famous related terms
- BPE (Byte Pair Encoding) —
BPE = greedy merge of most-common pairs. The dominant tokenization algorithm; originally a 1994 compression idea (Gage), popularized for neural machine translation by Sennrich et al., 2016. - WordPiece —
WordPiece ≈ BPE + a likelihood-based merge criterion. The variant used by BERT. - SentencePiece —
SentencePiece = a trainer for BPE or unigram models + whitespace treated as a normal character. It takes raw sentences rather than pre-tokenized words (escaping spaces as▁), which is why it’s popular for languages that don’t put spaces between words. - Unigram LM tokenization —
unigram = pick a vocab that maximizes corpus likelihood under a unigram model. An alternative to BPE used in SentencePiece. - tiktoken —
tiktoken = OpenAI's fast byte-level BPE implementation— the reference for GPT-family token counts. - Embeddings —
embedding = learned vector representation of a discrete thing— what tokens get turned into after tokenization: the float vectors the model actually does math on. - Context window —
context window = max number of tokens a model can attend to in one pass— the budget tokenization spends against. Always measured in tokens, not characters. - OOV (out-of-vocabulary) —
OOV = a word at inference time that the training vocab never saw— the problem byte-level subword tokenization makes go away, since the worst case is always “spell it out in bytes.”
Going deeper
- Neural Machine Translation of Rare Words with Subword Units (Sennrich, Haddow & Birch, 2016) — the primary source for “where did subword tokenization come from, and what problem was it solving?” (their answer: rare and unseen words in translation).
- Andrej Karpathy’s Let’s build the GPT tokenizer — the explainer to watch if the question is “could I implement this myself?”; it builds BPE from scratch and the merge loop stops being abstract about ten minutes in.
- Rabbit hole: the tiktoken repo — answers “what do my prompts actually look like to the model?” faster than any amount of reading, because you can encode your own strings and count.
A note on what I’m sure of: the algorithmic shape (BPE-style merges, byte-level fallback, deterministic encoding) and the practical consequences (cost, context, non-English overhead, the strawberry-
rfamily of bugs) are well-established. The relative quality and adoption of specific tokenizers shifts model-by-model and year-by-year — verify against the current model card rather than memorize.