Why LLMs can't count the r's in 'strawberry'
A model that can write a sonnet stumbles on a question a five-year-old gets right. The reason isn't intelligence — it's that the model never sees the letters.
On this page
The picture version
Five pictures for a reader who has only ever seen the screenshot. The prose below fills in the seams the pictures skip.
1 · The problem
A five-year-old gets this right. The sonnet machine doesn’t.
2 · The step you never see
Between “you type” and “the model reads” there is a shredder.
3 · Why it answers anyway
It half-remembers the spelling instead of reading it.
4 · The fixes, and the family
Give it letters, or give it a tool that has them.
5 · Keep this card
The whole thing on one index card.
Why it exists
You’ve probably tried it. You ask ChatGPT or Claude “how many r’s are in the word strawberry?” and watch a model that can debug your code, summarize a research paper, and write passable poetry confidently answer “two.” You correct it. It apologizes, recounts, and sometimes still gets it wrong. The screenshot has been a meme for years now, and it has been stubbornly hard to kill — new model generations keep half-fixing it rather than removing the cause.
The instinct is to call this a bug, or a sign that the model is “dumber than it looks.” Neither is quite right. The mistake isn’t in the model’s reasoning — it’s in what the model is allowed to look at. By the time the question reaches the network, the word “strawberry” isn’t there anymore. What arrives is a small handful of integers, and you’re asking it to count something inside a representation it doesn’t have.
This is the cleanest, most viral demonstration of a structural quirk in how every modern LLM reads text. It’s worth understanding because the same quirk shows up everywhere — arithmetic on long numbers, exact string edits, character-by-character transformations — and once you see it, half of the “weird LLM failure” genre stops looking weird.
Why it matters now
Essentially every chatbot, coding agent, customer-service bot and “AI summarizer” you’re likely to meet is built on a subword-tokenized model, so essentially all of them ship with this property. (The exceptions are real but rare in production: byte- or character-level architectures, and systems that quietly route character-level questions to a tool.) When users probe a model and find it failing on a kindergarten task, trust collapses faster than it should — because the failure looks like unreliability in general, when it’s actually a narrow, predictable artifact.
For people building on LLMs, this matters in a few specific places:
- Don’t ask the model to do character-level work directly. Counting letters, reversing strings, checking exact spelling, character-by-character edits. The model can fake it sometimes, but it’s the wrong tool. Use a code interpreter or a regex.
- “It got the easy thing wrong, can I trust the hard thing?” is a fair question with a non-obvious answer. A model that miscounts r’s may still be excellent at higher-level reasoning, because the two failure modes have different causes. The strawberry test isn’t a general intelligence test — it’s a tokenizer test.
- Reasoning models help, partially. Models trained to think out loud often catch this by spelling the word first, then counting. It’s not a fix to the input representation — it’s a workaround the model learned to apply. Sometimes it works; sometimes it doesn’t.
The short answer
LLM letter-counting failure = tokenizer hides letters + no character-level grounding
Picture to keep: you’re holding a word printed on paper; the model is holding a few sealed envelopes that were mailed in place of that word. Ask it to count the r’s and it’s guessing from what it remembers about envelopes — because it is not allowed to open them. Like that, except the envelopes aren’t opaque: the model has read a great deal about what’s inside envelopes like these, which is exactly why it answers “two” with confidence instead of saying it can’t see.
The model never receives the word “strawberry” as ten letters. It receives a short sequence of integer IDs that each stand for a chunk of the word. To count r’s it would have to recover the spelling from those chunks — and it was never explicitly trained to do that, only to predict plausible next tokens.
How it works
Here’s the naive model of what happens, and where it breaks. Naive model: you type a question, the model reads the question, the model answers. Where it breaks: there’s a step between “you type” and “the model reads,” and that step is lossy in exactly the way the question cares about.
Before any “thinking” happens, your text is run through a tokenizer — a fixed lookup that chops a string into pieces from a learned vocabulary of ~50k–200k entries. OpenAI publishes tiktoken, the byte-level BPE tokenizer library for its models, and most other frontier models use close cousins of the same approach.
For common English words, the tokenizer typically merges large chunks. “strawberry” gets split into a small number of subword pieces — often two or three, depending on the exact tokenizer and whether there’s a leading space. (The exact split isn’t quoted here on purpose: it varies by model, and Anthropic doesn’t fully publish Claude’s tokenizer, so “Claude splits it as X+Y+Z” isn’t a checkable claim. The OpenAI split is checkable in seconds with tiktoken or the public tokenizer playground.)
Whatever the exact pieces are, the important fact is what happens next: each piece becomes an integer ID, and that ID is looked up in a table to get a learned vector. The letters are gone at that point — nothing downstream ever gets them back. The string “strawberry” enters the tokenizer, but what reaches the first transformer layer is a handful of vectors that stand for chunks, something like [123, 4567, 890] after lookup. There are no letters in there. There is no r to count.
So when you ask “how many r’s are in strawberry?”, the model is being asked to answer a question about a representation it threw away at the door. The question makes sense to you, the human reading the prompt as characters. To the model, the prompt itself is a sequence of opaque chunk-IDs, and the word “strawberry” inside it is a couple of those chunks.
Why does the model so often get it almost right — answering “two” instead of throwing up its hands? The usual account is that sentences like “strawberry is spelled s-t-r-a-w-b-e-r-r-y” showed up somewhere in pretraining, leaving the model with fragmentary, indirect knowledge of how words spell. It can sometimes recall that knowledge, sometimes can’t, and sometimes recalls a slightly wrong version. So it produces a confident-sounding number that’s frequently off-by-one. That “fluent output from a partial, lossy memory” pattern is one of the ingredients in hallucination too — the family resemblance is real, though hallucination is a broader phenomenon with several other causes, so don’t read this as “the strawberry bug explains hallucination.”
A few seams worth seeing:
- It’s not really “hallucination” in the made-up-a-fact sense. The model isn’t fabricating a paper or inventing a quote. It’s failing at a perception task — what letters are in this token? — that it was never given the inputs to do well.
- Character-level models don’t have this problem. Architectures that operate directly on bytes or characters (ByT5, Charformer, more recent byte-level transformers) see every letter, so the specific blind spot disappears — they can still be wrong for all the ordinary reasons. They pay for the visibility in sequence length and compute. The case for going byte-native is real, but it hasn’t displaced tokenized models at the production frontier, and whether it eventually does is an open question in the field rather than a settled one.
- Tool use fixes it cleanly. Give the model a Python interpreter and ask it to count:
"strawberry".count("r")returns 3. Done. The fix isn’t smarter weights; it’s letting the model offload the character-level operation to something that actually sees characters. - Reasoning models partly mitigate it. Models trained to produce long chains of thought often handle this by first writing the word out letter-by-letter in their scratchpad — “s, t, r, a, w, b, e, r, r, y” — and then counting. Spelling out a word coaxes the tokenizer into emitting one-letter tokens (or close to it), which gives the rest of the forward pass actual letters to work with. It’s a behavioral workaround that lives in the chain of thought, not a fix to the input pipeline.
- The bug generalizes. “How many words in this paragraph?”, “reverse this string”, “what’s the 7th character?” — all the same family. Operations that need to see the substrate beneath the tokens are operations the model is structurally bad at.
Work one of those through, since asserting transfer isn’t the same as showing it. Ask a model to add two 12-digit numbers. Column addition is a character-level algorithm: you line up the ones digit with the ones digit and carry leftward. But a 12-digit number doesn’t arrive as twelve digits — it arrives as a handful of chunks, and the chunk boundaries in 847293617284 and 593018472936 have no reason to fall in the same places. So the model can’t line up the columns, because the columns aren’t in its input. Same diagnosis as strawberry, same two fixes: make it spell the digits out first (chain-of-thought), or hand the arithmetic to a calculator (tool use). And the same caveat: some arithmetic failures have nothing to do with tokenization — this is one mechanism among several, not a universal explanation for “LLMs are bad at math.”
The deep point: the strawberry question is a small, repeatable demonstration that the model and the user are looking at different objects. You see a string of letters. The model sees a sequence of subword IDs. Most of the time the gap doesn’t matter, because most language tasks don’t require seeing through the tokens. When the task does — counting, spelling, exact transformations — the gap is most of the story.
How much of it is the whole story is contested, and the contest cuts at this post’s thesis. Fu et al., Why Do Large Language Models (LLMs) Struggle to Count Letters? (arXiv, 2024), find that token frequency doesn’t explain the error pattern well, while the difficulty of the counting operation itself does — errors climb with repeated letters, which is exactly what “strawberry” has. Read that as a correction to the strong version of the claim: the tokenizer genuinely does remove the letters, and that much is mechanical, but “it’s only the tokenizer” is more than the evidence supports.
You started with letter-counting failure = tokenizer hides letters + no character-level grounding. What did this post add? — + it's a perception failure, not a reasoning failure. That distinction is the useful part: it tells you the fix is to change what the model can see (spell it out, hand it a code interpreter), not to reach for a smarter model. And it tells you what the strawberry screenshot does and doesn’t license you to conclude about everything else the model does.
Famous related terms
- Tokenization —
tokenization = learned vocabulary + deterministic split into subword pieces. The preprocessing step that throws the letters away. - BPE (Byte Pair Encoding) —
BPE = greedy merge of the most-common adjacent pairs. The algorithm behind many modern LLM tokenizers, OpenAI’s included. - Hallucination —
hallucination = next-token model + no built-in "I don't know" + a prompt the model can't actually answer. The strawberry miss is a tokenizer-flavored cousin of this. - How to spot hallucinations — practical heuristics; the strawberry test is one of the cheapest red flags for “this answer wasn’t grounded in something the model could actually see.”
- Chain-of-thought — letting the model spell things out before answering. The standard mitigation, and how reasoning models partly route around the tokenizer.
- Tool use / code interpreter —
tool use ≈ LLM + an external function it can call mid-generation. Hand off character-level work to something that actually sees characters. The clean fix. - Character-level / byte-level models —
byte-level model = transformer + bytes/characters as the input units. Architectures that skip tokenization entirely. No strawberry bug; longer sequences and higher compute cost.
Going deeper
- Sennrich, Haddow, Birch, Neural Machine Translation of Rare Words with Subword Units (2016) — the primary source; read it for why subword units beat whole words in neural NLP, which is the trade the strawberry bug is the far side of.
- Andrej Karpathy’s Let’s build the GPT Tokenizer, listed on his Neural Networks: Zero to Hero course page — the explainer, for “could I build one myself?” He writes a BPE tokenizer from scratch on camera.
- The OpenAI tiktoken repo — the rabbit hole: run your own words through it and find out which ones your model is blind inside.
Well-established: the mechanism itself — the tokenizer hides the letters, the model sees IDs, and character-level operations land outside the model’s input representation. Contested: how much of the letter-counting failure that mechanism accounts for, versus the difficulty of counting (see Fu et al. above). Not public: the exact token split for “strawberry” in any specific model — Claude’s tokenizer isn’t fully published. Run your own string through the tokenizer of the model you care about rather than trusting a number from a blog post.