Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why LLMs can't count the r's in 'strawberry'

A model that can write a sonnet stumbles on a question a five-year-old gets right. The reason isn't intelligence — it's that the model never sees the letters.

AI & ML intro May 2, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Five pictures for a reader who has only ever seen the screenshot. The prose below fills in the seams the pictures skip.

1 · The problem

A five-year-old gets this right. The sonnet machine doesn’t.

“how many r’s are in strawberry?” “There are two.” “Are you sure?” — it apologises, recounts… …and sometimes still gets it wrong the same model, minutes earlier debugged your code summarised a research paper wrote passable poetry ✓ ✓ ✓ so this isn’t “dumber than it looks” The mistake isn’t in the reasoning. It’s in what the model is allowed to look at.
By the time the question reaches the network, the word “strawberry” isn’t there any more. What arrives is a small handful of integers — and you are asking about something inside a representation the model doesn’t have.

2 · The step you never see

Between “you type” and “the model reads” there is a shredder.

what you typed strawberry tokenizer a few subword chunks str aw berry look up what the network receives [ 123, 4567, 890 ] then a vector per id. no letters anywhere. the exact split varies by model — the shape is what matters “how many r’s are in 4567?” There is no r in there to count. the question makes sense to you, reading the prompt as characters. to the model the prompt is a sequence of opaque chunk-ids.
Nothing downstream ever gets the letters back. The model is being asked about a representation it threw away at the door — which is why this is a perception failure rather than a reasoning one.

3 · Why it answers anyway

It half-remembers the spelling instead of reading it.

somewhere in pretraining “strawberry is spelled s-t-r-a-w-b-e-r-r-y” the usual account, not a measured one fragmentary, indirect memory sometimes recalled correctly, sometimes not, sometimes slightly wrong so the number comes out confident and off-by-one It is reconstructing, not reading. and how much of the failure is the tokenizer is genuinely contested: Fu et al. (2024) find error rates track the difficulty of the counting itself, rising with repeated letters — which is exactly what “strawberry” has. the letters really are hidden; “only the tokenizer” is more than the evidence supports.
Fluent output drawn from a partial, lossy memory is a pattern this shares with hallucination — the family resemblance is real, though hallucination is broader and has several other causes. Don’t read this as “the strawberry bug explains hallucination.”

4 · The fixes, and the family

Give it letters, or give it a tool that has them.

spell it out first “s, t, r, a, w, b, e, r, r, y” coaxes the tokenizer into emitting one-letter tokens, so the rest of the pass has letters to work with a behavioural workaround, not an input-pipeline fix hand it to something that sees characters "strawberry".count("r") → 3 the clean fix. not smarter weights — different eyes. and the same diagnosis covers a whole family: “reverse this string” “what’s the 7th character?” adding two 12-digit numbers column addition needs the ones digit lined up with the ones digit — but a 12-digit number arrives as a handful of chunks whose boundaries have no reason to match same two fixes; same caveat, that some arithmetic failures have nothing to do with tokenization
Operations that need to see the substrate beneath the tokens are operations the model is structurally bad at. The fix is to change what the model can see, not to reach for a bigger model.

5 · Keep this card

The whole thing on one index card.

letter-counting failure = the tokenizer hides the letters + no character-level grounding to get them back ∴ a perception failure, not a reasoning one so the fix is to change what it can see, not to reach for a smarter model
Picture to keep: you’re holding a word printed on paper; the model is holding a few sealed envelopes mailed in place of that word. Where it breaks: the envelopes aren’t quite opaque — the model has read a great deal about what’s inside envelopes like these, which is why it answers “two” with confidence instead of saying it can’t see.

Why it exists

You’ve probably tried it. You ask ChatGPT or Claude “how many r’s are in the word strawberry?” and watch a model that can debug your code, summarize a research paper, and write passable poetry confidently answer “two.” You correct it. It apologizes, recounts, and sometimes still gets it wrong. The screenshot has been a meme for years now, and it has been stubbornly hard to kill — new model generations keep half-fixing it rather than removing the cause.

The instinct is to call this a bug, or a sign that the model is “dumber than it looks.” Neither is quite right. The mistake isn’t in the model’s reasoning — it’s in what the model is allowed to look at. By the time the question reaches the network, the word “strawberry” isn’t there anymore. What arrives is a small handful of integers, and you’re asking it to count something inside a representation it doesn’t have.

This is the cleanest, most viral demonstration of a structural quirk in how every modern LLM reads text. It’s worth understanding because the same quirk shows up everywhere — arithmetic on long numbers, exact string edits, character-by-character transformations — and once you see it, half of the “weird LLM failure” genre stops looking weird.

Why it matters now

Essentially every chatbot, coding agent, customer-service bot and “AI summarizer” you’re likely to meet is built on a subword-tokenized model, so essentially all of them ship with this property. (The exceptions are real but rare in production: byte- or character-level architectures, and systems that quietly route character-level questions to a tool.) When users probe a model and find it failing on a kindergarten task, trust collapses faster than it should — because the failure looks like unreliability in general, when it’s actually a narrow, predictable artifact.

For people building on LLMs, this matters in a few specific places:

The short answer

LLM letter-counting failure = tokenizer hides letters + no character-level grounding

Picture to keep: you’re holding a word printed on paper; the model is holding a few sealed envelopes that were mailed in place of that word. Ask it to count the r’s and it’s guessing from what it remembers about envelopes — because it is not allowed to open them. Like that, except the envelopes aren’t opaque: the model has read a great deal about what’s inside envelopes like these, which is exactly why it answers “two” with confidence instead of saying it can’t see.

The model never receives the word “strawberry” as ten letters. It receives a short sequence of integer IDs that each stand for a chunk of the word. To count r’s it would have to recover the spelling from those chunks — and it was never explicitly trained to do that, only to predict plausible next tokens.

How it works

Here’s the naive model of what happens, and where it breaks. Naive model: you type a question, the model reads the question, the model answers. Where it breaks: there’s a step between “you type” and “the model reads,” and that step is lossy in exactly the way the question cares about.

Before any “thinking” happens, your text is run through a tokenizer — a fixed lookup that chops a string into pieces from a learned vocabulary of ~50k–200k entries. OpenAI publishes tiktoken, the byte-level BPE tokenizer library for its models, and most other frontier models use close cousins of the same approach.

For common English words, the tokenizer typically merges large chunks. “strawberry” gets split into a small number of subword pieces — often two or three, depending on the exact tokenizer and whether there’s a leading space. (The exact split isn’t quoted here on purpose: it varies by model, and Anthropic doesn’t fully publish Claude’s tokenizer, so “Claude splits it as X+Y+Z” isn’t a checkable claim. The OpenAI split is checkable in seconds with tiktoken or the public tokenizer playground.)

Whatever the exact pieces are, the important fact is what happens next: each piece becomes an integer ID, and that ID is looked up in a table to get a learned vector. The letters are gone at that point — nothing downstream ever gets them back. The string “strawberry” enters the tokenizer, but what reaches the first transformer layer is a handful of vectors that stand for chunks, something like [123, 4567, 890] after lookup. There are no letters in there. There is no r to count.

So when you ask “how many r’s are in strawberry?”, the model is being asked to answer a question about a representation it threw away at the door. The question makes sense to you, the human reading the prompt as characters. To the model, the prompt itself is a sequence of opaque chunk-IDs, and the word “strawberry” inside it is a couple of those chunks.

Why does the model so often get it almost right — answering “two” instead of throwing up its hands? The usual account is that sentences like “strawberry is spelled s-t-r-a-w-b-e-r-r-y” showed up somewhere in pretraining, leaving the model with fragmentary, indirect knowledge of how words spell. It can sometimes recall that knowledge, sometimes can’t, and sometimes recalls a slightly wrong version. So it produces a confident-sounding number that’s frequently off-by-one. That “fluent output from a partial, lossy memory” pattern is one of the ingredients in hallucination too — the family resemblance is real, though hallucination is a broader phenomenon with several other causes, so don’t read this as “the strawberry bug explains hallucination.”

A few seams worth seeing:

Work one of those through, since asserting transfer isn’t the same as showing it. Ask a model to add two 12-digit numbers. Column addition is a character-level algorithm: you line up the ones digit with the ones digit and carry leftward. But a 12-digit number doesn’t arrive as twelve digits — it arrives as a handful of chunks, and the chunk boundaries in 847293617284 and 593018472936 have no reason to fall in the same places. So the model can’t line up the columns, because the columns aren’t in its input. Same diagnosis as strawberry, same two fixes: make it spell the digits out first (chain-of-thought), or hand the arithmetic to a calculator (tool use). And the same caveat: some arithmetic failures have nothing to do with tokenization — this is one mechanism among several, not a universal explanation for “LLMs are bad at math.”

The deep point: the strawberry question is a small, repeatable demonstration that the model and the user are looking at different objects. You see a string of letters. The model sees a sequence of subword IDs. Most of the time the gap doesn’t matter, because most language tasks don’t require seeing through the tokens. When the task does — counting, spelling, exact transformations — the gap is most of the story.

How much of it is the whole story is contested, and the contest cuts at this post’s thesis. Fu et al., Why Do Large Language Models (LLMs) Struggle to Count Letters? (arXiv, 2024), find that token frequency doesn’t explain the error pattern well, while the difficulty of the counting operation itself does — errors climb with repeated letters, which is exactly what “strawberry” has. Read that as a correction to the strong version of the claim: the tokenizer genuinely does remove the letters, and that much is mechanical, but “it’s only the tokenizer” is more than the evidence supports.

You started with letter-counting failure = tokenizer hides letters + no character-level grounding. What did this post add? — + it's a perception failure, not a reasoning failure. That distinction is the useful part: it tells you the fix is to change what the model can see (spell it out, hand it a code interpreter), not to reach for a smarter model. And it tells you what the strawberry screenshot does and doesn’t license you to conclude about everything else the model does.

Going deeper

Well-established: the mechanism itself — the tokenizer hides the letters, the model sees IDs, and character-level operations land outside the model’s input representation. Contested: how much of the letter-counting failure that mechanism accounts for, versus the difficulty of counting (see Fu et al. above). Not public: the exact token split for “strawberry” in any specific model — Claude’s tokenizer isn’t fully published. Run your own string through the tokenizer of the model you care about rather than trusting a number from a blog post.