Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

How does an LLM 'see' an image?

You paste a screenshot into ChatGPT and it reads the text, describes the scene, answers questions. But the model only ever predicts text tokens — so how does a picture get into it at all?

AI & ML intermediate May 20, 2026 · updated Aug 24, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never thought about how a photo gets into a chatbot. The prose below fills in the seams the pictures skip.

1 · The problem

Your screenshot is pixels. The model only reads numbers.

your stack-trace screenshot a grid of pixel brightnesses THE MODEL ID ID ID … it only ever reads whole numbers from a fixed vocabulary ? no obvious way in
A photo has no words in it and the model has no notion of a pixel. Everything below is about how the picture is turned into the one thing the model does understand.

2 · The naive way

One token per pixel dies on arithmetic.

512 × 512 your screenshot at a modest size one token per pixel 262,144 tokens, from one image verdict OUT OF REACH a quarter-million tokens, one image attention cost grows with length²
Attention cost grows with the square of the sequence length, so a quarter-million-token image is hopeless. The unit has to get bigger than a pixel.

3 · The key idea

Cut the picture into squares and call each one a word.

cut into 16×16-pixel patches (drawn 8×8; really 32×32) one patch flatten + project a vector, like a word's embedding the whole grid 512 ÷ 16 = 32 a side 32 × 32 = 1,024 tokens, not 262,144
Each 16×16 square is flattened and projected into a vector — the visual analogue of looking a word up in the embedding table. The same screenshot now costs 1,024 tokens instead of 262,144.

4 · Where that breaks

A single square means nothing on its own.

one patch, alone — part of a letter? — a window border? — a code bracket? meaning lives in the neighbourhood, not in one square. PATCH VECTORS raw, context-free attention VISION ENCODER every patch now sees every other patch map across PROJECTOR into the language model's own vector space out come image tokens
Two extra stages sit between the patch and the sequence: an encoder that lets patches see each other, and a projector that lands them in the language model's own space. Only then are they tokens the model can read.

5 · The hidden cost

The grid is cut blind, so small detail gets smeared.

log:417 a line of your stack trace and the blind patch lattice 16-px patch 6-px line of text inside it one vector everything in the square, smeared into one vector SMALL DETAIL GETS SMEARED
A word boundary respects the word; a patch boundary will happily slice a letter in half. That mismatch is where the missed footnotes, watermarks and tiny captions come from.

6 · Keep this card

Image tokens sit in the sentence, and the model just reads on.

your screenshot cut project place IMG IMG What is this ? image tokens and words, in one sequence patches → vectors → tokens beside the words
Picture to keep: your screenshot cut up like a sheet of postage stamps, each stamp swapped for a word — then that pile of made-up words shuffled into the sentence you typed, and the whole thing read as one line of text.

Why it exists

You drag a screenshot of an error message into ChatGPT and it reads the stack trace back to you — file names, line numbers, the lot. You photograph a fridge and ask what you can cook. You paste a chart and it pulls the trend out. From the outside it feels like the model looked at the picture, and probably ran some text-recognition step on it along the way.

Neither of those is what happened. Hold on to that stack-trace screenshot; it’s the example this post keeps coming back to.

But step back to what an LLM actually is: a machine that takes a sequence of tokens and predicts the next one. Its whole universe is a list of integer IDs drawn from a fixed vocabulary of words and word-pieces. A photo is none of those things — it’s a grid of millions of pixel brightnesses. So there’s a real puzzle here: how does something with no notion of “pixel” end up answering questions about a picture?

The trick is almost suspiciously simple. You don’t teach the transformer to see. You turn the image into the only thing the transformer understands — vectors in a sequence — and drop them into the context window right next to the word vectors. The model then does the exact same thing it always does: attention over a sequence, predict the next token. What looks like “seeing” is image-shaped tokens flowing through the very same machinery that handles text.

Why it matters now

Vision is no longer a bolt-on. The widely used general-purpose models — GPT-4o, Claude, Gemini, the open Qwen-VL and Llama families — are natively multimodal: the same model that writes your code also reads the screenshot of the bug. Three concrete places this shows up:

The short answer

image input = picture cut into patches → each patch becomes a vector → those vectors are fed to the transformer as tokens, right beside the words

Picture to keep: your screenshot cut up like a sheet of postage stamps, each stamp swapped for a word — then that pile of made-up words shuffled into the sentence you typed, and the whole thing read as one line of text.

A picture is split into a grid of small fixed-size squares called patches. Each patch is flattened and run through a small learned function that turns it into a vector — the same kind of vector a word gets from the embedding table. Those vectors are placed in the sequence alongside the text vectors, and from that point on the transformer handles them exactly like word vectors — there is no separate “this one is an image” pathway. It just runs attention over all of them and predicts text. That’s the whole idea; the rest is detail about how the patch becomes a good vector, and how it gets moved into the language model’s own vector space.

How it works

Break 1: one token per pixel dies on arithmetic → patches

The naïve idea — feed the model one token per pixel — dies immediately. Your stack-trace screenshot, at a modest 512×512, is 262,144 pixels. Since attention cost grows with the square of the sequence length, a quarter-million-token image would be wildly out of reach. So instead the image is carved into a grid of patches — say 16×16 pixels each — and each patch becomes one token. Now that screenshot is a 32×32 grid: 1,024 tokens, not a quarter million. This is the core move from the 2020 paper that made the patch the standard unit for vision transformers, titled — literally — An Image Is Worth 16×16 Words.

Each patch is flattened from its little square of pixel values into a long list of numbers, then multiplied by a learned matrix that projects it down to the model’s vector width. That’s the patch embedding: the visual analogue of looking a word up in the embedding table — except that a word is a unit someone chose to be meaningful, while the patch grid is cut on a blind fixed lattice. A word boundary respects the word; a patch boundary will happily slice a letter in half. That mismatch is where most of the failure modes below come from. One extra ingredient is added — a positional encoding — because a bare set of patch vectors has lost all sense of where each patch sat, and “the cat is above the dog” depends entirely on that.

A photo

To the computer it’s just a grid of pixel brightnesses — no words, no objects, no “cat.”

Cut into patches

Carve the grid into fixed-size squares — e.g. 16×16 pixels. Each square is one “visual word.”

Patch → vector (+ position)

Flatten each patch and project it into a vector — like looking a word up in the embedding table. A position signal records where it sat.

…
Encode & project → image tokens

A vision encoder mixes the patches together; a projector maps them into the LLM’s vector space. Out come image tokens.

IMG IMG IMG IMG …
One sequence: image tokens + words → answer

The image tokens sit in the context right beside the words of your question. The transformer attends over all of them and predicts text — same loop as always.

IMG IMG … What is this ?
A tabby cat on a sofa.
An image is cut into patches; each patch is projected into a vector; a vision encoder and a projector turn those into image tokens in the LLM’s embedding space; the image tokens join the text tokens in one sequence and the model predicts an answer. Grid sizes and counts are illustrative.

Break 2: a lone patch means nothing → encoder, then projector

Patch embeddings on their own are raw, and they break in two separate ways. One 16×16 square of your screenshot might contain three dark strokes. Is that part of a letter, a window border, or a code-editor bracket? The patch alone can’t say — meaning lives in the neighbourhood, not the square. And even once you know, the vector is in a space the vision network invented, not one the language model has ever read. So two things usually happen.

First, a vision encoder — almost always a ViT, a transformer that runs attention over the patches — lets every patch look at every other patch. A patch that’s part of an eye gets context from the patches that form the rest of the face. The output is one context-aware vector per patch. LLaVA-style systems typically use an encoder pretrained by CLIP, so its vectors already sit in a space shaped by matching captions — not the language model’s space, which is what the next step is for, but a much better starting point than raw pixels. That’s a common choice, not a requirement of the recipe.

Second, a projector (often called a connector or adapter) maps those vision vectors into the language model’s own embedding space, producing the vectors I’ve been calling image tokens. The honest summary is that there isn’t one standard projector — this is the part that differs most across models:

Either way, the result is a handful to a few hundred vectors sitting in the LLM’s vector space. They get concatenated with the embedded text tokens, and the combined sequence flows through the transformer. From here it is exactly the decode loop you already know: attention mixes information across the whole sequence — letting a word like “What” pull from the image tokens — and the model predicts text one token at a time.

A caveat worth stating plainly: the exact architecture inside closed models like GPT-4o or Claude isn’t public. The patch → encoder → projector → shared sequence shape above is the well-documented open-model recipe (LLaVA, Qwen-VL, Flamingo, BLIP-2) and the standard account of how these systems work; treat the specific connector as “one of these families,” not a claim about any one vendor’s internals.

Why the failure modes look the way they do

Once you see the pipeline, the quirks stop being mysterious:

You started with image input = patches → vectors → tokens beside the words. What did the pipeline add that the compression line leaves out? — + an encoder that lets patches see each other, and a projector into the LLM's space. Which is what answers the question you arrived with: in the base recipe there is no OCR step to point at. Reading your stack trace isn’t a character-recognition module bolted on in front of the model — it’s something the same attention machinery does on the way through, learned rather than programmed. (How well it does it depends on training and on how the image was preprocessed, not on the pipeline shape alone.)

Check yourself

Before you go — you screenshot a dense log at 4K and the model misreads a line number. You retake the screenshot cropped tightly to just that one line, same monitor, so the file is now much smaller. It reads it correctly. Why would fewer pixels help?

Answer

Because what matters isn’t total pixels, it’s how much of the model’s fixed visual budget your line number gets. A big screenshot has to be fitted into that budget somehow — downscaled, tiled, or both — and either way one line in a dense log ends up spanning very few patches, smeared together with the lines above and below it. Crop tightly and that same line spans many patches, each carrying a piece of a digit. You didn’t add information; you stopped spending the budget on 200 lines you didn’t care about. (The exact preprocessing differs by provider, so how much this helps varies — but the direction is a property of patching, not of any one vendor.)

And one more — a model reliably describes a photo’s contents but keeps getting “how many chairs are in this room?” wrong, and it’s wrong by different amounts each time. Does that point at the patching, or at something else?

Answer

Probably integration, though the honest answer is “both can do this.” A chair usually spans many patches, so unlike tiny text it isn’t lost at the patching step — the evidence is in the sequence. What’s hard is tallying it: counting means tracking which patches belong to the same chair, not double-counting one that spans several, and not missing one that’s half occluded behind a table. Next-token prediction has no counter to keep, which is the same weakness as counting letters in a word. The instability is the tell — a resolution problem gives you a consistent blind spot (the same distant chair is invisible every time), while an integration problem gives you a different wrong number on each attempt. But chairs that are small, cropped, or badly occluded fail for resolution reasons too, so the two aren’t cleanly separable in a real photo.

Going deeper