How does an LLM 'see' an image?
You paste a screenshot into ChatGPT and it reads the text, describes the scene, answers questions. But the model only ever predicts text tokens — so how does a picture get into it at all?
On this page
The picture version
Six pictures for a reader who has never thought about how a photo gets into a chatbot. The prose below fills in the seams the pictures skip.
1 · The problem
Your screenshot is pixels. The model only reads numbers.
2 · The naive way
One token per pixel dies on arithmetic.
3 · The key idea
Cut the picture into squares and call each one a word.
4 · Where that breaks
A single square means nothing on its own.
5 · The hidden cost
The grid is cut blind, so small detail gets smeared.
6 · Keep this card
Image tokens sit in the sentence, and the model just reads on.
Why it exists
You drag a screenshot of an error message into ChatGPT and it reads the stack trace back to you — file names, line numbers, the lot. You photograph a fridge and ask what you can cook. You paste a chart and it pulls the trend out. From the outside it feels like the model looked at the picture, and probably ran some text-recognition step on it along the way.
Neither of those is what happened. Hold on to that stack-trace screenshot; it’s the example this post keeps coming back to.
But step back to what an LLM actually is: a machine that takes a sequence of tokens and predicts the next one. Its whole universe is a list of integer IDs drawn from a fixed vocabulary of words and word-pieces. A photo is none of those things — it’s a grid of millions of pixel brightnesses. So there’s a real puzzle here: how does something with no notion of “pixel” end up answering questions about a picture?
The trick is almost suspiciously simple. You don’t teach the transformer to see. You turn the image into the only thing the transformer understands — vectors in a sequence — and drop them into the context window right next to the word vectors. The model then does the exact same thing it always does: attention over a sequence, predict the next token. What looks like “seeing” is image-shaped tokens flowing through the very same machinery that handles text.
Why it matters now
Vision is no longer a bolt-on. The widely used general-purpose models — GPT-4o, Claude, Gemini, the open Qwen-VL and Llama families — are natively multimodal: the same model that writes your code also reads the screenshot of the bug. Three concrete places this shows up:
- Token cost balloons on images. Because an image becomes a block of tokens, a single high-resolution screenshot can cost more than a long paragraph of text — and you pay for it on every turn it stays in context. Knowing an image is tokens is how you predict your bill.
- The failure modes are specific and repeatable. Vision models miscount objects, fumble exact spatial relationships, and misread tiny text. Those aren’t random — they fall straight out of how the image gets chopped up, covered below.
- “Read this document” is now the same pipeline as “describe this photo.” Screenshot-to-answer, receipt scanning, UI agents that click around a screen — they all ride on the same image-into-tokens machinery, so its limits are their limits.
The short answer
image input = picture cut into patches → each patch becomes a vector → those vectors are fed to the transformer as tokens, right beside the words
Picture to keep: your screenshot cut up like a sheet of postage stamps, each stamp swapped for a word — then that pile of made-up words shuffled into the sentence you typed, and the whole thing read as one line of text.
A picture is split into a grid of small fixed-size squares called patches. Each patch is flattened and run through a small learned function that turns it into a vector — the same kind of vector a word gets from the embedding table. Those vectors are placed in the sequence alongside the text vectors, and from that point on the transformer handles them exactly like word vectors — there is no separate “this one is an image” pathway. It just runs attention over all of them and predicts text. That’s the whole idea; the rest is detail about how the patch becomes a good vector, and how it gets moved into the language model’s own vector space.
How it works
Break 1: one token per pixel dies on arithmetic → patches
The naïve idea — feed the model one token per pixel — dies immediately. Your stack-trace screenshot, at a modest 512×512, is 262,144 pixels. Since attention cost grows with the square of the sequence length, a quarter-million-token image would be wildly out of reach. So instead the image is carved into a grid of patches — say 16×16 pixels each — and each patch becomes one token. Now that screenshot is a 32×32 grid: 1,024 tokens, not a quarter million. This is the core move from the 2020 paper that made the patch the standard unit for vision transformers, titled — literally — An Image Is Worth 16×16 Words.
Each patch is flattened from its little square of pixel values into a long list of numbers, then multiplied by a learned matrix that projects it down to the model’s vector width. That’s the patch embedding: the visual analogue of looking a word up in the embedding table — except that a word is a unit someone chose to be meaningful, while the patch grid is cut on a blind fixed lattice. A word boundary respects the word; a patch boundary will happily slice a letter in half. That mismatch is where most of the failure modes below come from. One extra ingredient is added — a positional encoding — because a bare set of patch vectors has lost all sense of where each patch sat, and “the cat is above the dog” depends entirely on that.
To the computer it’s just a grid of pixel brightnesses — no words, no objects, no “cat.”
Carve the grid into fixed-size squares — e.g. 16×16 pixels. Each square is one “visual word.”
A vision encoder mixes the patches together; a projector maps them into the LLM’s vector space. Out come image tokens.
The image tokens sit in the context right beside the words of your question. The transformer attends over all of them and predicts text — same loop as always.
Break 2: a lone patch means nothing → encoder, then projector
Patch embeddings on their own are raw, and they break in two separate ways. One 16×16 square of your screenshot might contain three dark strokes. Is that part of a letter, a window border, or a code-editor bracket? The patch alone can’t say — meaning lives in the neighbourhood, not the square. And even once you know, the vector is in a space the vision network invented, not one the language model has ever read. So two things usually happen.
First, a vision encoder — almost always a ViT, a transformer that runs attention over the patches — lets every patch look at every other patch. A patch that’s part of an eye gets context from the patches that form the rest of the face. The output is one context-aware vector per patch. LLaVA-style systems typically use an encoder pretrained by CLIP, so its vectors already sit in a space shaped by matching captions — not the language model’s space, which is what the next step is for, but a much better starting point than raw pixels. That’s a common choice, not a requirement of the recipe.
Second, a projector (often called a connector or adapter) maps those vision vectors into the language model’s own embedding space, producing the vectors I’ve been calling image tokens. The honest summary is that there isn’t one standard projector — this is the part that differs most across models:
- A plain projection. The simplest approach applies a small network to each patch vector independently. The original LLaVA used a single learned linear projection from the vision features into the LLM’s embedding space (later LLaVA versions swapped in a two-layer MLP). Either way it’s sequence-preserving: one visual feature in maps to one image token out.
- A resampler / cross-attention. Some designs (Flamingo’s perceiver resampler, BLIP-2’s Q-Former) use a fixed, small set of learned query vectors that attend to the patches and pull out a fixed number of image tokens, regardless of how many patches went in. This caps the token count. BLIP-2’s Q-Former does more than resample — it’s the trained bridge between a frozen vision encoder and a frozen LLM — but the token-capping behaviour is the part that matters here.
Either way, the result is a handful to a few hundred vectors sitting in the LLM’s vector space. They get concatenated with the embedded text tokens, and the combined sequence flows through the transformer. From here it is exactly the decode loop you already know: attention mixes information across the whole sequence — letting a word like “What” pull from the image tokens — and the model predicts text one token at a time.
A caveat worth stating plainly: the exact architecture inside closed models like GPT-4o or Claude isn’t public. The patch → encoder → projector → shared sequence shape above is the well-documented open-model recipe (LLaVA, Qwen-VL, Flamingo, BLIP-2) and the standard account of how these systems work; treat the specific connector as “one of these families,” not a claim about any one vendor’s internals.
Why the failure modes look the way they do
Once you see the pipeline, the quirks stop being mysterious:
- Tiny text and fine detail get lost. Detail smaller than a patch can be smeared into a single vector. If a patch is 16 pixels wide and a line of text is 6 pixels tall, that text barely survives the projection. This is why models miss small captions, footnotes, or watermarks.
- High-resolution images get tiled or downscaled. To read detail without blowing the budget, systems often cut a big image into several tiles and run each through the encoder, then stitch the tokens together (LLaVA’s high-res variants do a version of this). More tiles means more tokens means more cost — which is why a detailed screenshot can cost hundreds to over a thousand tokens. How any given provider handles a large image is a moving target: resize-to-budget, tile, or preserve original dimensions are all choices vendors make differently and change between model versions. Check the current docs for the one you’re using rather than trusting a formula you read anywhere — including here.
- Counting and exact layout are hard. “How many people are in this photo?” asks the model to integrate information across many patches and keep a running count — something next-token prediction over patch vectors isn’t reliably good at, in the same family of weakness as why LLMs can’t count letters.
- There’s often no separate OCR step. In the open vision-language architectures above, reading text from an image emerges from the same patch-attention machinery, not a dedicated character-recognition module — which is why a model’s reading is fluent but occasionally hallucinates a character. (Some products do bolt a real OCR engine on top for documents, and the internals of closed models aren’t public, so this is a claim about the base VLM recipe, not every system you might use.)
You started with image input = patches → vectors → tokens beside the words.
What did the pipeline add that the compression line leaves out? — + an encoder that lets patches see each other, and a projector into the LLM's space. Which is what answers the question you arrived with: in the base
recipe there is no OCR step to point at. Reading your stack trace isn’t a
character-recognition module bolted on in front of the model — it’s
something the same attention machinery does on the way through, learned
rather than programmed. (How well it does it depends on training and on
how the image was preprocessed, not on the pipeline shape alone.)
Check yourself
Before you go — you screenshot a dense log at 4K and the model misreads a line number. You retake the screenshot cropped tightly to just that one line, same monitor, so the file is now much smaller. It reads it correctly. Why would fewer pixels help?
Answer
Because what matters isn’t total pixels, it’s how much of the model’s fixed visual budget your line number gets. A big screenshot has to be fitted into that budget somehow — downscaled, tiled, or both — and either way one line in a dense log ends up spanning very few patches, smeared together with the lines above and below it. Crop tightly and that same line spans many patches, each carrying a piece of a digit. You didn’t add information; you stopped spending the budget on 200 lines you didn’t care about. (The exact preprocessing differs by provider, so how much this helps varies — but the direction is a property of patching, not of any one vendor.)
And one more — a model reliably describes a photo’s contents but keeps getting “how many chairs are in this room?” wrong, and it’s wrong by different amounts each time. Does that point at the patching, or at something else?
Answer
Probably integration, though the honest answer is “both can do this.” A chair usually spans many patches, so unlike tiny text it isn’t lost at the patching step — the evidence is in the sequence. What’s hard is tallying it: counting means tracking which patches belong to the same chair, not double-counting one that spans several, and not missing one that’s half occluded behind a table. Next-token prediction has no counter to keep, which is the same weakness as counting letters in a word. The instability is the tell — a resolution problem gives you a consistent blind spot (the same distant chair is invisible every time), while an integration problem gives you a different wrong number on each attempt. But chairs that are small, cropped, or badly occluded fail for resolution reasons too, so the two aren’t cleanly separable in a real photo.
Famous related terms
- Multimodal model —
multimodal = one model that takes more than one input type (text + images) in the same sequence. The umbrella term for everything in this post. - ViT (Vision Transformer) —
ViT = transformer + self-attention over image patches instead of words. The encoder that turns patches into context-aware vectors. - Patch embedding —
patch embedding = flatten a square of pixels + project it into a vector. The image analogue of a word embedding. - CLIP —
CLIP = image encoder + text encoder trained so matching pairs land near each other. The common pretraining recipe that makes vision vectors “speak language.” - Projector / connector —
projector = small network mapping vision vectors into the LLM's embedding space. The seam where the two modalities are stitched together. - VLM (Vision-Language Model) —
VLM = vision encoder + projector + LLM. The full stack this post describes.
Going deeper
- An Image Is Worth 16×16 Words (Dosovitskiy et al., 2020) — the primary source, answering “where did ‘a patch is a word’ come from, and why did dropping convolutions work at all?”
- Visual Instruction Tuning / LLaVA (Liu et al., 2023) — the explainer, answering “what is the minimum you have to build to bolt vision onto an existing LLM?” — the answer being one small trained projection.
- Learning Transferable Visual Models From Natural Language Supervision / CLIP (Radford et al., 2021) — the rabbit hole, for “why can a vision encoder’s output line up with language before anyone trains a projector?”