Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

How does an AI 'see' a video?

Upload an hour-long video to Gemini and ask what happens at minute 40 — it answers. But no model 'watches' anything. It reads a flipbook, and the flipbook is missing most of the pages.

AI & ML intermediate Aug 6, 2026 · updated Aug 11, 2026 · 13 min read

On this page

The picture version

The whole idea in seven pictures, for someone who has never thought about what a video is to a computer. The prose below has the numbers, the caveats, and the seams the pictures skip.

1 · The problem

It finds the right minute of an hour-long lecture.

one hour of lecture you “where does she start talking about black holes?” asks MODEL never plays the video answers MM:SS
Ask about minute 40 of an hour-long recording and a timestamp comes back. From the outside it looks like the model sat down and watched the footage. It didn't.

2 · The naive way

Every frame costs more than the whole context.

· · · hand it every frame of ten minutes 10 min × 30 fps = 18,000 frames × 258 tokens per frame ≈ 4.6 million tokens what ten minutes of video needs 4,600,000 tokens what the biggest production context holds 1,000,000 tokens
Ten minutes at 30 frames a second is 18,000 frames, and at roughly 258 tokens a frame that is about 4.6 million tokens — several times the largest production context windows, for ten minutes of footage. An hour is pure fantasy.

3 · The trick

So throw away twenty-nine frames out of thirty.

one second of the lecture — 30 nearly identical frames keep 1, drop 29 1 fps ten minutes of footage: 18,000 → 600 frames and its bill: 4.6M → ~155,000 temporal redundancy a person mid-sentence looks the same at frame 301 and 302
Consecutive frames are nearly identical, so almost all of them can go — Gemini's documented default samples video at one frame per second. The content survives; every trace of smooth motion does not.

4 · The mechanism

The survivors become tokens, stacked in time order.

· · · 00:01 00:02 00:03 40:00 40:01 258 tokens per frame the model’s context, laid in time order look it up here “what happens at minute 40?”
Each surviving still goes through the ordinary image pipeline, and the blocks are laid into the context in time order, commonly with timestamps attached. That is what turns “minute 40” into a lookup rather than a search.

5 · The gap

Whatever happens between two stills never existed.

00:12 00:13 never sampled the sleight of hand, the wrist at impact, the collision one second
Anything that happens between two samples never reaches the context at all. The model isn't bad at fast motion — it never saw it, and no amount of extra resolution brings it back.

6 · The other stream

And the flipbook has no sound.

FRAMES silent — there is no sound in a picture tokenized image tokens SOUNDTRACK either or 32 tokens per second systems that tokenize audio too nothing many vision-language models
The sampled frames carry no sound, so knowing what was said depends on a second stream being tokenized alongside them — Gemini does it at 32 tokens per second; plenty of open models don't do it at all. When a video answer seems deaf, it often literally is.

7 · Keep this card

The whole thing on one index card.

VIDEO = ~1 frame per second + each frame tokenized like an image + laid in time order, with timestamps + a separate audio stream, or silence — the one people forget
Picture to keep: a contact sheet — one printed still per second of footage, in order, with the clock time written under each — and, if you're lucky, a transcript stapled to the back.

Why it exists

You upload a recording of an hour-long lecture to Gemini and ask “where does she start talking about black holes?” — and it hands you a timestamp. You ask your phone’s photo app for “the clip where the dog jumps into the pool” and it finds the exact video. From the outside it looks like the AI sat down and watched the footage.

But think about what a video actually is. Even a single image is expensive for a model: one frame becomes a few dozen to a few hundred tokens in the model’s context window. Video is that, thirty times per second. A ten-minute clip at 30 fps is 18,000 frames; at roughly 258 tokens a frame that’s about 4.6 million tokens — several times more than the largest production context windows, for ten minutes of footage. An hour would be pure fantasy. So the puzzle is real: how does a machine that can’t afford to look at every frame answer questions about an hour of video?

The answer is that it never looks at every frame. In the recipe that production systems document, the model reads video the way you’d skim a flipbook with most pages torn out: keep about one frame per second (Gemini’s documented default), turn each surviving frame into image tokens, and lay those token blocks into the context in time order, often with timestamps written next to them. “Watching” is attention over a sparse, ordered stack of stills. That lecture recording is the example this post keeps coming back to.

Why it matters now

Video input is now ordinary: Gemini takes hour-long uploads, ChatGPT’s Advanced Voice mode can share your phone’s camera, and computer-use agents work from streams of screenshots — video in all but name. Three places the flipbook mechanic shows up concretely:

The short answer

video input = ~1 frame per second sampled from the clip + each frame tokenized like an image + the blocks laid in time order (with timestamps) in one sequence

Picture to keep: a contact sheet — one printed still per second of footage, laid out in order with the clock time written under each — spread on a desk where the model can look at any of them at once. Not a screen playing; a wall of stills with the gaps between them simply gone.

That’s the whole trick. A video is decomposed into a sparse sequence of stills; each still goes through the exact image pipeline you may already know — patches → vision encoder → projector → image tokens; and the resulting blocks sit in the context one after another, so attention can relate “the man picks up the ball” at 00:12 to “the dog has the ball” at 00:19. The compression line is lossy in two honest ways: some systems also interleave audio tokens from the soundtrack, and some research models tokenize small stacks of frames together instead of single stills. Both refinements are covered below.

How it works

A video was never anything but frames (plus a soundtrack)

There is no “video” data type to see. Once decoded for playback, a video is a stream of still images — typically 24 to 60 per second, which your brain fuses into motion — plus an audio track stored alongside. (On disk, codecs mostly store differences between frames rather than full images; the full frames are reconstructed at decode time.) So “seeing video” reduces to two already-solved problems: turning images into tokens, and (optionally) turning audio into tokens. The only genuinely new problem is scale: there are far too many frames to afford.

Break 1: every frame is unaffordable → keep one per second

The naive design — decode the lecture and hand the model all 108,000 frames — died back in the hook, on arithmetic alone. The saving grace is that consecutive frames are nearly identical — a person mid-sentence looks the same at frame 301 and frame 302. This temporal redundancy is the same reason video files compress so well, and it means most frames can be dropped with little loss of content, even though all sense of smooth motion dies.

Production systems are unusually public about this step. Google’s Gemini docs state that video is sampled at 1 frame per second by default, each frame costing 258 tokens at standard resolution (66 at low resolution), with a 1M-token context fitting about an hour of video — three hours in low-resolution mode. Treat these as the right order of magnitude rather than a formula: Google’s own pages quote slightly different per-second figures in different places (~300 tokens/second on the video page, 263 in the token guide), and the numbers move as the models do. Run the arithmetic and the fantasy becomes tractable: that ten-minute clip drops from 18,000 frames to 600, from ~4.6M tokens to ~155,000 visual tokens (plus another ~19,000 if the soundtrack is tokenized too — more on that below). Still expensive — video is the most token-hungry thing you can put in a context — but possible.

Break 2: a still isn’t tokens either → the image pipeline, unchanged

Sampling leaves 3,600 stills of the lecture, which the model still can’t read — a transformer consumes vectors, not pixels. Fortunately this problem was already solved for single images. Every sampled frame now goes through the standard image pipeline: cut into patches, encoded by a vision transformer, projected into the language model’s embedding space. This post won’t re-derive that machinery — the image post covers it — but one consequence matters here: everything that’s true of image input is true of every frame. Tiny on-screen text smaller than a patch gets smeared; low-resolution mode makes that worse. A blurry frame yields blurry tokens.

Break 3: a bag of stills isn’t a video → time gets written back in

Tokenize each frame independently and you’ve thrown away the thing that made it footage. “Where does she start talking about black holes?” is unanswerable from an unordered pile of lecture stills, however many there are — order and timing are the question. Two things restore them. First, the frame blocks are laid into the context in chronological order, so sequence position itself carries “before” and “after,” the same way word order does. Second, systems commonly attach explicit timestamps to the frames — Gemini’s docs have you reference moments as MM:SS, which only works because the model can associate frames with clock time. That’s also how “what happens at minute 40?” becomes answerable: the frames sitting near the 40:00 labels are right there in the context to be looked up.

There’s a more radical research answer worth knowing about — a different family of models from the flipbook pipeline above. Video transformers from 2021 tokenize the video natively instead of sample-then-stacking: ViViT extracts spatio-temporal tokens, in one variant cutting the clip into tubelets, little space-time boxes, so a single token can carry motion, while TimeSformer keeps per-frame patches but splits attention in two — one pass across time between frames, another across space within a frame. That factorization exists because attention cost grows with the square of sequence length, and video-length sequences are enormous. A caveat, stated plainly: the internals of closed production models aren’t public, so whether any given assistant uses pure frame sampling or something tubelet-shaped under the hood is not something I can verify. The frame-sampling account above is what the public docs and the open-model recipes describe.

Break 4: the frames are silent → audio is a separate stream or nothing

Here’s the one that catches people out with a lecture recording. Everything above is a picture pipeline. Sampled frames carry no sound. If the system wants the model to know what was said, it must tokenize the audio track separately — Gemini does, at a documented 32 tokens per second, alongside the visual stream. Plenty of vision-language models, especially open ones, process only the frames — which produces a disorienting failure mode: the model describes the argument in the video perfectly and has no idea what either person said. When a video answer seems deaf, it often literally is.

Why the failure modes look the way they do

You started with video = frames sampled + tokenized like images + laid in time order. What did the failure chain add? — + a separate audio stream, or silence, and that’s the one people forget. Every other limit in the list above follows from the sampling rate, but “it didn’t know what she said” follows from a stream that may never have been tokenized at all.

Check yourself

Before you go — you ask about a two-minute clip of a card trick and the model describes the magician’s patter and hands perfectly, but gets the crucial sleight-of-hand wrong. Would uploading the same clip at 4K help?

Answer

No. Resolution and sampling rate are different knobs, and this is a sampling-rate failure: the sleight happened between two one-second samples, so no version of that frame exists in the context at any resolution. Higher resolution buys detail within a frame it already has — which is why it would help with small on-screen text and not here. The only fix is getting more frames around the sleight — sampling that stretch more densely, if your API exposes that control, or trimming the upload to those few seconds so a denser sample stays affordable. Note the patter came through: that tells you audio was tokenized, which is a separate stream from the frames that missed the trick.

And one more — you upload a three-hour recording and the model answers well about the opening and the closing but gets vague about the middle. Two different mechanisms in this post could each cause that. What are they, and how would you tell them apart?

Answer

One is lost-in-the-middle: recall degrades for material buried in the middle of a long context, regardless of what that material is. The other is resolution — three hours only fits by dropping to the low-resolution tier (66 tokens per frame instead of 258), so every frame is coarser, middle included. To separate them you have to change one variable at a time: cut the middle hour out, upload it alone, and pin the resolution setting to whatever the three-hour upload used. If answers get sharp with resolution held constant, the problem was position in a long context. If they stay vague, the detail was never encoded finely enough to begin with. (Upload the middle hour at standard resolution and you’ve changed both knobs at once — it’ll get better, and you’ll have learned nothing about which one mattered.)

Going deeper