Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why Git stores snapshots, not diffs

Git's reputation says 'version control = diffs.' Git's actual model says 'version control = snapshots, hashed.' That swap is the whole reason Git feels different.

Computer Science intro Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has used Git without ever wondering what it keeps. The prose below fills in the seams the pictures skip.

1 · The problem

Check out a commit from 2016. It happens instantly. It shouldn’t.

the mental model most people carry: a repo is a chain of changes to reach 2016, undo every change since — thousands of them, in order If that were true you would watch a progress bar. You don’t. It is instant, and so is switching back. so the mental model is wrong somewhere — and the place it is wrong is worth knowing
The chain-of-changes picture makes a clear prediction: reaching an old commit should cost time proportional to the history. Reality flatly contradicts it, which is the anomaly the rest of this post explains.

2 · What is actually stored

Not the changes. The whole contents of every file, every time.

three versions of one file, as Git keeps them the whole file,as it was in v1 the whole file,as it was in v2 the whole file,as it was in v3 a3f9c1…7b2e04…c81d5a… each label is a fingerprint computed from the contents themselves So reaching any version is a lookup, not a replay. That is the entire reason the checkout is instant.
Git stores complete contents, addressed by a fingerprint of those contents rather than by where they came from. Getting an old version means fetching one box, not unwinding a thousand steps.

3 · Why that isn’t as wasteful as it sounds

Identical contents get an identical fingerprint — so they are the same box.

commit 1commit 2commit 3 the files nobody touched stored once the one file that changed: three versions, three boxes You only pay for what actually changed — without storing a single difference.
Two files with identical contents produce the identical fingerprint, so they are the same stored object. Deduplication falls out of the addressing scheme rather than being a feature bolted on.

4 · So what is a commit?

A slip of paper listing which boxes were on the shelf that day.

one commit README.md → a3f9c1… main.py   → 7b2e04… utils.py  → c81d5a… and: the commit before this one It is not a copy of your project. It is a list of which versions were current at that moment. which is why a commit is tiny, and why the history is a chain of these lists rather than of changes
A commit records the fingerprint of every file version in that snapshot, plus a link to its predecessor. The snapshot is conceptual, not a duplicate — almost every entry points at a box that already existed.

5 · Keep this card

The whole thing on one index card.

a commit = a snapshot of the whole project + addressed by a fingerprint of its contents ∴ unchanged files cost nothing extra Git does compress old objects against each other on disk — but that is a storage optimisation underneath the model, not the model itself
Picture to keep: a warehouse where every box is labelled with a fingerprint of what is inside, not with where it came from — so two people storing identical contents get the same label, and therefore the same box. A commit is a slip of paper listing which boxes were on the shelf that day. Where it breaks: boxes here are never edited or thrown away, which is why “rewriting history” always means writing new ones.

Why it exists

Clone a repo with ten years of history, then check out a commit from 2016. It’s instant. Switch back to main. Instant. Create a branch. Instant — the progress bar doesn’t even get a chance to appear. Now compare that with the mental model most people carry in: “a repo is a chain of diffs, and checking out an old version means undoing every change since.” If that were true, jumping back four years would mean replaying thousands of patches, and it would feel like it.

It doesn’t feel like it, because that model is wrong. A Git commit names the complete state of the tree as it existed at that moment — every file, not a patch against the last one. When a file doesn’t change between two commits, Git doesn’t store a diff of zero bytes; it just points the new commit at the same file object the old commit was already pointing at. That’s why old commits are as cheap to reach as new ones: there is nothing to replay, just a pointer to follow.

The diff model isn’t a strawman — it’s how Subversion and CVS genuinely worked, which is where most people picked it up. Git threw it out, and that one swap is most of why Git feels qualitatively different from what came before.

We’ll follow one tiny repo the whole way down: two files, README.md and src/main.py, and two commits where only the README changes.

Why it matters now

Almost every codebase you’ll touch as a software engineer lives in Git. The mental model leaks into everything: why git log is fast, why branches are free, why git rebase can reorder history without corrupting it, why a detached HEAD is a normal state and not a disaster, why GitHub can show you any historical version of a file instantly without replaying patches.

It also matters because the snapshot model is what makes Git a content-addressed store — the same idea now showing up in IPFS, container image layers, Nix, and large-model weight stores. Git was an early mainstream example of “address things by what they are, not where they live,” and that idea keeps being rediscovered.

The short answer

git commit ≈ snapshot of the tree, addressed by the SHA-1 hash of its contents

Picture to keep: a warehouse where every box is labelled with a fingerprint of what’s inside it, not with where it came from. Two people who happen to store identical contents get the same label — so they’re the same box. A commit isn’t a copy of the warehouse; it’s a slip of paper listing which boxes were on the shelf that day. (Where the warehouse breaks down: boxes here are never edited or thrown away. Changing a file doesn’t open a box, it creates a new one with a new label — which is why “rewriting history” in Git always means writing new objects, never mutating old ones.)

A commit is a tiny object that points at a tree (a directory listing). The tree points at blobs (file contents) and at sub-trees. Every one of those objects is named by the hash of its own bytes. If two commits contain the same file, they end up pointing at the same blob — automatically, with no deduplication step.

There are no diffs in the storage model. Diffs are something Git computes on demand when you ask git diff or git log -p.

How it works

Try to invent this yourself, and watch each attempt fail.

Attempt 1: copy the whole project on every commit. Instant checkout, trivially simple, no patch replay. And obviously unusable: our repo’s src/main.py doesn’t change in commit 2, but you’ve now got two identical copies of it on disk. Ten thousand commits later you have ten thousand copies of every file that was never touched.

Fix 1: name every file by the hash of its own contents. Instead of “main.py, version 2,” store the raw bytes under a name computed from the bytes themselves — the SHA-1 of the content. Git calls that a blob: just bytes, no filename. Now duplicate content cannot be stored twice, because both copies compute the same name and land in the same place. There is no deduplication step to run; dedup is a consequence of the naming scheme. Add the same file to two repos in two countries and both call its blob the same name — the 40-character hex string you’re used to, in a SHA-1 repo, or 64 characters in a SHA-256 one.

But now nothing has a name or a path. A pile of content-addressed blobs isn’t a project — you can’t tell which blob was README.md. So: a tree, a tiny file listing entries (mode, name, hash), where each hash points at a blob or at another tree for a sub-directory. And a tree is itself content- addressed, which pays off immediately. Our repo:

T2 is a full snapshot of the project, and storing it cost one new blob and one new tree. The src/ subtree was unchanged, so it hashed to the same name, so it is the same object. The “diff” between the two commits is something Git derives at read time by walking T1 and T2 and noticing README.md differs.

But a tree has no history. Nothing says T2 came after T1, and nothing records who did it or why. So: a commit — a tiny record holding the hash of one tree, the hash(es) of its parent commit(s), an author, a committer, a timestamp, and a message. (A fourth type, tag, is an annotated pointer to a commit; it’s not load-bearing for the model.) The commit is hashed too, which means its name depends on its tree and its parent, recursively, all the way back.

That last part is where the properties fall out:

”But surely it can’t store full copies forever”

Right — and that’s the next failure. Content-addressing kills duplicate copies of unchanged files, but not near-duplicates. Edit one line of a 2 MB file a hundred times and you get a hundred whole 2 MB blobs, because a one-byte change produces a completely different hash.

It doesn’t actually cost that. The object types above describe the logical model — what Git tells the rest of itself it has. Underneath, Git has a second layer called packfiles. When a repo grows, git gc rolls many loose objects into a packfile and there it does delta-compress similar blobs against each other to save space.

The crucial detail: the deltas in a packfile are an internal storage trick, not the model. They aren’t tied to commit history — Git picks whichever pair of similar blobs compresses best, regardless of which commits they belong to. The logical layer is still snapshots-by-hash; the deltas are just zip-like compression underneath.

So the slogan “Git stores snapshots, not diffs” is true at the layer that matters for reasoning about Git, even though at the bytes-on-disk layer Git absolutely uses deltas to save space. The trick is that the deltas don’t define identity. The hash of the snapshot does.

Where the model shows its seams

A few places it gets weird:

You started with git commit ≈ snapshot of the tree, addressed by the hash of its contents. What did the two-file repo add? — + identity comes from content, so sharing is automatic and history is a chain of hashes. That single substitution is what buys instant checkout, free branches, and tamper-evidence at the same time: they aren’t three features Git implemented, they’re three consequences of naming things by what they are.

Going deeper