Why Git stores snapshots, not diffs
Git's reputation says 'version control = diffs.' Git's actual model says 'version control = snapshots, hashed.' That swap is the whole reason Git feels different.
On this page
The picture version
Five pictures for a reader who has used Git without ever wondering what it keeps. The prose below fills in the seams the pictures skip.
1 · The problem
Check out a commit from 2016. It happens instantly. It shouldn’t.
2 · What is actually stored
Not the changes. The whole contents of every file, every time.
3 · Why that isn’t as wasteful as it sounds
Identical contents get an identical fingerprint — so they are the same box.
4 · So what is a commit?
A slip of paper listing which boxes were on the shelf that day.
5 · Keep this card
The whole thing on one index card.
Why it exists
Clone a repo with ten years of history, then check out a commit from 2016.
It’s instant. Switch back to main. Instant. Create a branch. Instant — the
progress bar doesn’t even get a chance to appear. Now compare that with the
mental model most people carry in: “a repo is a chain of diffs, and checking
out an old version means undoing every change since.” If that were true,
jumping back four years would mean replaying thousands of patches, and it
would feel like it.
It doesn’t feel like it, because that model is wrong. A Git commit names the complete state of the tree as it existed at that moment — every file, not a patch against the last one. When a file doesn’t change between two commits, Git doesn’t store a diff of zero bytes; it just points the new commit at the same file object the old commit was already pointing at. That’s why old commits are as cheap to reach as new ones: there is nothing to replay, just a pointer to follow.
The diff model isn’t a strawman — it’s how Subversion and CVS genuinely worked, which is where most people picked it up. Git threw it out, and that one swap is most of why Git feels qualitatively different from what came before.
We’ll follow one tiny repo the whole way down: two files, README.md and
src/main.py, and two commits where only the README changes.
Why it matters now
Almost every codebase you’ll touch as a software engineer lives in Git.
The mental model leaks into everything: why git log is fast, why branches are
free, why git rebase can reorder history without corrupting it, why a
detached HEAD is a normal state and not a disaster, why GitHub can show you
any historical version of a file instantly without replaying patches.
It also matters because the snapshot model is what makes Git a content-addressed store — the same idea now showing up in IPFS, container image layers, Nix, and large-model weight stores. Git was an early mainstream example of “address things by what they are, not where they live,” and that idea keeps being rediscovered.
The short answer
git commit ≈ snapshot of the tree, addressed by the SHA-1 hash of its contents
Picture to keep: a warehouse where every box is labelled with a fingerprint of what’s inside it, not with where it came from. Two people who happen to store identical contents get the same label — so they’re the same box. A commit isn’t a copy of the warehouse; it’s a slip of paper listing which boxes were on the shelf that day. (Where the warehouse breaks down: boxes here are never edited or thrown away. Changing a file doesn’t open a box, it creates a new one with a new label — which is why “rewriting history” in Git always means writing new objects, never mutating old ones.)
A commit is a tiny object that points at a tree (a directory listing). The tree points at blobs (file contents) and at sub-trees. Every one of those objects is named by the hash of its own bytes. If two commits contain the same file, they end up pointing at the same blob — automatically, with no deduplication step.
There are no diffs in the storage model. Diffs are something Git computes on
demand when you ask git diff or git log -p.
How it works
Try to invent this yourself, and watch each attempt fail.
Attempt 1: copy the whole project on every commit. Instant checkout,
trivially simple, no patch replay. And obviously unusable: our repo’s
src/main.py doesn’t change in commit 2, but you’ve now got two identical
copies of it on disk. Ten thousand commits later you have ten thousand copies
of every file that was never touched.
Fix 1: name every file by the hash of its own contents. Instead of
“main.py, version 2,” store the raw bytes under a name computed from the
bytes themselves — the
SHA-1
of the content. Git calls that a blob: just bytes, no filename. Now
duplicate content cannot be stored twice, because both copies compute the
same name and land in the same place. There is no deduplication step to run;
dedup is a consequence of the naming scheme. Add the same file to two repos in
two countries and both call its blob the same name — the 40-character hex
string you’re used to, in a SHA-1 repo, or 64 characters in a SHA-256 one.
But now nothing has a name or a path. A pile of content-addressed blobs
isn’t a project — you can’t tell which blob was README.md. So: a tree, a
tiny file listing entries (mode, name, hash), where each hash points at a
blob or at another tree for a sub-directory. And a tree is itself content-
addressed, which pays off immediately. Our repo:
- Commit 1 → tree
T1→{README.md → blobA, src → tree S1}.S1→{main.py → blobX}. - Commit 2 → tree
T2→{README.md → blobB, src → tree S1}. SameS1. SameblobX.
T2 is a full snapshot of the project, and storing it cost one new blob and
one new tree. The src/ subtree was unchanged, so it hashed to the same name,
so it is the same object. The “diff” between the two commits is something Git
derives at read time by walking T1 and T2 and noticing README.md differs.
But a tree has no history. Nothing says T2 came after T1, and nothing
records who did it or why. So: a commit — a tiny record holding the hash of
one tree, the hash(es) of its parent commit(s), an author, a committer, a
timestamp, and a message. (A fourth type, tag, is an annotated pointer to a
commit; it’s not load-bearing for the model.) The commit is hashed too, which
means its name depends on its tree and its parent, recursively, all the way
back.
That last part is where the properties fall out:
- Branches are cheap. A branch is nothing but a stored commit hash — as a
loose ref, a 41-byte file holding 40 hex characters and a newline (64 hex in
a SHA-256 repository, or a single line inside
.git/packed-refsonce Git packs them away). Creating one is free; you’re not copying anything. git logis fast. Walking commit parents is just chasing pointers through a hash-keyed object store; no patch replay.- History is tamper-evident. Change one byte of one old file and every hash from that commit forward changes. Git’s identity is the chain of hashes — this is essentially the same idea blockchains were named for, and it predates them in mainstream tools.
- Rebase works. Re-writing history means producing new commit objects with new hashes; the old ones aren’t mutated, just orphaned until garbage collection.
”But surely it can’t store full copies forever”
Right — and that’s the next failure. Content-addressing kills duplicate copies of unchanged files, but not near-duplicates. Edit one line of a 2 MB file a hundred times and you get a hundred whole 2 MB blobs, because a one-byte change produces a completely different hash.
It doesn’t actually cost that. The object types above describe the logical
model — what Git tells the rest of itself it has. Underneath, Git has a second
layer called packfiles.
When a repo grows, git gc rolls many loose objects into a packfile and
there it does delta-compress similar blobs against each other to save space.
The crucial detail: the deltas in a packfile are an internal storage trick, not the model. They aren’t tied to commit history — Git picks whichever pair of similar blobs compresses best, regardless of which commits they belong to. The logical layer is still snapshots-by-hash; the deltas are just zip-like compression underneath.
So the slogan “Git stores snapshots, not diffs” is true at the layer that matters for reasoning about Git, even though at the bytes-on-disk layer Git absolutely uses deltas to save space. The trick is that the deltas don’t define identity. The hash of the snapshot does.
Where the model shows its seams
A few places it gets weird:
- Large binary files. Snapshots-by-hash is brutal for big binaries that change often: each version is a fully new blob, and packfile deltas don’t compress unrelated binary changes well. This is why Git LFS exists.
- Renames. Git doesn’t track renames as a first-class operation. A rename
is a delete plus an add of identical content;
git log --followandgit diff -Minfer renames after the fact — first pairing up files whose blob hashes match exactly, then scoring the leftovers for similarity, which is how a file that was renamed and edited still gets spotted. This works surprisingly well, and it falls down on heavy edits during a rename. - SHA-1. Git’s identity layer was built on SHA-1, which is no longer considered cryptographically safe. Migration to SHA-256 has been in progress for years; how far that has got in practice isn’t well documented — it is supported, but not the default on most hosting.
You started with git commit ≈ snapshot of the tree, addressed by the hash of its contents. What did the two-file repo add? — + identity comes from content, so sharing is automatic and history is a chain of hashes. That single
substitution is what buys instant checkout, free branches, and tamper-evidence
at the same time: they aren’t three features Git implemented, they’re three
consequences of naming things by what they are.
Famous related terms
- Content-addressed storage —
CAS = blob + hash(blob) as its name— the general pattern Git is one example of. - Hash table —
hash table = array + hash function— the same “use the hash to find the thing” idea, scoped to one process. See hash-table. - Merkle tree —
Merkle tree ≈ tree where each node is the hash of its children's hashes— a Git tree object is a Merkle node; this is what makes whole-repo integrity follow from the top commit hash.
Going deeper
- Pro Git, chapter 10 (“Git Internals”), free on git-scm.com — read this for the first-party answer to “what exactly is in a blob, tree, commit, and packfile,” with the byte layouts this post skipped.
git cat-file -p <hash>andgit cat-file -t <hash>in any real repo — five minutes of poking answers “is this actually true of my repo?” better than any explanation, this one included. Start atgit cat-file -p HEAD.- Merkle trees — the rabbit hole for why hashing a tree of hashes makes whole-repo integrity follow from one number at the top.