Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Write-ahead logging

Why databases write your change to a log before they write it to the actual table — and why crash recovery is impossible without it.

Data intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never lost data to a power cut. The prose below fills in the seams the pictures skip.

1 · The problem

Two changes have to happen together. The power doesn’t care.

your account − $100 their account + $100 power cut, here The first change made it to disk. The second didn’t. $100 has ceased to exist. and the app already told you it was done
A transfer is two changes that only make sense together, and a machine can stop between any two instructions. That $100 is the running example: the question is not how to make the crash impossible, but how to survive it.

2 · Why the obvious approach can’t be rescued

Change the data where it sits, and a crash leaves you unable to tell what happened.

before balance: 500 during: the machine dies mid-write balance: 5?? — half old, half new On restart you have a value that was never correct — and nothing anywhere records what it was, or was meant to be. you can’t undo it and you can’t finish it Making the write faster doesn’t help. The problem is that the old value is gone.
Overwriting in place destroys the only copy of the previous state, so an interrupted write leaves data that is neither the old value nor the new one. No amount of care about when you write fixes a design that has nothing to fall back on.

3 · The move

Write down what you’re about to do. Then do it.

1 · the notebook acct A: 500 → 400 acct B: 200 → 300 — committed — only then 2 · the real data updated at leisure — it no longer has to be urgent The notebook entry is what makes the promise, not the data page. Once it is safely written, the app can honestly say done.
The intended change is appended to a log and forced to durable storage before the real data is touched. Durability moves from the data to the log, which is what lets the database answer quickly and tidy up afterwards.

4 · What a crash looks like now

Read the notebook from the last known-good point and redo the rest.

already safely on disk committed, but not applied yet half-written replay these and drop this one — it never said “committed” Redoing a change that was already applied is harmless: the result is the same. So recovery doesn’t need to know how far it got. It just replays. a checkpoint marks how far back the replay has to start, so the notebook isn’t read from page one
Restart reads forward from the last checkpoint, reapplying anything marked committed and discarding anything that isn’t. Because reapplying a change twice lands in the same place, recovery doesn’t have to know what the crash interrupted.

5 · Why appending is safe when overwriting isn’t

A wall can be half-built. A notebook can’t be half-written.

these were finished and never touched again only the last one can be torn Old entries are never revisited, so a crash cannot damage them. A torn final entry is detectable, and simply ignored. it also means the log is written straight along the disk, which is the fastest way to write anything
Because the log only ever grows at the end, everything already written is beyond the reach of the next crash, and only the final entry can be incomplete — which a checksum can spot. The same property makes it fast: appending is sequential, and sequential is what storage is best at.

6 · Keep this card

The whole thing on one index card.

the rule = write down what you will do, durably + only then change the real data ∴ after a crash, replay the record The crash was never made impossible. It was made survivable. and the same log turns out to be useful for far more than recovery — replicas and backups read it too
Picture to keep: a builder who writes every instruction in a bound notebook before touching a brick — if they’re hit by a bus mid-wall, the next builder reads the notebook from the last known-good point and redoes whatever the wall doesn’t already show. The wall can be half-finished; the notebook can’t be half-written, because pages only ever get added to the end.

Why it exists

You tap “send” on a $100 bank transfer and the app says done. Somewhere in that half-second the bank’s database had to take $100 out of one account and put it into another — and if the power had cut in between, the obvious outcome is that the money simply stops existing. It doesn’t happen. That transfer is the running example for this post: two changes that must either both survive a power cut or both vanish, on hardware that offers no such guarantee.

Because the hardware really doesn’t. The naive fix is “write both updates at once,” but disks don’t work like that. A transaction routinely touches dozens of pages across multiple tables and indexes, and there is no primitive for “update twelve pages, all-or-nothing.” Even a single 8 KB page write on spinning rust or an SSD can tear if a crash interrupts it mid-write. So databases have to manufacture atomicity in software, and they all reach for the same trick: write what you’re about to do to a small append-only file first, fsync it, and only then start touching the real data pages. That file is the WAL.

This idea is older than Postgres. It’s the central recovery technique in the ARIES paper (Mohan et al., IBM, 1992), and pretty much every serious relational database — Postgres, Oracle, SQL Server, MySQL/InnoDB, SQLite — runs some descendant of it. Postgres is just where curious engineers most often run into it by name, because pg_wal/ keeps filling up their disk.

Why it matters now

Three things you almost certainly use depend on WAL being a real, durable, ordered stream:

If you’ve ever wondered why synchronous_commit = off makes Postgres faster but scarier, or why a busy database needs WAL on a fast disk even if the dataset fits in RAM — this is the layer you’re poking at.

The short answer

WAL = append-only log of intended page changes + fsync before the data page is allowed to move

Picture to keep: a builder who writes every instruction in a bound notebook before touching a brick. If they’re hit by a bus mid-wall, the next builder reads the notebook from the last known-good point and redoes whatever the wall doesn’t already show. The wall can be half-finished; the notebook can’t be half-written, because pages only ever get added to the end.

The rule, called the WAL rule, is one sentence: log first, data after. As long as the log entry describing a change is on stable storage before the modified data page is, the database can always reconstruct the truth after a crash. Either the WAL contains the transaction’s commit record — recovery redoes the change — or it doesn’t, and the transaction is treated as if it never happened. Atomicity falls out of that ordering. For the transfer: either the debit and the credit both make it (their commit record landed in the WAL before the crash) or neither effectively did (no commit record, so any half-written rows are skipped on replay).

How it works

Build it from the failure. Follow the $100.

Naive attempt: write the debit page, then the credit page. Why it breaks: the crash lands between them, and the money is gone. Doing them in the other order loses money too — it just invents it instead of destroying it.

Fix: write both, then declare the transaction complete. Why it breaks: there’s no moment at which “both” is atomic. The operating system and the drive are free to reorder, buffer, and partially complete those writes, and a crash can leave a single 8 KB page half-old and half-new.

Fix: stop trying to make the data pages atomic, and make one small write atomic instead. Before touching anything, append a record describing the intent — debit A, credit B — to a single sequential file, force it to durable storage, and only then let the data pages move at their leisure. The commit record’s presence in that file is now the single bit that decides whether the transfer happened. There is exactly one thing that has to survive the crash, and it’s an append to the end of a file.

A simplified life of one UPDATE under that rule:

  1. The change is computed in memory (in the buffer cache). Postgres produces a small WAL record describing the modification: which page, which tuple, what bytes changed.
  2. The WAL record is appended to an in-memory WAL buffer. At commit time (or sooner, if the buffer fills) it gets written to the WAL files in pg_wal/ and then fsync’d. Only after the fsync returns does Postgres tell the client “COMMIT successful.”
  3. The actual data page is now dirty in the buffer cache but has not been written to the table file yet. It can sit there for a while.
  4. Eventually a checkpoint forces all dirty pages to disk and records “as of WAL position X, the data files are caught up.” WAL files older than that are now eligible for recycling — unless archiving, a replica, or a replication slot is still holding them (more on that below).
  5. If the server crashes between steps 2 and 4, recovery starts from the last checkpoint, walks forward through the WAL, and reapplies every record whose effects didn’t make it to the data file. This is called the redo phase.

The shape of that lifecycle, with the fsync barrier in the middle:

sequenceDiagram
    participant C as Client
    participant B as Backend
    participant W as WAL (disk)
    participant D as Data file (disk)
    C->>B: UPDATE ... COMMIT
    Note over B: build WAL record,<br/>dirty page in cache
    B->>W: append record + fsync
    W-->>B: durable
    B-->>C: COMMIT ok
    Note over B,D: page stays dirty; written later
    B->>D: flush dirty page (at next checkpoint)

The commit acknowledgement crosses the wire before the dirty data page ever touches its file — that gap is exactly where WAL earns its keep.

The clever part is that step 2 is fast. WAL is a single sequential file (or small set of segment files) being appended to. Sequential I/O is what disks — even SSDs — are best at. Random updates across a 500 GB table would wreck a disk; one fat sequential append per transaction group does not.

A second clever part: WAL also lets Postgres lie about durability in a controlled way. With synchronous_commit = off, commits return before the WAL fsync completes. You get a much faster system, with the explicit trade that a crash can lose the last few hundred milliseconds of already-acknowledged commits. The database stays consistent — torn transactions never appear — you just lose the tail. That distinction (consistent vs. durable) only makes sense because WAL exists.

I should name a gap: I’m describing the standard ARIES-shaped picture and the public Postgres docs. Postgres has its own departures — notably, it does not use ARIES-style undo records the way DB2 does, because MVCC already keeps old row versions around, so rollback is “mark this transaction aborted, the old version is already there.” If you want the exact byte layout of a Postgres WAL record or the precise rules for full-page writes, the source (src/backend/access/transam/) and the docs are the ground truth, not this post.

You started with WAL = append-only log of intended changes + fsync before the data page moves. What did this post add? — + the commit record as the single deciding bit. The log and the ordering are the mechanism, but the reason the $100 can’t half-exist is narrower than that: recovery doesn’t ask “how much of this transaction made it to disk,” a question with many possible messy answers. It asks whether one small record is present in the log, a question with exactly two. Atomicity isn’t achieved by writing carefully. It’s achieved by reducing “did this happen?” to a single durable fact.

Advanced usage

The WAL rule is one sentence, but the machinery built on top of it piles up fast. A handful of things that matter in practice:

Full-page writes — the torn-page defense. A page write to disk isn’t atomic either. Postgres uses 8 KB pages, but the OS and storage layer write in 4 KB (or smaller) sectors. A crash mid-write can leave half the old page and half the new — a torn page. Replaying a delta-style WAL record against a torn page produces garbage. Postgres’s fix: after each checkpoint, the first time any page is modified, the entire page is written into WAL, not just the change. Recovery replays the full page first, overwriting whatever torn version is on disk, and only then applies the subsequent deltas. This is why WAL volume spikes immediately after a checkpoint, and why turning full_page_writes = off is only safe if your storage independently guarantees atomic page writes.

Group commit — sharing the fsync. An fsync that flushes 1 KB and one that flushes 1 MB cost roughly the same; the latency is dominated by the round trip to the disk’s durability barrier, not the bytes. So if a hundred transactions are trying to commit at the same moment, you can fsync the WAL once and acknowledge all of them. Postgres does this implicitly: committing backends pile into the same WAL flush, so one fsync satisfies many commits. It also exposes commit_delay and commit_siblings for explicit pacing (“wait this many microseconds for friends to show up, provided at least N other backends are mid-transaction”). It’s the cheapest throughput win in a write-heavy workload, and it falls out of WAL being a single shared append point.

WAL archiving and point-in-time recovery. Set archive_mode = on and an archive_command that copies completed WAL segments to a backup location, and Postgres will hand off each segment as it rotates. Take one base backup, keep archiving WAL forever, and you can recover the database to any moment covered by the base backup plus the WAL you still have — PITR. The recovery target can be a timestamp, an XID, or a named restore point. Mechanically it’s the same replay loop crash recovery uses, just fed from an archive instead of pg_wal/ and stopped early on purpose.

Replication slots — pinning the log. WAL segments are eligible for recycling once a checkpoint passes them, unless something is still holding them — archiving, wal_keep_size, a standby, or a slot. If a streaming replica or a logical-decoding consumer is lagging without such protection, the bytes they still need can disappear out from under them. A replication slot tells Postgres “do not recycle WAL past position X until this consumer acknowledges past it.” Useful and dangerous in the same breath: a forgotten slot on a dead consumer is the classic way to fill pg_wal/ until the disk runs out and the database refuses new writes.

Logical decoding — WAL as a change stream. The WAL is binary and tied to physical page layouts, but a logical decoding plugin lets Postgres decode committed table changes into a plugin-defined stream — typically row-level events (“INSERT into orders, values …”) — without anyone re-querying the table. This is what Debezium, pg_recvlogical, and most Postgres CDC tooling sit on. The WAL was designed for recovery; it turned out to also be an ordered, durable, at-least-once change stream (consumers handle the occasional re-delivery after a crash), and the ecosystem has been mining that second job ever since.

Check yourself

Before you go — a team sets synchronous_commit = off to speed up their payments service, reasoning that losing a few hundred milliseconds of writes in a rare crash is acceptable. What exactly have they agreed to, and is it what they think?

Answer

They’ve agreed that transactions the database already told the client had committed may not exist after a crash. Note where the damage lands: the database itself stays perfectly consistent — no half-transfers, no invented money, because the commit record either made it to the log or didn’t. What’s broken is the promise made outside the database. The payment service replied “success,” possibly sent an email, possibly called a partner API, for a transaction that recovery then decides never happened. Consistency and durability are separate guarantees and this setting trades away only the second — which is fine for analytics ingest and genuinely dangerous the moment anyone outside the database acted on the acknowledgement.

And one more — WAL volume spikes right after every checkpoint and then falls off. Given how full-page writes work, would making checkpoints more frequent help or hurt?

Answer

Hurt, for WAL volume — and this catches people, because more frequent checkpoints sound tidier. The first time a page is modified after a checkpoint, its entire 8 KB goes into the WAL rather than just the delta, as protection against torn pages. Halve the checkpoint interval and every hot page gets full-page-written twice as often, so WAL volume grows. What you buy in exchange is a shorter replay on restart, since recovery starts from the last checkpoint. It’s a real trade — recovery time against write volume — and neither direction is free, which is why the tuning parameters exist instead of a default that’s right for everyone.

Going deeper