Why monotonic time is different from wall-clock time
Wall-clock time tells you what to put on a calendar. Monotonic time tells you how long something took. Confusing them is how you get bugs that look like physics violations.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: one clock, kept correct
- Attempt 2: just never adjust it, then
- Attempt 3: two clocks from one counter
- Attempt 4: pick which monotonic you meant
- Show the seams
- Check yourself
- Famous related terms
- Going deeper
The picture version
The whole idea in six pictures, for a reader who has never wondered where a computer’s clock comes from. The prose below fills in the seams.
1 · The problem
A request that finished before it started.
2 · What actually happened
The clock moved between your two readings.
3 · Why no clock can fix it
Two questions. One clock can only answer one.
4 · The fix
One counter, two different things done to it.
5 · What it costs you
A clock nobody may correct can't be compared with anyone else's.
6 · Keep this card
The whole thing on one index card.
Why it exists
Someone opens your latency dashboard and there’s a bar hanging below zero. A request took negative twenty-three minutes. Nobody deployed anything. The service was healthy. The graph is just claiming that a thing finished before it started.
The code that produced it is the most natural code in the world:
start = time.time()
handle_request()
elapsed = time.time() - start
We’ll follow that one measurement the whole way down. Most days it works.
That night it didn’t, because the machine’s clock got adjusted underneath
it — by
NTP
nudging the wall clock backwards to correct drift, by a virtual machine
pausing and resuming, by an admin running date -s to fix a wrong clock, or
by a leap second being absorbed.
You probably assume the fix is a better clock — tighter NTP, a nicer time library, more decimal places. It isn’t, and no amount of clock accuracy would have saved that measurement. The problem is that two different questions were asked of one clock, and they need two different clocks:
- What time is it right now, in the human world? That’s wall-clock time: it must agree with calendars, time zones, NTP, and ultimately the rotation of the Earth. To stay agreeing, it has to be correctable — and corrections sometimes mean moving it backwards.
- How long has it been since X? That’s monotonic time: it must never go backwards and never jump. It doesn’t even need to be a meaningful date — it just needs to count forward from some arbitrary fixed point.
Modern operating systems expose both because trying to make one clock do both jobs is a contradiction. You can’t simultaneously promise “this matches civil time” and “this never goes backwards” — civil time itself goes backwards sometimes.
Why it matters now
The specific thing that has changed is where code runs. Ephemeral cloud
instances and GPU boxes spun up for a single job come up with a clock that
hasn’t been disciplined yet, and get yanked into shape by NTP shortly after.
That’s a window, on every freshly booted machine, in which any time.time()
subtraction can produce the negative bar above — and short-lived workloads
spend a disproportionate share of their lives inside it. (Containers don’t
add a second window: a pod reads the host’s CLOCK_REALTIME, so it inherits
whatever correction the host is in the middle of.)
It matters for everything where “elapsed time” is load-bearing:
- Timeouts (HTTP, database, RPC). A negative elapsed time can make a timeout fire instantly or never fire at all.
- Rate limiters and token buckets. If “now” goes backwards, you can grant the same budget twice or refuse legitimate traffic for hours.
- Retries with exponential backoff.
- Distributed traces and benchmarks where you subtract two timestamps to get a duration.
- Cache TTLs measured locally on a node.
The bugs are nasty because they’re rare, system-dependent, and they look like the laws of physics broke.
The short answer
monotonic time = "seconds since some fixed point" + "never goes backwards, never jumps"
Picture to keep: a wall calendar and a stopwatch, bolted to the same wall. Someone is allowed to walk up and correct the calendar whenever it disagrees with the rest of the world. Nobody is allowed to touch the stopwatch while it’s running — and it doesn’t know what day it is.
Wall-clock time answers “what time is it?”. Monotonic time answers “how much time has passed?”. They are different problems, so the OS gives you different clocks. Use wall-clock for things humans see (logs, scheduled jobs, “created_at”). Use monotonic for every duration, deadline, and timeout your code computes.
How it works
Attempt 1: one clock, kept correct
This is what produced the negative bar. There is one clock; a daemon keeps it agreeing with the outside world; you subtract two readings to get a duration.
Why it breaks: “kept correct” and “safe to subtract” are incompatible requirements. Correcting a clock that’s ahead means moving it backwards, and the moment it moves backwards between your two readings, your duration is wrong — sometimes negative, sometimes hugely positive if it was moved forwards. The clock did its job. Your subtraction was never entitled to assume otherwise.
Attempt 2: just never adjust it, then
Freeze the clock’s rate and let it free-run. Now subtraction is safe.
Why it breaks: every crystal drifts, so within days the machine
disagrees with the rest of the world about what time it is, and every
timestamp in your logs, every created_at, every cron schedule is wrong.
You fixed durations by breaking calendars.
Attempt 3: two clocks from one counter
Neither requirement can be dropped, so the OS stops trying to satisfy both with one number. Both clocks are usually derived from the same hardware source — on modern x86, typically the TSC. (Reads of either clock usually go through the vDSO, which is how they’re fast — but that’s the access path, not the clock.) The difference between the two clocks is entirely what the OS does with that counter before handing you a number.
For wall-clock time, the kernel keeps an offset from the counter to civil
time, and ntpd/chrony/systemd-timesyncd continuously adjust that
offset. There are two adjustment modes:
- Step — overwrite the clock instantly. Big jumps. This is what
clock_settimeanddate -sdo, and what NTP does on a fresh boot when it’s wildly wrong. - Slew — speed the clock up or slow it down by a tiny percentage so it converges on the correct time over minutes. This avoids backward jumps but means a “second” of wall-clock time is not always a real physical second.
For monotonic time, the kernel drops the civil-time offset entirely: no
steps, no date -s, no leap seconds. (It does not drop everything —
CLOCK_MONOTONIC is still frequency-adjusted by NTP, which is the
distinction the next section turns on.) On Linux that’s
clock_gettime(CLOCK_MONOTONIC, …). Swap
time.time() for time.monotonic() in the snippet above and the negative
bar becomes impossible.
Attempt 4: pick which monotonic you meant
One monotonic clock isn’t quite enough either, because “never goes backwards” leaves two questions open — should it be slewed with NTP, and should it tick while the machine is asleep? Linux answers by shipping several flavours:
CLOCK_MONOTONIC— never jumps, but can be slewed (counts a little slower or faster while NTP corrects).CLOCK_MONOTONIC_RAW— also never jumps, and is not slewed: raw, undisciplined hardware ticks. Useful when you want “seconds as this crystal counts them” rather than “seconds as the kernel currently defines them” — but note the direction of the trade. Because nothing is correcting it, over a long intervalCLOCK_MONOTONIC_RAWcan drift further from real elapsed time than the slewedCLOCK_MONOTONICdoes. Raw means untouched, not more accurate.CLOCK_BOOTTIME—CLOCK_MONOTONIC(slewing and all) but also includes time the machine spent suspended. Most “wall-clock-but-safe” choices end up reaching for this.
In application languages:
- Go’s
time.Now()carries both a wall and a monotonic reading in the same value, andt2.Sub(t1)uses the monotonic part automatically. This is one of the cleaner designs in the wild. - Python has
time.monotonic()andtime.perf_counter()separate fromtime.time(). Use the former two for durations. - Java has
System.nanoTime()(monotonic-ish) andSystem.currentTimeMillis()(wall-clock). ThenanoTime()Javadoc is explicit that it’s the one for measuring elapsed time, and that its value has no meaning as a date. - JavaScript has
performance.now()(monotonic, ms-resolution-ish) andDate.now()(wall-clock).
The pattern across all of them is the same: there are two functions because there are two questions.
Show the seams
A few things the textbook account skips:
- The monotonic epoch is per-machine, and arbitrary. Linux’s
CLOCK_MONOTONICis system-wide — every process on the box reads the same timeline, unless it has been put in its own time namespace — but its zero point means nothing outside that box, so a monotonic timestamp from one machine cannot be compared with one from another. Cross-machine timing has to use wall-clock plus careful sync, or give up on durations and order events causally instead. (Some language representations are narrower still: Go embeds a monotonic reading inside atime.Timevalue, so the safe comparison is between two readings your own process took.) - Suspend/resume is the trap.
CLOCK_MONOTONICon Linux does not advance while the system is suspended — that’s still true today, not a historical quirk. A laptop that sleeps for an hour sees almost no elapsed time onCLOCK_MONOTONICbetween sleep and wake.CLOCK_BOOTTIMEis identical toCLOCK_MONOTONICexcept that it does include suspended time, which is precisely why it exists. Whether your language’smonotonic()ticks during suspend depends on which clock the runtime picked, and the answer is not always documented. - Leap seconds are messy. When a leap second is inserted, civil time has to absorb an extra second somewhere. Different systems handle this differently: some step backwards by one second, some “smear” it across hours so each second is fractionally longer, some pretend it didn’t happen. Monotonic clocks ignore leap seconds entirely, which is another reason to use them for durations.
- Virtual machines lie convincingly. A VM that gets paused by the hypervisor for 200 ms and then resumed will, depending on configuration, see either a 200 ms gap on its monotonic clock or no gap at all. There isn’t a universally correct answer, which is why benchmarks inside VMs need to be read with care.
- Reading the hardware yourself is a different promise. Threads on
different cores executing
rdtsccan see slightly inconsistent values if the cores’ counters aren’t perfectly synchronised, which is why the kernel validates its clocksource at boot and falls back to a slower one when the TSC doesn’t hold up. Going throughclock_gettime(CLOCK_MONOTONIC)buys you that validation; hand-rollingrdtscfor a fast timer does not. The guarantee is attached to the API, not to the counter underneath it.
You started with monotonic time = "seconds since some fixed point" + "never goes backwards, never jumps". What did the negative-latency bar add
to that line? — + the fixed point is arbitrary and local, which is the
price of the guarantee. A clock that refuses to be corrected can’t be
compared with anyone else’s, which is why monotonic time is useless for
logs and mandatory for durations. “What time is it” and “how long did this
take” are not the same question, and the universe doesn’t owe us one clock
that answers both.
Check yourself
Before you go — you fix the dashboard by switching to time.monotonic(),
and the negative bars stop. Then someone asks you to correlate that
latency spike with a spike on a different machine. Can you subtract one
box’s monotonic timestamp from the other’s?
Answer
No. Each machine’s monotonic clock counts from an arbitrary, unrelated zero — often boot — so the difference between two of them is a meaningless number that will nonetheless look plausible. Cross-machine correlation has to use wall-clock timestamps (accepting their sync error, typically milliseconds under decent NTP) or a logical clock that orders events by causality instead of by time. The rule of thumb: monotonic for durations within one machine, wall-clock for anything two machines have to agree on.
And a diagnosis: a laptop app measures a 30-second timeout with
CLOCK_MONOTONIC. The user closes the lid, goes to lunch, and reopens it
an hour later. Has the timeout fired?
Answer
On Linux, no — CLOCK_MONOTONIC does not advance while the system is
suspended, so from the app’s perspective almost no time passed.
That is exactly why CLOCK_BOOTTIME exists: same never-backwards
guarantee, but it includes suspended time. Which one your language’s
monotonic() maps to is a runtime decision and isn’t always documented —
worth checking before you rely on a long timeout surviving a sleep.
Famous related terms
- NTP —
NTP = "ask a time server what time it is" + "slew or step the local clock toward that"— the source of most wall-clock adjustments. - Leap second —
leap second ≈ "an extra second inserted into UTC to keep it aligned with Earth's rotation"— the reason wall-clock time can repeat or skip. - Logical clock / Lamport timestamp —
Lamport clock = counter + "I saw a message from you, bump mine past yours"— what you reach for when neither wall-clock nor monotonic is enough, e.g. ordering events across machines. - TSC —
TSC = CPU register + "ticks every cycle"— the hardware most monotonic clocks ultimately read. - vDSO —
vDSO ≈ "syscalls without the syscall"— whyclock_gettimeis fast enough to call in tight loops.
Going deeper
man 2 clock_gettime— the primary source for exactly what each Linux clock promises, and the only way to settle “does this one tick during suspend?” for a specific flavour.- The “Monotonic Clocks” section of Go’s
timepackage documentation — the clearest worked example of how a language can bolt monotonic semantics onto an existing timestamp type sot2.Sub(t1)is safe by default. - Google’s “leap smear” writeup — the rabbit hole, for why much of the industry chose to spread a leap second across hours rather than let civil time repeat a second.