Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why syscalls are expensive

A function call costs a few cycles. A system call costs hundreds — sometimes thousands. The gap isn't sloppy engineering; it's the price of the user/kernel boundary.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never thought about what sits between their program and the disk. The prose below fills in the seams.

1 · The problem

Same work. Same disk. Wildly different time.

one byte at a time read(fd, buf, 1) 1,073,741,824 trips to the kernel for one 1 GB log file 64 KB at a time read(fd, buf, 65536) 16,384 trips to the kernel same file, same parsing Nothing about the work changed.
All that changed is how many times you asked the kernel for something. Everything below is about what one of those asks costs, and why.

2 · Why it isn’t a function call

There is a wall, and it has exactly one door.

USER MODE — ring 3 your program. May not touch disks, network cards, page tables, or another process’s memory. a wall the hardware enforces with exactly one door syscall in out KERNEL MODE — ring 0 may write any byte of RAM and command any device — which is exactly why user code must never be let in unchecked.
The wall is the reason one machine can safely run more than one program. The door is the only sanctioned way through it, and going through costs what the door costs.

3 · What the door actually does

The hardware flips a bit. Everything else is distrust.

between your read() and the kernel’s handler 1 flip the privilege level — cheap, tens of cycles 2 swap to a kernel stack — the user’s can’t be trusted 3 save every one of your registers 4 swap page tables, where KPTI is on — roughly a hundred cycles 5 check every pointer you passed in — then do the actual work …then reverse all of it on the way back out
Only step 1 is the CPU instruction. Steps 2–5 exist because the kernel may not believe a single register, pointer or flag the user set up.

4 · The half you can’t see

You come back to a machine that forgot you.

before the call cached data lines cached translations trained predictors one syscall after it returns cached data lines cached translations trained predictors This half never shows up in the benchmark. you feel it as your own code running slowly for a while afterwards
Timing the syscall in isolation measures the crossing and misses the wreckage — which is a large part of why syscall cost is consistently underestimated.

5 · The only lever

Nobody made the door cheaper. They knock less.

The toll is fixed. Only the multiplier is yours. cost = toll × crossings buffered I/O ask for 64 KB, not for 1 byte 65,536× fewer knocks io_uring one crossing carries hundreds of I/Os the toll, amortized the vDSO the answer is already in your address space no crossing at all
Three different answers to one question, and none of them speeds up the door. The vDSO is the tell: when the kernel can publish the answer into your address space in advance, the cost vanishes entirely — because the cost was never the answer, it was the crossing.

6 · Keep this card

The whole thing on one index card.

SYSCALL COST = a guarded door you must go through + everything warm goes cold behind you × how many times you knock …and only the last term is under your control
Picture to keep: every syscall walks your CPU through a security airlock and back. The airlock is quick. The mess is that everything your code had warmed up is cold on the other side, and you come back into a room someone else has rearranged.

Why it exists

You’ve written the unbuffered loop at least once: read a byte, process it, read the next byte, all the way through a 1 GB log file. It takes what feels like forever. Then you wrap the file handle in a buffered reader — same disk, same file, same parsing logic — and it finishes in a fraction of the time. Nothing about the work changed. All you changed was how many times you asked the kernel for something. That’s this post: what exactly you were paying for a million times, and why it costs what it costs.

We’ll keep that loop as the running example. In it you write read(fd, buf, 1) and it looks like any other function call. It isn’t. A normal function call is a handful of cycles — push some registers, jump, pop, return. A syscall is, on a modern x86-64 box, on the order of hundreds of nanoseconds, and often more once you count the indirect costs that hit after the call returns. One survey that timed a minimal syscall across fifteen different x86-64 hosts found the mode switch alone ranging from under 100 ns on the fastest machine to several hundred on the slowest, the spread tracking CPU generation and which speculative-execution mitigations were enabled. Against a function call’s handful of cycles, that is two to three orders of magnitude — and it isn’t because kernel programmers are bad at their jobs.

The analogy that holds: a function call is asking the person next to you to pass the salt; a syscall is asking airport security to fetch something from your checked bag. The ID check, the escort, the re-locking of the door cost far more than the fetch. Where the analogy breaks: airport security is slow because humans are slow, whereas a syscall is slow by construction — a faster CPU doesn’t remove the work, because the work is the checking.

The gap exists because a syscall is not really a function call at all. It’s a controlled crossing of a hardware-enforced security boundary. The CPU is in one of two modes — user mode or kernel mode — and a huge amount of machinery exists to stop user code from forging its way into kernel mode. A syscall is the one sanctioned door, and going through it costs what the door costs.

The pain points the boundary is solving:

All of that costs cycles. The “expense” is the integrity tax.

Why it matters now

It shows up the moment your workload is bottlenecked on lots of small I/O:

The AI-era version: when people batch tokenization, batch GPU launches, or use shared-memory IPC instead of sockets, “amortize the syscall” is usually in the unstated reasons. It’s also why the cost of a single syscall is a useful unit of “how much work is this small operation actually worth doing.”

The short answer

syscall cost = flip the CPU into kernel mode + save and restore all your registers + swap page tables + everything the CPU had warmed up goes cold

Picture to keep: every syscall walks your CPU through a security airlock and back. The airlock is quick; the mess is that everything your code had warmed up — cached lines, cached address translations, trained branch predictors — is cold on the other side, and you come back into a room someone else has rearranged.

A function call stays in user mode and shares everything with the caller. A syscall flips the CPU into kernel mode, saves a pile of state, switches the process’s address-space mapping (on kernels with the post-Meltdown isolation turned on), and disturbs the CPU caches the user code was depending on. You pay all of that twice — once going in, once coming out. The next section names each of those four terms properly.

How it works

Ask the question backwards: what would it take to make read(fd, buf, 1) as cheap as a function call, and what breaks each time you try?

Attempt 1: just call the kernel’s function

Link the kernel’s read into your process and call it. Zero overhead.

Why it breaks: the kernel’s code runs with the authority to write any physical page and command any device. If your process can jump into it, your process has that authority — a buggy loop scribbles over another process’s memory, and a malicious one owns the machine. The privilege separation isn’t a policy layered on top; it’s the reason an operating system can host more than one program.

Attempt 2: a hardware door instead of a call

So make the CPU enforce it. On x86-64 the user-space syscall path is the syscall instruction, which does roughly this in microcode:

  1. Save the user RIP (instruction pointer) and RFLAGS into specific registers.
  2. Load the kernel entry point from a CPU model-specific register (MSR_LSTAR).
  3. Switch the CPU’s CPL from ring 3 (user) to ring 0 (kernel).
  4. Mask interrupts according to MSR_SFMASK.

That much is cheap-ish — tens of cycles. If the story ended here, your byte-at-a-time loop would be fine.

Why it breaks: the hardware changed the privilege level and nothing else. The CPU is now running kernel code on the user’s stack, with the user’s registers, holding the user’s pointers — all of which the user controls and none of which the kernel may believe.

Attempt 3: distrust everything on entry

So the entry stub does the work the hardware didn’t:

Why it’s still expensive after you’ve counted the cycles

Add all that up and you have a number you could put in a microbenchmark. That number is still an undercount, because the crossing leaves wreckage behind it — microarchitectural collateral damage your user code pays for after the call has already returned:

So a benchmarked syscall time is really “the direct crossing, plus a tail of cold-cache pain that you’ll feel as your user code runs slowly for a little while afterwards.” Only the first part shows up in a benchmark that times the syscall in isolation, which is a large part of why syscall cost is consistently underestimated — and why the honest form of the question is never “how many nanoseconds” but “how much did my program slow down.”

Show the seams

You started with syscall cost = flip the CPU into kernel mode + save and restore all your registers + swap page tables + everything the CPU had warmed up goes cold. Which of those terms explains why the buffered version of your loop won? — none of them individually; what won was × the number of crossings. Every item in that line is a fixed toll, so the only lever most programs have is crossing less. That’s the whole design logic behind buffered I/O, io_uring, and the vDSO: nobody made the door cheaper, they made you walk through it fewer times.

The deeper idea is the same one behind virtual memory: the OS has no power over user code except at moments when the hardware hands control over. The syscall is one of those moments, deliberately constructed. Everything expensive about it is the cost of making that handover safe.

Check yourself

Before you go — you profile a service and find it spends 30% of its time in syscalls. A colleague proposes switching to io_uring. When would that help a lot, and when would it barely move the needle?

Answer

It helps when the 30% is many small operations — thousands of tiny reads or writes per second, where the per-crossing toll dominates the work being requested. io_uring doesn’t make a crossing cheaper; it lets one crossing carry many operations, so the toll gets divided. It barely helps when the time is a few large, slow calls — a read that actually waits on a disk or a network, where the syscall overhead is a rounding error next to the I/O itself. The diagnostic question isn’t “how much time in syscalls” but “how many syscalls, and how much real work per syscall.”

And a prediction: clock_gettime is documented as a syscall, yet a loop calling it millions of times a second doesn’t show the cost this post describes. Why not?

Answer

Because on Linux it usually isn’t a syscall. The vDSO maps a small piece of kernel-maintained code and data into every process’s address space, so clock_gettime can read the current time in user mode without ever crossing the boundary — for the clocks and architectures that fast path supports; anything it can’t serve falls back to a real syscall and pays full price. It’s the exception that proves the rule: the cost was never in “getting the answer from the kernel,” it was in the crossing — so when the kernel can safely publish the answer into your address space in advance, the cost disappears entirely.

Going deeper