Why syscalls are expensive
A function call costs a few cycles. A system call costs hundreds — sometimes thousands. The gap isn't sloppy engineering; it's the price of the user/kernel boundary.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: just call the kernel’s function
- Attempt 2: a hardware door instead of a call
- Attempt 3: distrust everything on entry
- Why it’s still expensive after you’ve counted the cycles
- Show the seams
- Check yourself
- Famous related terms
- Going deeper
The picture version
The whole idea in six pictures, for a reader who has never thought about what sits between their program and the disk. The prose below fills in the seams.
1 · The problem
Same work. Same disk. Wildly different time.
2 · Why it isn’t a function call
There is a wall, and it has exactly one door.
3 · What the door actually does
The hardware flips a bit. Everything else is distrust.
4 · The half you can’t see
You come back to a machine that forgot you.
5 · The only lever
Nobody made the door cheaper. They knock less.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’ve written the unbuffered loop at least once: read a byte, process it, read the next byte, all the way through a 1 GB log file. It takes what feels like forever. Then you wrap the file handle in a buffered reader — same disk, same file, same parsing logic — and it finishes in a fraction of the time. Nothing about the work changed. All you changed was how many times you asked the kernel for something. That’s this post: what exactly you were paying for a million times, and why it costs what it costs.
We’ll keep that loop as the running example. In it you write
read(fd, buf, 1) and it looks like any other function call. It isn’t. A
normal function call is a handful of cycles — push some registers, jump,
pop, return. A syscall
is, on a modern x86-64 box, on the order of hundreds of nanoseconds, and
often more once you count the indirect costs that hit after the call
returns. One survey that timed a minimal syscall across fifteen different
x86-64 hosts found the mode switch alone ranging from under 100 ns on the
fastest machine to several hundred on the slowest, the spread tracking CPU
generation and which speculative-execution mitigations were enabled. Against
a function call’s handful of cycles, that is two to three orders of
magnitude — and it isn’t because kernel programmers are bad at their jobs.
The analogy that holds: a function call is asking the person next to you to pass the salt; a syscall is asking airport security to fetch something from your checked bag. The ID check, the escort, the re-locking of the door cost far more than the fetch. Where the analogy breaks: airport security is slow because humans are slow, whereas a syscall is slow by construction — a faster CPU doesn’t remove the work, because the work is the checking.
The gap exists because a syscall is not really a function call at all. It’s a controlled crossing of a hardware-enforced security boundary. The CPU is in one of two modes — user mode or kernel mode — and a huge amount of machinery exists to stop user code from forging its way into kernel mode. A syscall is the one sanctioned door, and going through it costs what the door costs.
The pain points the boundary is solving:
- Untrusted code can’t be allowed to touch hardware (disks, network cards, page tables, other processes’ memory) directly. One buggy program would take down the whole machine.
- The kernel runs with full privileges — it can write any RAM, talk to any device. You absolutely do not want user code to trick the CPU into running user instructions in that mode.
- The transition has to be safe both directions. Going in, the kernel must not trust a single register, pointer, or flag the user set up. Coming out, the kernel must not leak its own state.
All of that costs cycles. The “expense” is the integrity tax.
Why it matters now
It shows up the moment your workload is bottlenecked on lots of small I/O:
- A web server doing one
readand onewriteper byte of payload will saturate on syscalls long before the network or disk. - A database doing a
preadper 4 KB row is paying syscall overhead on every page. - An AI inference server streaming tokens out over HTTP
is doing a
writeper token chunk, which is fine until you have thousands of concurrent streams. - The reason
io_uringexists at all is “syscalls per I/O operation is too expensive for modern NVMe and 100 GbE; let’s submit and reap many at once.”
The AI-era version: when people batch tokenization, batch GPU launches, or use shared-memory IPC instead of sockets, “amortize the syscall” is usually in the unstated reasons. It’s also why the cost of a single syscall is a useful unit of “how much work is this small operation actually worth doing.”
The short answer
syscall cost = flip the CPU into kernel mode + save and restore all your registers + swap page tables + everything the CPU had warmed up goes cold
Picture to keep: every syscall walks your CPU through a security airlock and back. The airlock is quick; the mess is that everything your code had warmed up — cached lines, cached address translations, trained branch predictors — is cold on the other side, and you come back into a room someone else has rearranged.
A function call stays in user mode and shares everything with the caller. A syscall flips the CPU into kernel mode, saves a pile of state, switches the process’s address-space mapping (on kernels with the post-Meltdown isolation turned on), and disturbs the CPU caches the user code was depending on. You pay all of that twice — once going in, once coming out. The next section names each of those four terms properly.
How it works
Ask the question backwards: what would it take to make read(fd, buf, 1)
as cheap as a function call, and what breaks each time you try?
Attempt 1: just call the kernel’s function
Link the kernel’s read into your process and call it. Zero overhead.
Why it breaks: the kernel’s code runs with the authority to write any physical page and command any device. If your process can jump into it, your process has that authority — a buggy loop scribbles over another process’s memory, and a malicious one owns the machine. The privilege separation isn’t a policy layered on top; it’s the reason an operating system can host more than one program.
Attempt 2: a hardware door instead of a call
So make the CPU enforce it. On x86-64 the user-space syscall path is the
syscall instruction, which does roughly this in microcode:
- Save the user
RIP(instruction pointer) andRFLAGSinto specific registers. - Load the kernel entry point from a CPU model-specific register
(
MSR_LSTAR). - Switch the CPU’s CPL from ring 3 (user) to ring 0 (kernel).
- Mask interrupts according to
MSR_SFMASK.
That much is cheap-ish — tens of cycles. If the story ended here, your byte-at-a-time loop would be fine.
Why it breaks: the hardware changed the privilege level and nothing else. The CPU is now running kernel code on the user’s stack, with the user’s registers, holding the user’s pointers — all of which the user controls and none of which the kernel may believe.
Attempt 3: distrust everything on entry
So the entry stub does the work the hardware didn’t:
- Swaps stacks. The user stack pointer can’t be trusted (could be
garbage, could point at kernel memory); the kernel switches to a per-CPU
kernel stack via
swapgsand a load from a per-CPU area. - Saves user registers. All general-purpose registers go onto the kernel stack so the syscall handler can use them, and so the user state can be restored exactly.
- Switches the page table — on kernels with KPTI enabled, which is the default on affected CPUs. User mode and kernel mode now have different page tables, so Meltdown-style attacks can’t speculatively read kernel memory from user mode. Every syscall writes CR3 on the way in and again on the way out; the kernel’s own documentation puts a CR3 write at “on the order of a hundred cycles,” required at every entry and exit. What that write costs after depends on the CPU: without hardware PCID support, each CR3 write flushes the entire TLB, so every syscall throws away the machine’s address translations. With PCID — present on the parts you are likely running — the CPU tags entries by address space and skips the wholesale flush. It’s the single biggest reason “how expensive is a syscall” has no one answer.
- Validates user pointers. Any pointer the user passed in (
bufinread(fd, buf, n)) has to be checked: is it actually in the user’s address space? Is it readable/writable? If the user passed a kernel address, the syscall has to refuse rather than let the kernel happily dereference it. - Does the actual work.
- Reverses everything on return — restore registers, swap CR3 back,
sysretqback to user mode.
Why it’s still expensive after you’ve counted the cycles
Add all that up and you have a number you could put in a microbenchmark. That number is still an undercount, because the crossing leaves wreckage behind it — microarchitectural collateral damage your user code pays for after the call has already returned:
- TLB flushes from CR3 writes mean the next several user-mode memory accesses miss the TLB and pay full page-walk cost. Your hot loop’s carefully-warmed translations are gone.
- L1 / L2 cache pollution. The kernel ran code and touched data; that evicted some of yours. When you come back to user mode, the first accesses miss caches that were hot a microsecond ago.
- Branch predictor and indirect-branch state get disturbed, mostly as a side effect of running other code. On some CPUs and configurations the kernel also flushes or constrains that state deliberately — the IBPB / IBRS family — though which barriers fire on which boundary is a per-CPU, per-mitigation-setting question, not a flat per-syscall tax. Either way, predictors that had learned your loops have to relearn.
- Speculative-execution mitigations add real overhead per crossing on affected CPUs, and they hit hardest exactly where there’s least real work to hide them: the syscalls that do almost nothing in the kernel pay the largest proportional penalty. Brendan Gregg’s measurements on production cloud workloads when KPTI landed in 2018 put the cost at roughly 2% at 50,000 syscalls per second per CPU, climbing with the syscall rate, and about 5% on a syscall-heavy database benchmark. Note what that pair is really saying: the tax is charged per crossing, so it only becomes visible when you cross a lot.
So a benchmarked syscall time is really “the direct crossing, plus a tail of cold-cache pain that you’ll feel as your user code runs slowly for a little while afterwards.” Only the first part shows up in a benchmark that times the syscall in isolation, which is a large part of why syscall cost is consistently underestimated — and why the honest form of the question is never “how many nanoseconds” but “how much did my program slow down.”
Show the seams
- The numbers are not stable, and that is a fact about the world, not a gap in the reporting. Syscall cost depends on CPU generation, on whether the part has PCID, on which mitigations the kernel enabled at boot, and on the microcode revision underneath all of it — and the last two change on a schedule nobody publishes in advance. Across the fifteen-host survey above, the same minimal syscall differed by roughly an order of magnitude between the fastest and slowest machine. There is no citable “the” number; the only figure that means anything is one you measured on the box you are actually running on.
- Not every syscall is the same price.
getpidis nearly pure overhead and is the canonical “cost of the boundary itself” microbenchmark.readfrom a hot page-cache page is overhead-dominated.readthat hits disk is dominated by the disk, not the syscall. Don’t optimize a syscall that’s already dwarfed by what it’s calling. - vDSO
cheats. A few “syscalls” —
gettimeofday,clock_gettime,getcpu,time— usually don’t cross the boundary at all. The kernel maps a small piece of its own code and data into every process, and that code answers from kernel-maintained pages in user mode. It’s a fast path, not a guarantee: which symbols exist is architecture-specific, and a request the fast path can’t serve — an unsupported clock, say — quietly falls back to a real syscall. That’s why high-rate timing code usually feels free. io_uringdoesn’t make syscalls faster — it makes there be fewer of them. Submission and completion go through shared-memory ring buffers; one syscall can submit hundreds of I/Os. The boundary cost hasn’t gone away, it’s been amortized.- Neither containers nor VMs double the crossing — but for different
reasons. A container is a process with fancier namespaces; its syscalls
go straight to the host kernel like any other. A VM has its own kernel,
so a guest syscall is handled inside the guest and never reaches the
hypervisor either. The virtualization tax is a different event, the VM
exit: the guest traps out to the hypervisor for things the hardware
won’t let it do itself — certain privileged instructions, some interrupts,
device I/O that has to be emulated. Same shape as a syscall, one level
up, and triggered by device and interrupt work rather than by your
read.
You started with syscall cost = flip the CPU into kernel mode + save and restore all your registers + swap page tables + everything the CPU had warmed up goes cold. Which of those terms explains why the buffered version of your
loop won? — none of them
individually; what won was × the number of crossings. Every item in that
line is a fixed toll, so the only lever most programs have is crossing
less. That’s the whole design logic behind buffered I/O, io_uring, and
the vDSO: nobody made the door cheaper, they made you walk through it fewer
times.
The deeper idea is the same one behind virtual memory: the OS has no power over user code except at moments when the hardware hands control over. The syscall is one of those moments, deliberately constructed. Everything expensive about it is the cost of making that handover safe.
Check yourself
Before you go — you profile a service and find it spends 30% of its time in
syscalls. A colleague proposes switching to io_uring. When would that
help a lot, and when would it barely move the needle?
Answer
It helps when the 30% is many small operations — thousands of tiny reads
or writes per second, where the per-crossing toll dominates the work being
requested. io_uring doesn’t make a crossing cheaper; it lets one crossing
carry many operations, so the toll gets divided. It barely helps when the
time is a few large, slow calls — a read that actually waits on a disk or
a network, where the syscall overhead is a rounding error next to the I/O
itself. The diagnostic question isn’t “how much time in syscalls” but “how
many syscalls, and how much real work per syscall.”
And a prediction: clock_gettime is documented as a syscall, yet a loop
calling it millions of times a second doesn’t show the cost this post
describes. Why not?
Answer
Because on Linux it usually isn’t a syscall. The vDSO maps a small piece of
kernel-maintained code and data into every process’s address space, so
clock_gettime can read the current time in user mode without ever
crossing the boundary — for the clocks and architectures that fast path
supports; anything it can’t serve falls back to a real syscall and pays full
price. It’s the exception that proves the rule: the cost
was never in “getting the answer from the kernel,” it was in the crossing —
so when the kernel can safely publish the answer into your address space in
advance, the cost disappears entirely.
Famous related terms
- vDSO —
vDSO = kernel code page mapped into user space + lets a few "syscalls" run without crossing— whyclock_gettimeis essentially free. io_uring—io_uring = shared-memory submission + completion rings + one syscall amortized over many I/Os— Linux’s answer to syscall-per-IO.- KPTI —
KPTI = separate page tables for user and kernel + swap on every crossing— the post-Meltdown tax on syscall cost. - Context switch —
context switch ≈ syscall + scheduler picks a different process to resume— strictly more expensive than a syscall, for the same reasons plus a fresh address space. - VM exit —
VM exit = guest traps out to the hypervisor + hypervisor handles it and resumes the guest— same shape as a syscall one level up, but triggered by privileged instructions and emulated devices, not by the guest’s own syscalls. - eBPF —
eBPF = verified bytecode the kernel runs in-kernel on your behalf— partly motivated by “if I could just run my filter inside the kernel, I wouldn’t pay the boundary cost per packet.”
Going deeper
- The Linux
arch/x86/entry/entry_64.Ssyscall entry stub — the primary source, and the fastest way to answer “what is actually being done on my behalf between thesyscallinstruction and my handler?” Every line is paying for something on the list above. - The kernel’s Page Table Isolation documentation — the primary source for the mitigation half of the bill: what CR3 costs, and exactly when PCID does and doesn’t save you a full TLB flush.
- Brendan Gregg’s writeups on KPTI / Meltdown overhead — the best answer to “how much did the mitigations really cost,” measured on production workloads rather than microbenchmarks.
- LWN’s coverage of
io_uring(2019 onward) — the rabbit hole, for how the “stop crossing so often” argument turns into an actual API.