On the Costs of Syscalls (2021)
gms.tf
gms.tf
According to some benchmarks I've seen (sorry no link), a system call in the middle of a memory heavy inner loop can impact performance for 100 us before performance is back to steady state. This is 100x longer than the system call alone. Of course there is work being done, but at a reduced throughput. This is in line with my practical experience from working with performance sensitive code.
These things are difficult to benchmark reliably and draw actionable conclusions from the results.
This is not a criticism to the author of this article, the article clearly describes the methodology of the benchmarks and does not suggest any wrong conclusions from the data.
Table 1 on page 3 is absolute gold, it quantifies the indirect costs by listing the number of cache lines and TLB entries evicted. The numbers are much larger than I remembered.
According to the table, the simplest syscall tested (stat) will evict 32 icache lines (L1), a few hundred dcache lines (L1), hundreds of L2 lines and thousands of L3 lines, and about twenty TLB entries.
After returning from said syscalls, you'll pay a cache miss for every line evicted.
Also worth noting that inside the syscall, the instructions per clock (IPC) is less than 0.5. When the CPU is happy, you generally see IPC figures around 2 to 3.
Anyway, a 20-line example of a program written against said interpreter is https://github.com/c-blake/batch/blob/1201eefc92da9121405b79... but that only needs the wdcpy fake syscall not the conditional jump forward (although that could/should be added if the open can succeed but the mmap can fail and you want the close clean-up also included in the batch, etc., etc.).
I believe Cassyopia (also mentioned in Soares) hoped to be able to analyze code in user-space with compiler techniques to automagically generate such programs, but I don't know that the work ever got beyond a HotOS paper (i.e. the kinda hopes & dreams stage) and it was never clear how fancy the anticipated batches being. The Xen/VMware multi-calls Soares2010 also mentions do not seem to have inline copy/jumps, though I'd be pretty surprised if that little kernel module is the only example of it.
Do you some intuition of what causes this? I don't have how much experience with this kind of work, but 100µs is enough time to do hundreds of random memory accesses. How can a single syscall do so much damage to the cache?
https://en.wikipedia.org/wiki/Translation_lookaside_buffer
(Also, worth pointing out that they didn't claim a delay of 100µs, just that some (presumably, much smaller, on the order of ns?) delays can show up up to 100µs later before "steady state" is fully restored.)
Isn't that exactly what changed for the Meltdown/Spectre mitigations? That it is invalidated now?
But I don't claim any specific knowledge of on these mitigations.
It takes a while of normal operation until these are populated again with the hot data, until then the system overall throughput is reduced.
I've been on a code-deleting rampage lately, killing off many minor features that add a lot of complexity for seemingly little gain. Glad I don't have anyone interrogating me like this!
> in order to quickly fix this pressing matter, I've attached code that I believe will fix your issue. It uses a custom caching mechanism and fits well into the systemd ecosystem.
pid_t systemd_getpid(void) {
return 0;
}But dick answers is the specialty of the systemd development team. I wouldn't want to get into a contest against them.
"I ran the benchmark on a heterogeneous set of hosts, i.e. different kernels, operating systems and configurations"
In the fine print: the "different OS:es" were two Linux Distributions which are (if I understand correctly) not even that far apart tech wise? (RHEL and Fedora).
I would like to see this extended to other processor families.
I had two loops with a large number of operations. One called a trivial Win32 function (GetACP), the other made a System Call (NtClose) as well as the same trivial Win32 function. I used QueryPerformanceCounter to time the loops.
From there, I estimated the number of machine cycles per iteration (using my clock speed of the time), and subtracted the "trivial" version which made no system calls.
We use getpid() in logging and maybe it's unnecessary.
It's a fantastic write up on not just how system calls work, but the motivation behind why operating systems even implement them in the first place. It's not specifically about Linux, but the book is clearly heavily inspired by early *nix designs and it's still applicable.
You can read it for free here: https://pages.cs.wisc.edu/~remzi/OSTEP/cpu-mechanisms.pdf
https://www.matheusmoreira.com/articles/linux-system-calls
LWN has the kernel's perspective well covered by their articles on the anatomy of Linux system calls:
https://lwn.net/Articles/604287/
https://lwn.net/Articles/604515/
https://lwn.net/Articles/604406/
They are thoroughly dissected in these articles. I also use them as a reference.
> One big reason for this is that some system calls have additional code that runs in glibc before or after the system call runs.
That's not really the fault of the system calls. The problem is the C library itself. It's too stateful, it's full of global data and action at a distance. If you use certain system calls, you invalidate its countless assumptions and invariants, essentially pulling the rug from under it.
I've found programming in freestanding C with Linux system calls and no libc to be a very rewarding experience. Getting rid of libc and its legacy vastly improves C as a language. The Linux system call interface is quite clean. Don't even have to deal with the nonsense that is the errno variable.
Yes, everybody knows you can get from A to B on foot. But for multiple reasons, and over and extended period of time, we developed other methods of doing so. That doesn't invalidate the experience of doing it on foot, but it does make the walk into a very conscious choice that probably isn't the one most people are going to make most of the time.
libc wraps system calls. If you use libc and try to make your own system calls, you're going to collide with internal details of libc.
Not using libc is fine. Not making your own system calls via asm is fine. Pick one.
It's not! I wrote somewhat at length about that very question in this article:
It's the policy of (some/most/all) libc implementations. Don't like it? Find a different libc or do without libc.
> I have altered the ABI. Pray I do not alter it any further.
Linux is actually the odd one. It's the only one with a stable language-agnostic system call ABI. It's stable because breaking the ABI makes Linus Torvalds send out extremely angry emails to the people involved until they fix it. Because of this stability, you actually can have alternative libc implementations like musl and you can also rewrite literally everything in Rust or Lisp if you want. Only on Linux can you do this.
For some reason people try extra hard to make it look like glibc is some integral part of the kernel. Diagrams on Wikipedia showing glibc enveloping the kernel like it's some outer component. Look up Linux manuals and you somehow get glibc manuals instead. The truth is they're completely independent projects. The GNU developers have zero power to force anyone to use glibc on Linux. You can make a freestanding application using system calls directly and literally boot Linux into it. I have made it my goal to create an entire programming language and ecosystem centered around that exact concept.
This is demonstrated by the rest of your comment, in which you note that linux has a stable, language agnostic system call ABI, and that other libc implementations exist.
TFA was ONLY about Linux syscalls, so what other platforms do or do not is irrelevant in context.
Do they apply to all possible uses of syscalls? They do not.
Does OpenBSD limiting how you can call a syscall help at on performance (or hurt)?
Fun fact: this is already an issue for the vdso.
gettimeofday(2) can use the CPU builtin cycle counters, but it needs information from the kernel to convert cycles to an actual timestamp (start and scale). This information too can change during the runtime.
To not have userspace trampled over by the kernel, the vdso contains an open-coded spinlock that is used when accessing this information.
I learnt about this while debugging a fun issue with a real-time co-kernel where the userspace thread occasionally ended up deadlocking inside gettimeofday and triggering the watchdog :-)
Unrelated, but what does "open-coded" mean? I never seem to find an obvious answer online.
In my understanding, open-coded means something akin to "manually inlined". Or: written inline while an acceptable alternative exists as a function.
I would appreciate if someone explained why it’s not a vdso rather than just repeating it’s not necessary like that’s a sufficient explanation.
Also given https://ipfs.io/ipfs/QmdA5WkDNALetBn4iFeSepHjdLGJdxPBwZyY47i... while Linus may have changed tack in the decades since I would not expect much support from stating getpid is performance critical to you.
For sure, right up there with naming things and off-by-1s.
The only time your PID would change from a previous value obtained from that very same function is in the child, after a call to fork(), which is itself so heavy, that it would dwarf your calls to getpid
This sounds so mad it needs more context. You don't need to detect when a process is forked, it's an action issued from inside the process?
Also, what sort of raving lunatic had a busy loop checking the current process PID? Who allowed such a person access to a compiler? That seems like the bigger problem.
Somehow I doubt the code for getpid as a vdso would meaningfully impact the amount of memory used on a running system nor would I expect any meaningful impact to the number of page table entries. Do you have any supporting evidence for your claim otherwise?
Why aren't the pages shared like other dynamic libraries?
We could use the rseq area for PID and TID.
Having said that I agree that it's probably not worth it to put getpid() into vDSO.
If we're going to add all sorts of dross to the vdso, like getpid(), who knows what else will end up there?
Or clone3, but with PID namespaces the answer to "what is the PID of a process" is a trickier question.
getpid was a fairly common way to check for fork prior to mechanisms like madvise WIPEONFORK (Linux 4.14, released in 2017). systemd existed prior to 2017 and likely needed to support older kernels past 4.14's initial release. And the getpid cache was something glibc had prior to April 2017, likely to support this kind of use case. It was even documented behavior[1].
So, anyway, what systemd was doing wasn't totally stupid. It became a lot more expensive when glibc broke it on them and then it needed to be improved. But before that it wasn't as objectionable as people in this thread are suggesting.
(Also: pthread_atfork requires linking pthreads, which is or at least historically was seen as a significant burden on applications which do not use pthreads.)
L2:bat/examples$ BATCH_EMUL=1 ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 3.815 *$2 + 695.6826 *$3
bootstrap-stderr-corr matrix
0.8872 -0.7856
0.04174
L2:bat/examples$ BATCH_EMUL=1 ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 4.528 *$2 + 695.3265 *$3
bootstrap-stderr-corr matrix
1.174 -0.7640
0.05302
L2:bat/examples$ a (695.3265 +- 0.05302)-(695.6826 +- 0.04174)
-0.356 +- 0.067 # i.e. run-to-run consistent
L2:bat/examples$ ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 779.67 *$2 + 48.861 *$3
bootstrap-stderr-corr matrix
1.550 -0.7495
0.2023
L2:bat/examples$ ./closes|awk '{print $3,$2,$1*$2}'|fitl -s0 -c,=,n,b -b100 1 2 3
$1= 787.77 *$2 + 48.745 *$3
bootstrap-stderr-corr matrix
1.350 -0.6207
0.1635
L2:bat/examples$ a (48.861 +- 0.2023)-(48.745 +- 0.1635)
0.12 +- 0.26 # i.e. run-to-run consistent
L2:bat/examples$ a (695.3265 +- 0.05302)/(48.745 +- 0.1635)
14.265 +- 0.048 # i.e. batching makes it 14X faster!
One can go further to get a mean +- std.err of ~695.52 +- ~0.04 nanosec hot-cache time per syscall overhead or even do point-weighting there, but it's probably more important to check fit residuals for serial autocorrelation (very do-able with https://github.com/c-blake/fitl, EDIT: oh, and if you care `a` is basically shown in https://github.com/SciNim/Measuremancer/pull/12 which uses standard error propagation).The bigger point would not be ever more careful measurement of the cost if you do not, but rather the ease of "batchifying" if you do of any highly regular call interface similar to syscalls with high latency and also the "follow on" API design, granularity-wise. The inner code of a mini-assembly language letting you implement, say, open/fstat/mmap/close all in one syscall crossing is only about 30 lines of C. { Yes, yes..It could use multi-CPU-arch and syscall auditing integration.. and sure, ebpf & io_uring alter the Linux landscape these days }. Also, @exDM69 is very correct that pure hot cache-hot loop numbers are only one part of the cost story (https://news.ycombinator.com/item?id=39188551).