The Greatest Bug of All
wilshipley.com
wilshipley.com
I had no idea who she was, and now the entire engineering deptartment thinks I'm nuts for her.
But a mapped region certainly can be faster for random access reads from a file. Doing a seek() followed by a read() takes two system calls (and thus four context switches) in addition to any I/O overhead. Faulting the page in from an existing mapping instead takes only one interrupt handler at worst, and potentially completes without kernel intervention at all if the page is already mapped.
Given a particular access pattern, whether the disk gets read or not is independent of the choice to use seek+read or mmap. System calls that don't block don't cause a context switch (1). Therefore, the occurrence of context switches due to disk reads is also independent of the choice to use seek+read or mmap, so trying to avoid context switches isn't a reason to use one or the other.
(1) modulo the timeslice expiring during the system call, which could actually help the system call case since there's no need to take an interrupt from user space when the timeslice expires due to a timer: the system call can be predicted since it's in the code stream, but the interrupt can't be. Anyway, that's all in the noise.)
I'm glad we agree on my second comment, in any case. :)
Anyway, in any common parlance of the term "context switch" I've ever seen, the amount of state switched to take a system call is much less (edit: oops, originally said 'greater' ;) than the amount of state switched in a context switch. At the very least, a context switch should mean the save and restore of user-level register state, which isn't necessary for a system call. You certainly wouldn't want to do that just to make a system call on a typical RISC architecture with 30-something registers! After all, from the perspective of the caller, it's just a special procedure call. The compiler and system call entry point can even use the same argument-passing convention.
I wrote at least part of (maybe most, can't remember anymore :) the "recommended" context switch handler for IA-64. It was significantly more expensive than the system call handler. IA-64 dedicated quite a bit of hardware to making system calls cheap. If I recall correctly, besides a few privileged registers, the only thing we changed was the location of the register stack engine's backing store. That's just a single user-level register. (You don't want to spill registers with potentially privileged state into user space.)
Anyway, this is way off-topic now and I don't think anybody's interested other than the three of us, so that's the last I'll say about it.
This mmap vs. read argument is as old as the hills. I picked a side a long time ago; keep it simple.
I guess I'm stunned that this turns out to have been such a controversial notion...
I'm sorry, this isn't controversial. All I said was, "mmap isn't faster than read in typical use cases". But then we all got really specific talking about ESP and CS and CR3 and now we have something to go back and forth on, which is kind of fun, and I might learn something new. Didn't mean to snipe at you.
Given that main memory reads are pushing 100 cycles on a modern box, all those things add up. A context switch (my usage: meaning a bounce to a non-local, non-current execution environment) is really expensive.
Yes, cache is important. But there's a world of difference between the "flush the caches and the TLB" behavior that system calls used to incur and the "expensive relative to local variable access" behavior we're talking about here.
"A bounce to non-local, non-current execution environment" literally doesn't mean anything. You could be talking about the CR3 change that swaps page hierarchies, the CPL3->0 change that allows privileged instructions, the CS change that got you there, or even the ESP change to "swap stacks". Which of these are expensive?
I'm chasing you down because it's fun, not because I think you don't know what you're talking about.
But, let's make this relevant: if system call overhead is drowned by disk I/O overhead, then one argument for mmap goes away.
I'm sorry, but you seem to have a wildly inflated idea of the speed of system calls on modern OSs. The real world just doesn't work like that. This is my last post on this thread.
Want the punch line? OSX getpid averages under 100 cycles; snprintf (no IO) averages over 1000.
Breakpoint 1, 0x96210c44 in getpid ()
(gdb) disp/i $eip
2: x/i $eip 0x96210c44 <getpid>: call 0x96210c49 <getpid+5>
(gdb) stepi
0x96210c49 in getpid ()
2: x/i $eip 0x96210c49 <getpid+5>: pop %ecx
0x96210c4a in getpid ()
2: x/i $eip 0x96210c4a <getpid+6>: lea 0xa6427a3(%ecx),%ecx
0x96210c50 in getpid ()
2: x/i $eip 0x96210c50 <getpid+12>: mov (%ecx),%eax
0x96210c52 in getpid ()
2: x/i $eip 0x96210c52 <getpid+14>: test %eax,%eax
0x96210c54 in getpid ()
2: x/i $eip 0x96210c54 <getpid+16>: jle 0x96210c57 <getpid+19>
0x96210c57 in getpid () 2: x/i $eip 0x96210c57 <getpid+19>: mov $0x14,%eax
0x96210c5c in getpid ()
2: x/i $eip 0x96210c5c <getpid+24>: call 0x961e8cf4 <_sysenter_trap>
s = (s >> 1) ^ (-(s & 1) & 0xd0000001);
I also timed it the way you asked, with wall clock time over 100M iterations (10M runs too fast to time). Surprising even me, getpid wins by almost a full second.Got another challenge for me to write code for?
Typically kernel pages (and the TLB entries which cover them) are "pinned." You don't want to take a TLB exception in the middle of an interrupt service routine, let alone a page fault. Most architectures couldn't even handle the nested exception, and would panic.
I haven't kept up with x86 on Linux; does Linux enable the global page extensions? I would assume it does for kernel mappings.
Also, as I pointed out below, typically you don't even need to change any TLB entries on kernel entry, because the address space doesn't change and the kernel TLB mappings are usually pinned. In other words, any load from or store to a kernel page will always hit the TLB.
On most architectures, you probably don't even need to swap registers; not the general-purpose ones, anyway.
Sorry, but you're confusing switching processes with switching privilege, and greatly overestimating the cost of a system call.
I agree in general, but... I have run into a few platform bugs in my time. For example, at one point realloc in the standard C library that came with Visual Studio would break if you (IIRC) allocated a bunch of blocks that were over 16k, and realloced them to be smaller than 16k. And I had to figure that out the hard way, because the bug was making our code crash.
Also, on most game consoles, I'd say about a quarter of the time you think something is a platform bug, it really is. Seriously. Compiler bugs and SDK bugs just aren't all that uncommon there.
Remember that the very first thing you do, when looking
at any bug, before you even start thinking about it, and
long before you look at your code, is replicate it. You
can't debug what you can't replicate, and user reports
are usually lacking in some details that your trained eye
will catch.
So if you don't have a repro, you don't investigate the bug? Not a good way to treat your customers.His point is that you always try to reproduce the bug before you assume you understand it.
His comments don't say anything about what to do if you're unable to reproduce. He's just suggesting in which order he thinks you should take certain actions.
It's altogether too easy to jump straight to a patch, submit it, and miss the forest for the trees because you never actually went through the use flow and realized you were patching a symptom, not the underlying problem.
One will do well to learn debugging things from least amount of information - log files, crash dumps etc. Doing this consistently even for the easy bugs with repro will teach developers to put more information into log files and make data structures easier to discover within dumps. Then once hard problems come you will be ready.
Your first paragraph presents a straw man. Sure, in extreme cases, reproducing a bug is not mandatory. Again, the author does not claim that reproducing a bug is mandatory, but rather that it is a useful practice.
It is not a virtue to try to base your work on the least amount of information.
You say that "when the hard problems come you will be ready". But really when the hard problems come you will look at them assuming that the logged information is enough to solve them. No programmer can know ahead of time what information to log, so by purposefully blindfolding yourself from experiencing the bug directly, you might miss the bigger picture.
Indeed, I can't think of a reasonable logging mechanism that the author could have thought of ahead of time that would have helped with this particular bug. Emphasis on "reasonable".