C Runtime Overhead
ryanhileman.info
ryanhileman.info
strace -tt /tmp/test
15:38:33.245549 execve("/tmp/test", ["/tmp/test"], [/* 16 vars */]) = 0
15:38:33.246436 arch_prctl(ARCH_SET_FS, 0x601070) = 0
15:38:33.246617 set_tid_address(0x6010a0) = 17227
15:38:33.246875 exit_group(0) = ?
15:38:33.247060 +++ exited with 0 +++
Aah, the promise of the Exokernel[1].
[1] http://u.cs.biu.ac.il/~wiseman/2os/microkernels/exokernel.pd...
Serious application development for consumer computers was done in straight Assembly. Anything higher than that was "scripting". :)
The first autocode and its compiler were developed by Alick Glennie in 1952 for the Mark 1 computer at the University of Manchester and is considered by some to be the first compiled programming language. His main goal was to make the programming of Mark 1 machine, known for its partcularly abstruse machine code, comprehensible. Although the resulting language was much clearer than the machine code, it was still very machine dependent.
Ron Minnich presented a paper at the Plan9 conference
Using Currying and process-private system calls to break the one-microsecond system call barrier [1]
Which is a technique he's been using on Blue-Gene. There are some more papers linked from the IBM Research website. [2]
[0] http://www.networkworld.com/article/2317965/tech-primers/os-...
Sorry the conference repository server is awol, all I could find is this :
[1] https://www.yumpu.com/en/document/view/7777923/using-curryin...
[2] http://researcher.watson.ibm.com/researcher/view_group_pubs....
Of course you can use Rust without stdlib, but that leaves you with fairly barebones language.
My favorite part is my x86_64 linux syscall interface, in which I define a short assembly stub (syscall.c) and simply list syscall names (syscall.list) to wire them up. The calling convention between user-space functions and the kernel was close enough I was able to reduce the inline asm to two instructions (`mov rax, syscall num; syscall;`).
What are the functions in libc that you truly need and use?
How many of those routines are you incapable or unwilling to rewrite yourself?
I have asked myself these questions on many occasions.
The cruft laden operating system I use (which I imagine by many people's standards is "bare bones") is hopelessly dependent on "libc". But I find that the userland programs I truly need can be written using only a minimal subset of the many functions that are compiled into it.
Could I write any of these in assembly and create my own "libasm" to do the syscalls? This is a project I have contemplated many times.
To motivate myself to keep writing more assembly in the age of scripting language du jour madness, I write a small program in assembly language and then write the same program in C and then step through each compiled program (e.g., with ald). Seeing some of the overhead is an effective motivator.
But note my goal is not to achieve "better" than what someone else has written.
My goal is having more control over the result and to practice assembly language programming. And... to eliminate some of the "overhead" I see.
Big difference between my handwritten assembly and what gcc generates. Regardless of which is "better" one is more succinct.
Perhaps I drifted a bit from the original blog post; I think the author may have just been trying to find a speed increase, rather than chasing after some sort of minimalist aesthetic.
For example, I wanted to use C++11's std::array to return fixed-lengths arrays by value on the stack. This was in a minimal API that originally included NO other headers. In MSVC, std::array included a tree of 100 other headers including gems such as istream, exception, new, malloc, and float.
Yes, I'm aware that none of this stuff gets compiled into the binary, but it shows a disregard for lean dependencies that is emblematic of a bigger problem.
But yeah, it drags in a whole tree of headers. Minimizing this is possible, and we've taken steps to do so in the past, but it's a lot of work for possibly minimal benefit - user translation units tend to drag in many STL headers, especially if they're using precompiled headers like they should. We think our time is better spent fixing correctness bugs and implementing features.
Or, <algorithm> gets most of its subtree via <memory>, which it needs for a few algorithms that use temporary buffers. That one is beyond your control... The standard should add an <algorithm_lightweight> header containing only algorithms that don't need extra memory. Or you could add your own such header to use internally.
Users must be able to include <iterator> by itself, and get istream_iterator.
We do break things up into internal headers, we just haven't done that as much as possible. Come to think of it, equal() is defined in one of our central headers, so we should probably take advantage of that in <array>.
[0] https://github.com/lpsantil/rt0
[1] http://asm.sourceforge.net/asmutils.html
[2] https://github.com/leto/asmutils/blob/master/src/httpd.asm