Writing C software without the standard library
weeb.ddns.net
weeb.ddns.net
The syscall/errno stuff has always seemed unusual, inelegant, and inefficient --- instead of just returning a negative error code directly, the function returns the vague "an error has occurred" -1, and you have to then check errno separately after that. It only adds insult to injury when you realise that the kernel itself isn't doing it, but the syscall wrappers. And thanks to POSIX standardising this mechanism, the alternative will likely never get much adoption; of course, if you write your own syscall wrappers like this article, then you can skip that bloat.
For now this guide is linux-only, but I will be writing a windows version when I feel like firing up a virtual machine.
Unfortunately the Windows syscalls are not officially documented and even less stable than on Linux, changing even between service packs.
http://j00ru.vexillium.org/ntapi/
http://j00ru.vexillium.org/ntapi_64/
At least on Linux the first few (i.e. the oldest, most common and useful) syscalls have not really moved around over the years:
IIRC raw syscalls are an officially supported kernel API, that's why you can have alternate libc implementations (e.g. musl), and Linux is an oddity in that, on most systems even if the syscalls are fairly stable there are no actual guarantees with respect to them, and the only officially supported interface to the kernel is the standard library. OSX does not allow statically linking libSystem for that reason for instance.
It makes me think twice about e.g. using Go as part of the build workflow in a big project if it's going to randomly stop working on various OS X versions.
What would have been the reasons for targeting syscalls directly instead of the C standard library?
Since the standard library probably assumes (and requires) a C stack, linking against the standard library would require cgo (or some other specific workaround) on non-linux platforms.
IIRC raw syscalls are an officially supported kernel API
The term you are looking for is We don't break userland. Once a systemcall goes live it is engraved in stone. This is why if you look though the syscall table you see call, call2, call_ext.Sure, Linux makes promises about its syscalls. But that's not the only way to get robust backwards compatibility and as GP says it's the oddity in that respect.
You can get pretty close though; it's possible to skip the C runtime and most of the other user space libraries and call into ntdll directly. Many of the functions it exports are fairly thin wrappers over the system calls.
And non-reentrant.
errno is usually defined as TLS (thread local storage).
It's a mess.
Obvious in hindsight, not so much in the 1970s.
VMS sounds amazing, but somehow I'm still thinking this might be nostalgia.
If multithreading were pervasive nobody in their right mind would choose to use a global variable for error status codes.
Sure, there may be some overlap or interleaving between Standard Libraries, syscalls, and posix, but they are definitely not the same.
Still I like the idea. This is something that should be covered in a CS 102 type course. I know way to many cs guys who have no idea how to debug, let alone how their is being implemented.
Many linux developers believe "code is portable" means "can run on different linux distros".
Edit: To be fair, he said architectures, and i think he meant CPU architectures, i.e. AMD64 and i386.
The BSDs are all exactly the same for every architecture, sane.
That's not portable even across versions of the same OS, many OS don't support raw syscalls and don't make any guarantees about them. That's why you can't statically link libc on OSX for instance, libSystem is literally the system's interface and necessarily dynamically linked.
The proper abstract interface on non-linux systems is the standard library.
These are nowadays supported by Windows and few BSDs too, who needs more portability than that :)
Well, it is hard to see an alternative to that outside supplying your own OS with the library...
POSIX only specifies a set of C library functions. Once you stop using them, you are on your own.
You get the same sorts of benefits though.
You can get reduced executable size and a lack of dependencies on various msvc*.dlls, without giving up much of the functionality of the C standard library and without having to write it all yourself.
So you can't really go without it.
I discovered a little book by him after wading thru recursive descent parsers chapter by Herb Schildt... I don't want to get down on Schildt, I enjoyed that too, he showed some cool stuff..
but Plauger had me thinking that programming could be a craft, an art, a noble intellectual pursuit.
Of course, if your target archs are all supported by libc, porting is much easier with libc.
This is why lots of embedded, secure and defense software doesn't use standard libraries, and printf will instead be called sio04583 or something....
Also.. Don't forget that you can get gcc to get rid of unused code when compiling static executables, and use sstrip (yes, two 's' - it's a different program) to strip even more, if it's an ELF binary...
* The space savings are moot as other processes such as the daemons are going to load libc into virtual memory anyway, and the kernel shares libc's page among all processes.
* This adds a lot of LOC you have to maintain, instead of shoving it off on the compiler/libc vendor, this increases the chance of bugs.
* This will prevent the use of VDSOs to optimize high volume system calls like gettimeofday.
* It's still probably good to know how these happen, even if you're not doing them yourself.
* The only place this would really see benefit is in a single process environment, however in those cases I would suggest a unikernel anyway for simplicity sake.
The tooling answer to this, it seems, would be to support statically linked libraries. But again, I would want to see numbers before personally worrying about this.
* Your executable is going to be loaded as a whole page regardless of how large it is, on most platforms this means you'll need at least 4k of User VM.
* You'll need a page table, which has its own overhead. If someone was going to push back on the libc assertion I'd expect it here, a the PTEs for libc and and the VDSOs cannot be shared between processes (as far as I'm aware).
* I would expect it to in theory RUN faster assuming it was a small toy program like the example, this is because there is less work to be done even with shared pages.
I think there is a strong argument that this page is often already paged into memory from everyone using it. However, if the function you used from it would have fit in the pages you were already using for your application, I could imagine some benefit.
I continue to stress, though, that this is just imagined. Numbers would be first thing I would have to collect before acting on this. (And I hope it doesn't sound like I am tasking you or anyone else with this. That is not my intent.)
A good libc developer could actually put these into separate pages based on how often they change. In other words, if a static only is ever set once then coalesce it into a page with other statics that are only set once and are not process dependent.
Why that's important: because of fork, when fork creates a new process it sets the parent processes pages read only and then preforms copy on write when they are modified. In theory you can share both the binary and some of the statics between all the processes.
That said techniques to allow a binary to fit under page size do have value; as the unused extra page can be used for other things.
Why can't I just call gettimeofday() which will access the vdso page mapped into my address space?
The only bit i'd disagree would be with games, at least on AAA games since i have some experience there and -at least at the engine level- there is still a lot of low level wizardry being done there.
> I've even seen functionality regression for the sake of modern and
> "responsive" design in many cases.
>
> For example, you used to be able to browse youtube favorites in
> pages. Whoops, not anymore! Now you need to scroll down and
> painfully wait for the site to make its stupid animation and load
> the next page.
> What's worse is, you can't just skip through pages. You have to
> load EVERYTHING and it all stays loaded, so after 10-20 pages the
> browser starts to lag and hog cpu just to scroll.
I see this issue more and more, scrolling is broken in so many sites,
even popular ones like FB.
Plus, often search doesn't look in the "scrollable" content, i.e. you
need to load (=scroll) all these "pages" and rely on the browser
search function.
Many small things like this, e.g. snippets of code that don't fit the
DIV box and you need to scroll horizontally. And news pages lagging on
a i7 w/ 16Gb or RAM. It's all so sad.That's a feature. They want all this valuable search data so the make you search via the site instead of the browser.
Web is especially good example. With all its short-comings, it's still sandboxed and portable platform with very capable (and ugly) programming language for delivering same codebase across desktop and mobile devices (even servers, unfortunately). There are many hits and misses like TODO apps built on Electron, but that's to be expected with any technology and isn't a reason to dismiss all the good applications.
To me, consuming all the hardware resources of each new generation seems to be the price we have to pay to build more complex software in larger and larger quantity.
What's bothering me is that it doesn't seem to be converging in desired unification of different applications and services. Everyone is building their own stacks and environment without much care for interoperability. For example, few years ago it seemed clear that XMPP is a way forward in chat clients that everyone should adopt to enable federation. Nowdays we are expected to have 3+ chat clients installed on phone and remember which contact is on which service. I really don't understand how users put up with it.
I keep meaning to finish up my Gopher Client and release it upon the world...
https://www.weegeeks.com/upload/Jar-Gopher-Browser-Full-Scre...
I don't consider variable-width fonts, nice margins and dynamic word wrapping to be a waste of resources.
There's a sane middle ground between Electron and the 70s.
Sadly, neither of the things I mentioned above exist yet. People are working on languages that fit that space, but most of them aren't done yet. Nim is the only one that's officially released.
As for desktop UI libraries, your choice these days seem to be between bad and worse. I'm working on my own desktop UI library, but I don't know what the hell I'm doing. It'll be a learning experience, if nothing else.
You might also want your UI library to be cross-platform, which will limit your choice of languages, especially if you ever want to port to mobile.
I'm hopeful about Go GUIs though. Go is an easy language to use and performant enough for complex GUI apps. Google abandoned gxui but it seems like they are backing Shiny at least a little.
He explicitly calls out people "writing poorly performing code on top of pre-made engines". I'm pretty sure he's aiming most of his ire at non-indies using Unity3D or Unreal.
Hatoful Boyfriend (remake that uses Unity), a simple graphical choose your own adventure with no animation at all that is painfully slow on my 3-4 year old laptop (although this must be intentional for some reason, it is hard to imagine how that could be achieved purly by abuse of unity)
Lethis - Path of Prgress a 'hipster retro' sim city in a steam punk world that fails to work at all with Intel integrated graphics.
That being said, I'm happy the web isn't still just plaintext or minimally styleable
[0] https://git.musl-libc.org/cgit/musl/tree/src/misc/syscall.c
[1] https://git.musl-libc.org/cgit/musl/tree/src/internal/syscal...
[2] https://git.musl-libc.org/cgit/musl/tree/src/internal/syscal...
[3] https://git.musl-libc.org/cgit/musl/tree/arch/x86_64/syscall...
glibc's source feels like archeology: there's so much history and remnants of bygone eras.
musl feels like a structured, well-engineered and specified piece of architecture.
This isn't a knocj against glibc. But it was grown, not built.
musl and the team's great documentation have been incredibly handy when I've been building tightly constrained applications.
_syscall5:
mov %r9, %r10
_syscall3:
mov %rcx, %rax
syscall
ret
And then: extern unsigned long _syscall3(
unsigned long, unsigned long,
unsigned long, unsigned long);
extern unsigned long _syscall5(
unsigned long, unsigned long, unsigned long,
unsigned long, unsigned long, unsigned long);
#define syscall0(NUM) _syscall3(0,0,0,NUM)
#define syscall1(NUM,A) _syscall3(A,0,0,NUM)
#define syscall2(NUM,A,B) _syscall3(A,B,0,NUM)
#define syscall3(NUM,A,B,C) _syscall3(A,B,C,NUM)
#define syscall4(NUM,A,B,C,D) _syscall5(A,B,C,NUM,0,D)
#define syscall5(NUM,A,B,C,D,E) _syscall5(A,B,C,NUM,E,D) typedef unsigned long int u64;
typedef unsigned int u32;
...
if you define your own types like this you may need to revise them when you switch architecture or even compiler.Now you could argue that this is part of the standard library, but I actually see it as a part of the standard C language.
int main() {
assert(sizeof(signed char) == 1);
assert(sizeof(short) == 2);
assert(sizeof(int) == 4);
assert(sizeof(long long) == 8);
return 0;
}
I'm not interested in the language lawyering, because yes I know the standard provides more freedom to compilers. I just think those definitions are very universal for any real computer that would otherwise run my software. Please don't bring up Windows 3.1, that's about as relevant to most of us a PDP-11.And for what it's worth, using typedefs based on the above provides more readable printf strings. This is hideous:
int main() {
int64_t portable = 123;
printf("Ugly: %" PRId64 "\n", portable);
return 0;
}
Where this is acceptable: int main() {
long long palatable = 123;
printf("Better: %lld\n", palatable);
return 0;
}That leaves out either the 2-byte or 4-byte int from the standard data-types, but you can gain that back by simply using a compiler attribute or compiler-defined type to allow access too it. While that sounds non-standard, it really wouldn't be that bad because it could simply be used in `stdint.h` to expose the standard `int16_t` and `int32_t` types, which could be used like normal.
All that said, POSIX requires `CHAR_BIT == 8`, which then basically ensures what you wrote to be true. So if you're willing to target POSIX (or POSIX-supporting systems) then such a thing is perfectly fine.
The `stdint.h` types are still better though, IMO, because they make your intentions a lot more clear.
Well, I've already learned something new. I assumed that convention was from the compiler. This is a great resource.
cpp <<< "#include <stdio.h>"|grep size_t
which is super convenient
[1] https://github.com/abh/djbdns/blob/master/str_len.c [2] http://cr.yp.to/djb.html
This is almost what Duff's device solves, except then you need to know the length beforehand.
My assumption is that DJB tested this locally and found enough of a speedup that it was worth it, considering the very low added complexity and risk of major degradation / defects on untested platforms.
There's no doubt that this is a fun project though - if you or someone-else enjoys this type of stuff, you should definitely try your hand at writing a simple Unix kernel or similar, you'd probably enjoy it.
On that note though, the writers aversion to inline assembly is unfortunate. It's a necessary evil for this type of programming. The syntax is ugly, but it's not really that hard to get used too (Especially since the large majority of inline assembly is just a few lines long, or even just one line long). In particular, the syscall wrappers can be done in a one-line piece of inline assembly, and then you can avoid the function-call overhead for the syscall by placing the inline assembly in a `static inline` function in your headers (Or a macro if you prefer), as well as avoid the extra .S file (Which IMO is the better part - it's always easier when you don't have to mix different languages like that).
I would also add that, while I used to share the aversion for AT&T asm syntax the author does, virtually all of the assembly code out there related to linux is written in AT&T, so it's worth it to get used to it and at least be able to read it. On that note, you can use the Intel syntax in inline assembly though, if you prefer, so even if you hate AT&T with a passion you can still write inline assembly ;)
When I wrote that comment I asked myself should I have written "malloc" or "malloc/free" - surely one implies the other.
Of course, you need to be careful as if you write code like that in a language without garbage collection, it's inherently not reuseable - retrofitting deallocation is often really painful because it gets easy to adopt patterns that make object ownership etc. unclear when you don't have to ensure it's easy to deallocate in the right order.
Arenas are really nice if you're allocating a lot of objects of the same size, whereas malloc() must be prepared to handle a lot of different memory usage patterns.
TeX basically uses a special purpose implementation of malloc/free, with a static array as backing instead of memory requested from the OS with mmap(2) or sbrk(2). The main reason is portability (the original version was released in 1978 using WEB/Pascal).
One of the reasons that I love HN is how informative you all are!
The main loops are all assembler, but the support code is in C, but the C code is used as a more expressive assembler, and linking with libc requires way too much memory.
All this means that I don't even have things like memcpy() available. In a way it's a quite liberating way to program, since you are in full control of the hardware.
I guess yesterday's computers is today's embedded hardware.
There's also the case where your API accepts a pointer and a size, but you don't want to have lingering pointers into the caller's memory, so you have to copy the data over to the "inside" of the API. This kind of design is perhaps less common in demo software, but certainly plausible in embedded products which at least try to be somewhat optimized.
Much of the C code is used during precomputation of data before the actual time-critical code is run. This involves copying lots of data in order to set it up so that as little computation as possible is performed in the actual time-critical parts.
Is this ever an real issue, even on any embedded system in the last 20 years?
(I've never needed to remove stdlib yet).
It's a statement of one's professional competency.
Ask Cisco when they cut the Linksys routers' RAM in half a few years ago. Every byte counts. Component cost savings add up when you make a few million of them.
But sometimes it can take so much extra work that it's not worth it. So gotta do the cost benefits analysis or hell if it's something you're interested in doing just do it anyway.
Of course, removing libc won't be your first (or second or third or ...) step for removing bloat from a mature codebase.
An implementation of malloc/free
Functions to parse and print floats (somewhat system dependent)
Assembly implementations of any trigonometric functions used
While there is code that goes to that effort (The Go runtime comes to mind), it's quite a pain for "normal" code.
"It's often necessary to either push useless data or simply align the stack pointer when the pushed values don't happen to be aligned."
That's kind of hand-wavy. How do we "simply align the stack pointer"?
Of course, you now need to keep track of the old stack pointer so you can restore it. Most code saves the old stack pointer into the fp, so you can just do sp = fp to undo any pushes without needing to care about how much was pushed; but it's cleaner and more efficient to have the compiler arrange things on the stack so that everything's already aligned and you don't need to do it programmatically.
...why, yes, I have spend the past couple of months with my head buried inside a compiler backend; why do you ask?
And there is no portability. It only works with the specific architecture's calling convention and the specific c compiler.
Is this faster than a (const) mov ?
Also, the reason it is "faster" is that the encoding is 1 byte, vs. 9 bytes (in 64 bit) for "mov rbp, 0" - roughly, 1 for "mov rbp,", 8 more for a 64 bit "0".
Another reason why it was faster was that the processor recognized it and avoided partial flags stalls after an "inc". But in 64-bit code you rarely have "inc" at all, so it matters less. On the other hand, a few years ago XOR had a false dependency on the register you're clearing; I'm not sure it is still that way on more recent processors.
"Also, I like how returning 0 is "xor eax, eax"."
https://news.ycombinator.com/item?id=13052503
That led to multiple replies on it.
focused on windows, because grafix. also, linux guys are more likely to use C anyway.
I only remember a comment about avoiding exceptions, which I might have read first in the context of micro controllers.
https://www.google.de/search?q=site%3Apouet.net+c%2B%2B+exce...
https://www.google.de/search?q=site%3Apouet.net+c%2B%2B+stdl...
How do we summon Dang? :)
tldr; practicality and longevity should trump netiquette.
Hacker News is very technical and code heavy. Seems to make sense to me that some may want to communicate / discuss code itself. I could even see it opening up more conversations like "this is my implementation of X; thoughts?" or "do this in any other language" challenges.
On the other hand backslash escaping (the lack of it and emphasis clusterfuck drives me nuts), block quotes, inline code/monospace and links I really miss.
Section titles, lists and tables would be nice as well, though not deal breakers.
Considering the maintainers have repeatedly refused to make any improvement to comment formatting in almost 10 years since HN was created, I'm not holding my breath though, it's obvious nobody cares about comment authorship and craft.
The "hello world" example is just the first step to annihilate your capacity of understanding how thinks works by relying on institutional black magic, that maybe wrong
(see all the scanf bugs that have been living in C code for so long and all bugs coming from respecting the old's man wisdom)