Can I use a system call?
justine.lol
justine.lol
I would love to know how this table is generated and what the test process is. What's the best way to run tests like this across such a wide variety of OSes? Maybe vagrant images?
https://www.youtube.com/watch?v=GUQUD3IMbb4&t=85s -- link at the bottom of the file
https://en.wikipedia.org/wiki/Swedish_Rhapsody_(numbers_stat...
Don’t you need to request heap space from operating system via system call?
Something else to consider is that there may not even be an operating system the program was run on.
Sure, in this case you can just implement malloc the same way a kernel implements me. My failure of understanding is when there _is_ a kernel that your program needs to interact with.
That way if you need to add new platform, it's only one place; you're shielded against weird platform bugs (Apple goes one way, Microsoft the other), so the app behaves the same; you can tweak system to your particular use.
For performance reasons (avoiding a syscall) almost all malloc implementations only use sbrk / mmap occasionally and keep the malloc calls in userspace.
It’s not that libc is used to fix kernel problems/compatibility issues, but rather, the kernel ABI is not a supported, stable public API.
On macOS, there are cases where using the syscall ABI directly will result in your code breaking when interacting with any other code that does correctly use libc, due to out-of-sync userspace state maintained by libc.
For example, the fork(2) implementation in macOS’ libSystem will invalidate the cached copy of the current process’ pid used by getpid(2).
If you fork(2) by directly trapping to a syscall, the cached pid won’t be invalidated; any future calls to getpid(2) through libSystem will return the stale pid.
If you’re only using syscalls and no libc at all, you’re fine (on Linux, and only Linux).
Yes. As a glibc maintainer once said, if you use clone you're on your own. And that is fine.
Meanwhile OpenBSD was discussing ideas to extend their system-call-origin verification to only allow syscalls from libc.
It's true that we don't need to load the vDSO into the address space of the process since the kernel does it for us. We do need to manually link the vDSO functions though: the kernel uses the auxiliary vector to pass the address of the vDSO ELF image to the process, allowing it to parse the ELF header and find the addresses of the functions. In most cases, the libc will do this while initializing itself prior to calling main.
Nope, it is built into the kernel image - at least on my system, but there doesn't seem to be a config option to disable that.
> When mapped into the process address space, it's just a normal ELF image. The functions all conform to the normal C ABI for the platform.
That's purely for convencience because it allows ASLR for the vDSO and there is no point in making it a new ad-hoc format if every C library already needs an ELF dynamic linker - if the vDSO was a thing before ELF it would surely have a different format.
> We do need to manually link the vDSO functions though: the kernel uses the auxiliary vector to pass the address of the vDSO ELF image to the process, allowing it to parse the ELF header and find the addresses of the functions. In most cases, the libc will do this while initializing itself prior to calling main.
You need to look up the entry points in the vDSO before you can call them but that is really no different than checking the kernel version before deciding which syscalls to invoke. And no, the vDSO is not loaded like any other .so - it is very much a special case [0] even if the libc ends up reusing some of the normal .so code and structures. For example, it needs to be loaded even for statically linked binaries [1].
[0] https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/sys...
[1] https://sourceware.org/git/?p=glibc.git;a=commit;h=1e8bdc3a2...
What’s so crazy about an OS vendor wanting to exercise control over that userspace shim layer, rather than assuming that user programs will be poking directly at the kernel?
You can argue that having multiple libc implementations is good, sure, but it’s not the only way of doing things and it’s not without tradeoffs.
Looking at this through a KERNEL32 lens, I can actually somewhat see the alternative point of view: in NT < 4 it was mostly an in-process RPC proxy for the subsystem server in CSRSS, and one can argue that libc on Unix-likes is in the same position (cf the vDSO on Linux). You can build RPC to a trusted server either way: define a stable wire protocol; or require the client to load proxy code into its address space. COM on Windows and hardware-accelerated graphics on Linux both take the second approach.
But then why the hell does the proxy also have opinions about memory allocation, assignment of TLS slots, or floating-point formatting? (KERNEL32 has them on the first two points as well, mind you.) History[1] aside, does this really look like good engineering? (KERNEL32 makes a bit of a point there, but given how much HeapAlloc sucks as an allocator compared to the interoperability benefits it brings, I’m not sure it should be taken too seriously.)
[1] https://utcc.utoronto.ca/~cks/space/blog/unix/UnixAPIAndCRun...
I think my ideal system would look like Linux in some ways. Any given Linux system has a 'standard' shared library ecosystem, with a single shared libc, and usually a shared C++ standard library as well. Executables and libraries that participate in that ecosystem can assume they'll use the same standard libraries and so can pass those libraries' types across module boundaries. Alongside them are 'hermit' executables, like Go programs or C programs statically linking to musl, that avoid the standard libc. But those executables typically also avoid all other system shared libraries (or greatly constrain their usage), so the difficulty in using their APIs (due to mismatched standard library types) isn't a big deal. Now, Linux sort of forces this approach by having the dynamic linker itself be part of libc. But that has some downsides, such as forcing the vDSO to work differently from every other library. I think I'd prefer to have a universal dynamic linker, like on Windows, but to still have 'ecosystem shares a libc, while hermits stick to themselves' as a convention.
As described elsewhere in this thread, that is the policy of MacOS X and Windows.
It is also a policy which OpenBSD is moving towards:
https://lwn.net/Articles/806863/
According to the Go devs, this is also the policy on Solaris-family OSs (Go uses libc there):
https://github.com/golang/go/issues/24357#issuecomment-37300...
AFAIK, it is only Linux which explicitly supports the kernel system call interface, and so where alternative libcs are possible.
You can do this on Windows.
The syscall numbers change sometimes.
I use it to evade EDR hooks :)
Could you elaborate on this? Specifically I'm not understanding what exactly is "explicit" with regards system calls on Linux compared to other OSes.
Libc is seen as user land by the Linux kernel, it provides an API for programs to abstract them away from syscalls. Libc dies syscalls for programs.
By "never breaking userland", Linux promises to never break libc by never changing syscall numbers/args/types - new syscalls are added in ways that don't break older syscalls.
Other platforms don't make this explicit promise and reserve the right to shuffle syscalls around, change them, remove them, etc as desired - making maintaining a libc for them a pain in the ass, and programmers are told to only use the libc as syscalls don't come with guarantees.
You can (and I often do) program on windows using only syscalls, but your program tends to need rebuilding across major versions as the syscall interface isn't stable.
Fuck me typing that on a phone was painful.
You're allowed, of course, it's just not supported. We pop out of the woodwork to note that it's not supported and will break.
Linux's approach of treating the syscall ABI as a long-term supported, stable interface is somewhat unusual.
> Cosmopolitan Libc is a libc, so it can't very well depend on six other c libraries
At least on macOS, trapping to the kernel without going through libSystem first is simply not a supported use-case.
"Apple does not support statically linked binaries on Mac OS X. A statically linked binary assumes binary compatibility at the kernel system call interface, and we do not make any guarantees on that front. Rather, we strive to ensure binary compatibility in each dynamically linked system library and framework."
More to the point, platform vendors decide what their ABI boundaries are. Historically Unix vendors (back in the late 80s and early 90s) had no ABI boundaries... the expectation was every new OS release would require a recompile of all software. Obviously Linus has very different ideas about stability than the other Unix like systems of the time (which was a good thing), and focused on syscall stability. That made a lot of sense since he only wanted to maintain a kernel, not a full OS distribution.
When modern macOS and Windows developed their ABIs they both were relatively mature OS distributions including a dynamic linker and default runtime libraries, and the ABI boundary chosen was well above the kernel as that is an easier place to define and maintain it.
How do we actually know that placing the boundary above the kernel makes it easier to define and maintain?
Or: Life can be easier for the kernel team if they're allowed to make breaking changes. Then task the lib team with writing shims / wrappers / etc. to fix all the problems which that causes. Then the manager of the lib team may have a perfect reason to boost his headcount. Then...
Without it, you end up maintaining system calls that nobody should use.
Now, you could argue that moves the mess to the c library, which would still have deprecated functions that used to call old system calls, but now are built on top of better ones, but there’s more flexibility there. Application programmers can, one by one, move to newer c libraries that remove that cruft.
Having a stable syscall interface is one thing... not supporting static linking is another... Even on linux, statically linking libc, GTK, ... is not a great idea.
I also had a Mathematica binary (statically compiled except for libc) that ran on Linux from 1998 to 2010, including X windows (at some point, somebody moved the X files to a different location, so I had to set an env var).
(I'm not arguing for nor against syscalls-as-kernel-API, I'm honestly just curious. Without knowing too much about it, it seems pretty sensible to support statically linked executables, but I might be missing something. I guess linking against an "as old as you need to support" libc dynamically isn't the end of the world, but it does constrain the build environment somewhat.)
I can't speak for Windows, but on macOS, you build for older targets by passing `-(mmacos|ios)-version-min=` to the compiler to specify your deployment target version; this is used to determine symbol visibility, API visibility, toggled #ifdefs, etc.
The provided version is also stored in the Mach-O load commands of your executable, and will used to select compatible symbols (and enable compatibility shims) when loading your binary's image at runtime.
No need for static linking — or jumping through hoops to build against an "as old as you need to support" set of installed libraries.
For example, look at [1] and search for e.g. 0x01aa, to see how the meaning of that syscall number changes in a pretty systematic way over releases.
1: https://github.com/j00ru/windows-syscalls/blob/master/x86/cs...
Use of DLLs is baked VERY deeply into windows. The kernel is not even a monolithic file. It is an exe like most others pulling in dlls that implement other functionality. It even features like API-set DLL redirections which mean the bootloader for the kernel needs to implement a fair bit of the PE loader functionality of the kernel just to load the kernel. (I've no real clue if there is shared code between the two loaders, or if they are two separate loaders that implement similar things. Many of the options/features of the full kernel executable loader are not really needed in the bootloader.)
So it is not much of a surprise that making the numbers stable to support fully statically linked executables was not really something they cared about. From their perspective you could always just pull in the ntdll.dll for your syscalls like they intended (or more likely a higher level win32 dll that uses the syscall).
Except you're not doing that consistently though.
On OpenBSD, mincore(2) was removed 3 years ago, and it's UNIMPL syscall 78 was eventually recycled by a different system call: mquery(2), which your library calls expecting mincore and passes bogus arguments to.
I would be very surprised if there aren't more serious mistakes lurking in your library, in fact I know there are.
https://github.com/openbsd/src/commit/54e4f6b9a1dc183e1dc7c4...
https://github.com/openbsd/src/commit/1d60349d0b961891264d42...
https://web.archive.org/web/20220921050525/https://justine.l...
First the code to say hello:
const char \*const str = "Hello World!\n";
HANDLE standardOutput = GetStdHandle(STD_OUTPUT_HANDLE);
WriteFile(standardOutput, str, strlen(str), NULL, NULL);
Then the compiler switches: /merge:.rdata=.text cuts out an entire section 0.5K
/nocoffgrpinfo cuts out about 256 bytes from your .text section
/emittoolversioninfo:no is supposed to omit the RICH header, but doesn't seem to work anymore.
/stub:stub.bin will let you replace the default DOS EXE stub with something smaller. Using a 64 byte file will get the EXE header and PE header to fit in the first 512 bytes (when combined with merging the .rdata and .text sections)
Set "Entry Point" to main. This completely bypasses the C standard library and CRT, none of that gets initialized or called.Then you end up with an EXE containing 3 sections: The EXE header (512 bytes), the .text section (512 bytes, but only 141 bytes actually used in there), and the .idata section (512 bytes, for importing DLLs, only 41 bytes actually used)
Is that really the case? Given it’s specifically intended as a highly portable libc, in some cases maybe it could be implemented by calling through to the blessed system libc, rather than kernel calls.
I mean, that likely wouldn’t fit with your goals for the project, and it would likely need some horrible linker hackery, but in principle it seems technically possible. (I’ll believe you if you say it’s completely impossible, though!)
Waitable Timers on Win32 let you request time values in units of 200 nanoseconds. See `CreateWaitableTimer`, and `SetWaitableTimer`.
`Sleep` and `SleepEx` can be implemented using a waitable timer, just use `WaitForSingleObjectEx`.
If you need to wait for an object (semaphore, etc) using a time unit other than milliseconds, you can use `WaitForMultipleObjectsEx` with one of them being a waitable timer.
---
The next question is if Win32 can actually deliver those requests for precise times or not. From the testing I did a while ago where I was simulating Sleep, I got actual sleep times rounded to about 4ms. So much for requesting nanosecond level precision.
"Your program will also boot on bare metal too. In other words, you've written a normal textbook C program, and thanks to Cosmopolitan's low-level linker magic, you've effectively created your own operating system which happens to run on all the existing ones as well."
"Annoyingly complicated".
This is of course a little annoying at times.
There's not really a strong argument for not linking libc. Even if you want to implement your own libc, you can still do so with linker tricks and calling through to the supported syscall wrappers in the real libc.
However, if, for some reason, you really want to re-implement libc, you still can, but your library will have to link against the system-provided libc for the syscall entry points, because that's the only stable interface to the kernel.
There's no reason that should be a problem for a hypothetical libc-reimplementor.
Sure there is. The libc sucks. Freestanding C actually turned out to be a superior language because there's not as much legacy weighing it down. There's many systems languages out there, nobody should be forced to link to C stuff.
While you cannot link directly to libsystem_kernel.dylib, if you choose to ignore everything in libsystem_c.dylib and only use the syscall wrappers reexported from libsystem_kernel.dylib via libSystem it has practically the same effect on macOS (in fact, the resulting binary will be identical to what a hypothetical binary linked to just a single libsystem_kernel.dylib would be except for a single `LC_LOAD_DYLIB` command).
Such a binary would have the system lib c initialized sitting in its address space, but for the rare binary that really wants its own libc that doesn't seem unreasonable.
It's annoying that libc combines two completely different things - userspace utility functions, system calls - but unfortunately, that's where history has brought us.
Not on Linux.
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
The system call interface is not only stable but also language agnostic. The correct place for this inferface isn't in some C library, it's in the language itself. We could have compilers directly targeting this. GCC could add a system_call keyword that emits code conforming to that ABI. Dynamic languages could have a JIT compiler that does the same thing.
That said, the vDSO is much smaller than any libc.
In this case that means the api designer gets a little bit of wiggle room in the process userspace where it might be more appropriate to say shim some calls so that they're no longer 1:1 with a syscall. The obvious example is that malloc() can call syscalls for you and is the "default" way to allocate memory you might want to provide but It's more of it's own little runtime.
My favorite example was in my time at LinkedIn. Because the internal Kafka team had thick client libraries, making the company wide change to enable encryption was as simple as pushing out new libraries and deprecating old ones.
If OpenBSD wanted to, they could implement something like io_uring with support for all kernel functionality, port libc to use that and ditch conventional syscalls entirely (simulating blocking where needed), without user-space knowing anything changed.
Under Linux, which is actually rather unusual in providing ABI guarantees, you're stuck maintaining bug-for-bug compatibility for every syscall you've ever written, and cannot change user-space no matter how good that change would be for either or both.
But instead you have to maintain bug-for-bug compability in libc for every API you've ever written. In case of macOS where the kernel and libc/libSystem developments are closed and done by the same entity is makes zero difference (imho).
Also, I understand catiopatio's argument in the sibling comment about doing more in user-space than in kernel in case of a bug, but it breaks down the moment thin wrappers around syscalls exist there. You link to libc/libSystem and use every thin wrapper in existence - now no syscall can be changed (if I understand correctly how macOS works, never used one)
It is not instead. With syscall ABI you need to maintain both, without it you only need to maintain libc (a majority of which is dictated by POSIX anyway).
> In case of macOS where the kernel and libc/libSystem developments are closed and done by the same entity is makes zero difference (imho).
As above, maintaining two contracts is harder than one.
Proprietary parts aside, most OS's have their kernel and user-space developed together, and that is exactly what allows them to work this way.
This is used to adopt new features, change or deprecate old features with a brutal efficiency that Linux cannot compete with - e.g., when OpenBSD implemented pledge in both kernel and all relevant tools.
This is one of the reasons that these projects can keep up or in some cases surpass Linux (FreeBSD networking is seen as superior, and you used to get better performance from running Linux binaries on FreeBSD through its compatibility layer) despite having much smaller groups of maintainers and users.
There are no alternative libc implementations for *BSD, unlike Linux the system provided libc is the only libc that matters.
When people talk about the supposed requirement to depend on the platform libc dso, what they're actually talking about is Apple frowning upon statically linked binaries. https://developer.apple.com/library/archive/qa/qa1118/_index... They don't forbid it though, like Microsoft and Fuschia do. Fuschia for instance uses RIP origin detection. Windows does it by changing the RAX ordinals fortnightly. Apple simply asks that we say, hey, if you use Cosmo there's some risk Apple might break our binaries. We take proactive steps to avoid that happening with Cosmo. For example, we don't do some of the things Go did, like reverse engineering the memory layout of Apple's time functions. Cosmo sticks to the APIs that are shared by UNIXes in general, e.g. gettimeofday(), rather than depending on Apple's own internal designs, e.g. Mach system calls. I don't believe Apple can rightfully claim APIs that aren't their own, as being their own implementation detail which they can change at will. We do our best to respect Apple's boundaries, so I believe the risk of breakage with Cosmo on Apple should be minimal.
Apple doesn't differentiate between "UNIX" and "Apple" when it comes to API; a particular API is either public with stability guarantees, or it is private and unstable.
> I don't believe Apple can rightfully claim APIs that aren't their own, as being their own implementation detail which they can change at will.
Apple absolutely does claim this, and if that means the library ABI has to change to accommodate a change, the dynamic linker and symbol tricks are leveraged to keep things working for code built against the earlier ABI.
If an engineer comes up with a really clever trick to make gettimeofday() just a tiny bit faster, but this requires breaking syscall ABI, they will absolutely do that.
> I believe the risk of breakage with Cosmo on Apple should be minimal.
Using system-private interfaces on Apple platforms means there are no guarantees here. You'll be OK, sometimes, for some releases. It mostly worked for Go, for a while.
It's not "two megacorps" disagreeing, it's "supported interface" vs "unsupported, unstable, system-private interface".
On paper. Have they acted on it? Have they actually broken an API that used to come standard with UNIX?
The underlying syscalls that support them have occasionally changed and broken existing apps that bypassed libSystem. For example, Sierra broke all go apps that called `gettimeofday` (the exact syscall jart used above an example!) because the go compiler emitted direct syscalls: https://github.com/golang/go/issues/16606
I'd like to know, did Apple actually break a "private" C ABI when this ABI actually implemented a standard UNIX function? I know they could, but did they?
By standard UNIX function I take it that you mean a function defined via POSIX and part of one of the various specifications used for UNIX certification (for the moment lets ignore the fact there are multiple revisions and optional extensions). It is important to note that the specifications says essentially nothing about:
* Binary formats
* Libraries (static or dynamic)[1]
* What symbols are in what library
It is all written in terms of what source code should compile, and how that compiled code functions. Everything else such as calling conventions, syscall interfaces, what is library code vs a syscall, etc is an implementation detail.
So given the above, I am not entirely sure what you mean by a `"private" C ABI when this ABI actually implemented a standard UNIX function.` Do you mean has Apple ever changed an internal function called by a function specified in POSIX? IOW, if your question is does Apple reserve the right to implement `stat()` as a call to `stat_internal()` and then change the arguments to "stat_internal()" ? Absolutely.
If you mean has Apple ever changed a function that is part of POSIX but it considers private? Those don't really exist on macOS, if POSIX allows it and it is part of the standard that has passed conformance it is by definition public and the C ABI level interfaces for as exposed by libSystem are stable (which is not to say that all of those interfaces are great, but they are standard and supported). IOW, the standard specifies that `stat()` exists, and it is by definition public.
That is not to say incompatible changes have never had to happen (for example, when UNIX conformance was originally implemented a lot of existing functions required incompatible changes to pass the test suites). All of that is handled via symbol versioning and redirecting new binaries to different symbols than the older binaries used, which maintains both binary compatibility for old binaries and allows new source to compile in the correct (conformant) way. This is why if you inspect libsystem_kernel.dylib you see variants of symbols like:
* _recvmsg
* _recvmsg$NOCANCEL$UNIX2003
* _recvmsg$UNIX2003
The old ones keep working with the existing semantics for older binaries, the headers have magic in them redirect to the newer ones when targeting the appropriate minimum OS version, and the userspace libraries have multiple entry points that provide both sets of semantics (often implemented in the userspace shim, sometimes by dispatching to the kernel with different syscalls).
[1]: Despite that, at this point POSIX does specify some of the semantic of `dlopen()` and `dlsym()`, which is pretty insane when you think about.
If you by "allowed" means "is meant to work", then yeah it means it's not allowed on many of these platforms.
OpenBSD has infrastructure to control which memory regions syscalls are allowed to stem from, specifically designed to block syscalls from anything by libc. See https://lwn.net/Articles/806776/
I do not recall if it was enabled, but when someone does not provide you an ABI it very much means that you are not supposed to try to write code against it. If you do it anyway, you will have to live with the resulting instability.
for all things bad about windows, one cannot say that they didn't take backward compatibility quite seriously
https://j00ru.vexillium.org/syscalls/nt/64/ (click show all)
People should study the difference between libc and glibc and history of it. We've been down this path before.