There's nothing severe about it. Most systems are exactly that: systems of which the kernel is only one part, syscalls are rarely if ever intended to be called directly nilly-willy.
The issue is that unlike windows unices have never enforced this.
There's nothing severe about it. Most systems are exactly that: systems of which the kernel is only one part, syscalls are rarely if ever intended to be called directly nilly-willy.
The issue is that unlike windows unices have never enforced this.
It makes sense for systems where libc is tightly coupled and coversioned with the kernel, e.g. BSDs, but Linux always relied on third-party C libraries and supported static binaries, etc.
You could argue that BSD made the mistake of intending to have a Windows-style C library compat guarantee but not enforcing it, but that was not in scope for Linux. The philosophy has always been syscall-level compat (and there are lots of famous threads with Linus re-enforcing this to others who would presume that things should be “fixed in user space”).
So it’s hardly reasonable to generalize based on some BSD concerns; Linux is WAI and represents the most common Unix-like system people use today by far.
There’s a pretty good argument that this level of compat, while the source of some problems, has also made other things much easier: consider container images that are bundled with their own system libraries. (You could certainly invent schemes to inject these libraries, but dealing with link and library level compatibility seems even more complex to deal with than system call-level compatibility.)
Isn't that the kernel default? Even if you use system calls directly, file descriptors still inherit by default.
To the extent that Go's default for file descriptors today is !inherit (I'm unfamiliar, but if so, it's a good choice), the Go runtime must already add O_CLOEXEC to bare syscalls. There's no reason to believe it incapable of adding the flag to libc syscalls instead.
You are thinking of the older way, where fcntl(fd, F_SETFD, FD_CLOEXEC) must be used after open(), leaving a short window in which the file descriptor may be inherited.
The newer way passes the O_CLOEXEC to open() and there is no fcntl() call. This is atomic with respect to inheritability: The kernel returns a non-inheritable file descriptor to libc, and libc returns it to the application.
Other syscalls that return a file descriptor have similar flags, so they are atomic too.
These flags and behaviours are exactly the same, whether done by calling through libc as most programs do, or direct kernel syscalls bypassing libc, as Go and a few other programs do.
This syscall level behavior is POSIX-specified[1] since at least the 2008 edition[2]:
> O_CLOEXEC > If set, the FD_CLOEXEC flag for the new file descriptor shall be set.
What that means is, any C program or Go program that passes the O_CLOEXEC flag to open(2) on a POSIX 2008 conforming system (including Linux and the BSDs, for example), will atomically create the fd without inherit behavior. There is no "short window" and hasn't been for more than a decade. The Go runtime must use that flag to provide that property; there is no other way on these systems. Libc users are of course able to use the same flag.
[1]: https://pubs.opengroup.org/onlinepubs/9699919799/functions/o...
[2]: https://pubs.opengroup.org/onlinepubs/9699919799.2008edition...
You're incorrect or maybe just misleading about libc created file descriptors inheriting, as stated. Either way, it is unrelated to using libc for syscalls vs bare machine traps.
libSystem is now used when making syscalls on Darwin, ensuring forward-compatibility with future versions of macOS and iOS.
> solaris/macos approach … of calling libc stubs
OpenBSD does have somewhat different constraints and they seem to think this will work for them.
What containers are you referring to? Because this is definitely not how Zones work on Solaris.
There are two types of Zones in Solaris 11; "Kernel Zones" which run their own independent version of Solaris and "non-global Zones" which are automatically kept at the same version as the host.
Windows has to virtualise containers with an incompatible OS version.
Not as far as I'm aware. Windows Sandboxes don't work that way nor do other technologies I'm aware of. What are you referring to?
Windows containers have to run under hyper-v for incompatible kernel versions.
Linux is not a system. Linux is a kernel, linux has distributions, it is not a single coherent system where the kernel and standard library are co-developed. That's the entire point I'm making.
> So it’s hardly reasonable to generalize based on some BSD concerns
It's not "some BSD concerns", it's pretty much every non-linux unix. What's not reasonable is generalising "syscalls are a perfectly fine interface" which is almost exclusively a Linux exclusivity. Or don't claim compatibility with anything other than linux, that's also a perfectly fine choice.
> There’s a pretty good argument that this level of compat, while the source of some problems, has also made other things much easier: consider container images that are bundled with their own system libraries.
Last time I checked, Go did not run exclusively on linux. If it did, raw syscalls would indeed not be a concern (though even then they try to have their cake and eat it, as e.g. they want to do raw syscalls yet benefit from vDSO, which has been an issue in the past because their assumptions did not hold: https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/)
It does seem that there's a very reasonable security concern about doing so from +w+x memory, or from non-PIE regions, or without first explicitly authorizing the calling range.
My question still stands regarding what the security concern of opt-in PIE -w+x code making direct syscalls is.
Edit: (Of course I do understand that the BSDs (unlike Linux) do not guarantee a stable syscall ABI and as such performing them directly is strictly a bad practice.)
Still, I don’t think there’s anything wrong with letting an application mark part of its code safe for syscall execution, versus enforcing libc only. Seems like the exact same thing as the execute bit. Moreover, some systems genuinely have a stable syscall ABI - I think Linux would be considered one.
Windows pretty much enforces it in the sense that syscall numbers can change between minor updates, so raw syscalls breaks extremely often.
> Moreover, some systems genuinely have a stable syscall ABI - I think Linux would be considered one.
Linux is not a system. It's a kernel, with userlands you can bolt on. That's why it has a stable syscall ABI: that's the only interface Linux can have if it intends to provide an interface.
That's not true; Solaris and I suspect other *nixes of the past made it explicitly clear that libc was the interface for userland, not the kernel.
I don't think this is true, at least on Linux.
On Linux, the system call interface¹ is the documented interface to user space. Even the commonly used vDSO² is a stable interface. This is important because it means the popular C libraries are not part of the Linux kernel interface. Although glibc is often portrayed³ as some kind of Linux kernel wrapper, they are entirely separate projects. Linux manuals⁴ also make it seem like they are one and the same:
> The Linux man-pages project documents the Linux kernel and C library interfaces that are employed by user-space programs.
These same manuals also document systemd as if it was part of Linux. I went there expecting low level documentation useful for writing one's own init system and got systemd documentation instead. It's very confusing in my opinion. Why are external projects documented in the Linux manuals?
Anyway, these kernel features are used by C libraries to implement all their functions. Using C libraries is the traditional way to build a Linux user space but it is certainly not the only way. Compilers could emit these system calls directly, avoiding the need for a runtime library. A programming language virtual machine could be built directly on top of Linux system calls. It is possible to create freestanding programs that run on Linux with zero dependencies.
Incompatibilities are caused by these user space libraries, not by the system calls themselves. For example, glibc maintains a lot of thread local state and will not work correctly if the program calls clone(). A program that does not link to glibc does not have this limitation.
Although low level, Linux system calls are in many ways a simpler interface: their behavior is more precisely documented compared to POSIX; there is no need to deal with errno; there is no hidden C library functionality that's hard to understand; freestanding programs do not contain references to hundreds of hidden standard library symbols that implement obscure functionality.
The kernel itself containd a nolibc.h header[5] that it apparently uses for its own tools.
[1]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[2]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[3]: https://en.wikipedia.org/wiki/File:Linux_kernel_System_Call_...
[4]: https://www.kernel.org/doc/man-pages/
[5]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
My assertion pertains to systems. Linux is not a system.
> Even the commonly used vDSO² is a stable interface.
https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/
> [snip completely irrelevant everything else]