The Go runtime scheduler's way of dealing with system calls
utcc.utoronto.ca
utcc.utoronto.ca
For dynamic binaries, we continue to to permit the main program exec segment because "go" (and potentially a few other applications) have embedded system calls in the main program. Hopefully at least go gets fixed soon.
We declare the concept of embedded syscalls a bad idea for numerous reasons, as we notice the ecosystem has many of static-syscall-in-base-binary which are dynamically linked against libraries which in turn use libc, which contains another set of syscall stubs. We've been concerned about adding even one additional syscall entry point... but go's approach tends to double the entry-point attack surface.
https://marc.info/?l=openbsd-tech&m=157488907117170&w=2
[edit for convenience of readers - read the above linked thread - I just grabbed the go part]
Unfortunately our current go build model hasn't followed solaris/macos approach yet of calling libc stubs, and uses the inappropriate "embed system calls directly" method, so for now we'll need to authorize the main program text as well. A comment in exec_elf.c explains this.
If go is adapted to call library-based system call stubs on OpenBSD as well, this problem will go away. There may be other environments creating raw system calls. I guess we'll need to find them as time goes by, and hope in time we can repair those also.
[/edit]
The idea is to extend this protection to only allow system calls from expected address ranges, so that a successful exploit can't simply make raw calls but instead has to track down an existing authorized one (and thus contend with ASLR). To that end, the new call-once syscall msyscall(2) is added. The linker uses it to register libc.so with the kernel after randomly mapping it into the current process.
Yes.
> (And perhaps ban versions of libc that have bugs?)
There is only one OpenBSD libc.
> We've been concerned about adding even one additional syscall entry point
I don't understand the need for such a severe "only libc syscalls ever" approach.
What would be the security concern with allowing syscalls only from preauthorized (ie msyscall(2)) regions, making initial region authorization opt-in (instead of opt-out), allowing the program to call msyscall(2) itself, and rejecting any statically linked (ie non-ASLR'd) regions for authorization?
There's nothing severe about it. Most systems are exactly that: systems of which the kernel is only one part, syscalls are rarely if ever intended to be called directly nilly-willy.
The issue is that unlike windows unices have never enforced this.
It makes sense for systems where libc is tightly coupled and coversioned with the kernel, e.g. BSDs, but Linux always relied on third-party C libraries and supported static binaries, etc.
You could argue that BSD made the mistake of intending to have a Windows-style C library compat guarantee but not enforcing it, but that was not in scope for Linux. The philosophy has always been syscall-level compat (and there are lots of famous threads with Linus re-enforcing this to others who would presume that things should be “fixed in user space”).
So it’s hardly reasonable to generalize based on some BSD concerns; Linux is WAI and represents the most common Unix-like system people use today by far.
There’s a pretty good argument that this level of compat, while the source of some problems, has also made other things much easier: consider container images that are bundled with their own system libraries. (You could certainly invent schemes to inject these libraries, but dealing with link and library level compatibility seems even more complex to deal with than system call-level compatibility.)
Isn't that the kernel default? Even if you use system calls directly, file descriptors still inherit by default.
To the extent that Go's default for file descriptors today is !inherit (I'm unfamiliar, but if so, it's a good choice), the Go runtime must already add O_CLOEXEC to bare syscalls. There's no reason to believe it incapable of adding the flag to libc syscalls instead.
You are thinking of the older way, where fcntl(fd, F_SETFD, FD_CLOEXEC) must be used after open(), leaving a short window in which the file descriptor may be inherited.
The newer way passes the O_CLOEXEC to open() and there is no fcntl() call. This is atomic with respect to inheritability: The kernel returns a non-inheritable file descriptor to libc, and libc returns it to the application.
Other syscalls that return a file descriptor have similar flags, so they are atomic too.
These flags and behaviours are exactly the same, whether done by calling through libc as most programs do, or direct kernel syscalls bypassing libc, as Go and a few other programs do.
This syscall level behavior is POSIX-specified[1] since at least the 2008 edition[2]:
> O_CLOEXEC > If set, the FD_CLOEXEC flag for the new file descriptor shall be set.
What that means is, any C program or Go program that passes the O_CLOEXEC flag to open(2) on a POSIX 2008 conforming system (including Linux and the BSDs, for example), will atomically create the fd without inherit behavior. There is no "short window" and hasn't been for more than a decade. The Go runtime must use that flag to provide that property; there is no other way on these systems. Libc users are of course able to use the same flag.
[1]: https://pubs.opengroup.org/onlinepubs/9699919799/functions/o...
[2]: https://pubs.opengroup.org/onlinepubs/9699919799.2008edition...
You're incorrect or maybe just misleading about libc created file descriptors inheriting, as stated. Either way, it is unrelated to using libc for syscalls vs bare machine traps.
libSystem is now used when making syscalls on Darwin, ensuring forward-compatibility with future versions of macOS and iOS.
> solaris/macos approach … of calling libc stubs
OpenBSD does have somewhat different constraints and they seem to think this will work for them.
What containers are you referring to? Because this is definitely not how Zones work on Solaris.
There are two types of Zones in Solaris 11; "Kernel Zones" which run their own independent version of Solaris and "non-global Zones" which are automatically kept at the same version as the host.
Windows has to virtualise containers with an incompatible OS version.
Not as far as I'm aware. Windows Sandboxes don't work that way nor do other technologies I'm aware of. What are you referring to?
Windows containers have to run under hyper-v for incompatible kernel versions.
Linux is not a system. Linux is a kernel, linux has distributions, it is not a single coherent system where the kernel and standard library are co-developed. That's the entire point I'm making.
> So it’s hardly reasonable to generalize based on some BSD concerns
It's not "some BSD concerns", it's pretty much every non-linux unix. What's not reasonable is generalising "syscalls are a perfectly fine interface" which is almost exclusively a Linux exclusivity. Or don't claim compatibility with anything other than linux, that's also a perfectly fine choice.
> There’s a pretty good argument that this level of compat, while the source of some problems, has also made other things much easier: consider container images that are bundled with their own system libraries.
Last time I checked, Go did not run exclusively on linux. If it did, raw syscalls would indeed not be a concern (though even then they try to have their cake and eat it, as e.g. they want to do raw syscalls yet benefit from vDSO, which has been an issue in the past because their assumptions did not hold: https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/)
Still, I don’t think there’s anything wrong with letting an application mark part of its code safe for syscall execution, versus enforcing libc only. Seems like the exact same thing as the execute bit. Moreover, some systems genuinely have a stable syscall ABI - I think Linux would be considered one.
Windows pretty much enforces it in the sense that syscall numbers can change between minor updates, so raw syscalls breaks extremely often.
> Moreover, some systems genuinely have a stable syscall ABI - I think Linux would be considered one.
Linux is not a system. It's a kernel, with userlands you can bolt on. That's why it has a stable syscall ABI: that's the only interface Linux can have if it intends to provide an interface.
It does seem that there's a very reasonable security concern about doing so from +w+x memory, or from non-PIE regions, or without first explicitly authorizing the calling range.
My question still stands regarding what the security concern of opt-in PIE -w+x code making direct syscalls is.
Edit: (Of course I do understand that the BSDs (unlike Linux) do not guarantee a stable syscall ABI and as such performing them directly is strictly a bad practice.)
That's not true; Solaris and I suspect other *nixes of the past made it explicitly clear that libc was the interface for userland, not the kernel.
I don't think this is true, at least on Linux.
On Linux, the system call interface¹ is the documented interface to user space. Even the commonly used vDSO² is a stable interface. This is important because it means the popular C libraries are not part of the Linux kernel interface. Although glibc is often portrayed³ as some kind of Linux kernel wrapper, they are entirely separate projects. Linux manuals⁴ also make it seem like they are one and the same:
> The Linux man-pages project documents the Linux kernel and C library interfaces that are employed by user-space programs.
These same manuals also document systemd as if it was part of Linux. I went there expecting low level documentation useful for writing one's own init system and got systemd documentation instead. It's very confusing in my opinion. Why are external projects documented in the Linux manuals?
Anyway, these kernel features are used by C libraries to implement all their functions. Using C libraries is the traditional way to build a Linux user space but it is certainly not the only way. Compilers could emit these system calls directly, avoiding the need for a runtime library. A programming language virtual machine could be built directly on top of Linux system calls. It is possible to create freestanding programs that run on Linux with zero dependencies.
Incompatibilities are caused by these user space libraries, not by the system calls themselves. For example, glibc maintains a lot of thread local state and will not work correctly if the program calls clone(). A program that does not link to glibc does not have this limitation.
Although low level, Linux system calls are in many ways a simpler interface: their behavior is more precisely documented compared to POSIX; there is no need to deal with errno; there is no hidden C library functionality that's hard to understand; freestanding programs do not contain references to hundreds of hidden standard library symbols that implement obscure functionality.
The kernel itself containd a nolibc.h header[5] that it apparently uses for its own tools.
[1]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[2]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[3]: https://en.wikipedia.org/wiki/File:Linux_kernel_System_Call_...
[4]: https://www.kernel.org/doc/man-pages/
[5]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
My assertion pertains to systems. Linux is not a system.
> Even the commonly used vDSO² is a stable interface.
https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/
> [snip completely irrelevant everything else]
Who preauthorizes them, and what prevents more from getting added?
Seeing as msyscall() is itself a system call, any calls to it must themselves originate from an authorized region.
I'm aware of the issues. I was a part of the discussions in the hut when they were being implemented.
I don't see how that's the case; my original question was essentially asking what that security hole is. Provided that syscall ability is opt-out by default and that only code subjected to ASLR is permitted to be authorized, it doesn't seem terribly risky to allow additional such regions to be registered. An exploit has to contend with ASLR either way; either by locating libc, or by locating some other authorized region within the current process.
AFAIU, the entire security benefit here is due to ASLR alone. If an exploit manages to track down libc, it can go right ahead and make all the system calls it wants. (Unless there's some other piece to the puzzle that I've missed? Is there something special about libc in particular?) As such, I still don't understand how the called-once restriction is supposed to meaningfully increase security - by the time you've found the msyscall() function, you've also found _all the others_ anyway.
It has to create the appropriate gadgets to generate function call sequences, and generating gadgets is hard.
Go, as of version 1.12,uses libSystem in Darwin to make syscalls.
Perhaps it is just me, but it seems all this user space rigamarole to map bits of execution onto cores points to an overall architecture “smell”. This should be performed and enabled by the OS.
You can see the seams between the OS and the go runtime tear a little whenever a library acquires an ownership lock where the thread id is recorded. In Go, computation moves freely between threads, so that lock doesn’t work (at least without special instructions to the runtime to lock that goroutine to a thread).
The whole POSIX threading model seems broken in this context.
Oof that’s horrible, any pointers to the logic behind it? I’m curious the rationale.
https://golang.org/doc/faq#Do_Go_programs_link_with_Cpp_prog...
The idea that Go's scheduler design is somehow inherited from Plan 9 is ridiculous.
1. stackless (async/await) where the operation becomes an inspectable object that you can choose what to do with (awaiting being suspend for completion) as taken by C#, C++, Python, JS, PHP, Swift and Rust
2. "with stack" where you pretend its not async; but this means when something else uses the thread you need to get the suspended operation's stuff off the thread; usually by not using the thread's stack at all and having it in the heap and just jumping into and out of these "off-thread" stacks; as used by Go and being looked at for Java (as Project Loom)
Interesting paper on it http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p136...
> While fibers may have looked like an attractive approach to write scalable concurrent code in the 90s, the experience of using fibers, the advances in operating systems, hardware and compiler technology (stackless coroutines), made them no longer a recommended facility.
Disadvantage of stackless is the extra boilerplate (e.g. async/await everywhere); though it also gives more control as the consumer of the operations (e.g. fanout and wait for many; or continue not waiting for the result at all)
Advantage of the "with stack" approach is it looks the same as non async code as its all hidden (goroutines aside); which is why Java is no doubt looking at doing it as there is a large body of code that would need to be rewritten so "hiding it" is easier to avoid that.
C# had/has teething issues when async/await was introduced as it kept the initial thread blocking methods; and added the async and they don't mix very well, you need really to go one way or the other when developing.
Javascript leapt at async/await as it was all async anyway, but callback based which makes for horrible code to follow; so it made everything much cleaner.
LLVM IR has async/await and coroutines, but most real-world VMs and language static compilers cannot depend on such intrinsics because of their memory and execution models. For example, Pony's ORCA has unique memory barrier and execution models that wouldn't work with this approach, although it uses LLVM for compilation down to metal. This is why LLVM is a loose framework and collection of tools split into "middleware" passes, rather than a single monolith.
PS: According to its paper, ORCA is supposedly one of the fastest GCs for most use-cases. It beat Zulu's C4, Erlang BEAM and another one in a deathmatch. It's too bad it can't be extracted as a separate project or integrated into OpenJDK or LLVM without lots of work. Of course, no GC is better (I'm staring at you, Rust. :).
0. next unit of work (pointer/counter; instruction, function pointer, etc.)
1. task-local heap (pointer or structure)
2. operand stack (pointer or structure; for stack-oriented VMs only)
The problem there is that you need to very carefully size your stack as a mis-sizing will lead to a risky stack overflow. I'm not sure it's necessary either as allocating a "large stack" but using very little of it means most of it is never committed, and thus only costs memory mappings.
A million threads is an extreme case, though. No system can reliably spawn that many threads that are actually doing something interesting without a very large amount of memory. When you leave the realm of microbenchmarks you have to expect that an unknown quantity of threads will have deep call stacks at any given time, so you really need to give yourself leeway to avoid the risk of OOM.
So I don't think it's fair to say POSIX threads are comparable or whatever if they don't have this property.
It depends on domain, but in my domain of high-RPS network servers, they absolutely do.
> 10M goroutines would mean 20GB just for stacks.
When I deploy to metal, the average host has ~512GB of RAM, and not much cotenancy.
> The 2kB minimum stack size is on the same order of magnitude as the 10kB POSIX thread stack size.
There's also a question of cost to create and destroy; it's very common for goroutines to live for O(µs). I don't know how POSIX threads compare here.
The vast majority of the "expense" is irrelevant as it's virtual memory and unlikely to ever be touched (and thus committed).
For example, try setting 'ulimit -s 128' (128kB stack limit) and see how many C programs crash. Then try, say, 16. Go's default is 8 kB, raised from 4 kB in 1.2: https://golang.org/doc/go1.2#stack_size
Linux's default userspace stack limit is 8 megabytes for a reason — programs really do use it.
Not really, the 8MB limit was added back in '95 from a previous limit of "essentially none"[0] with a justification of
> Limit the stack by to some sane default: root can always increase this limit if needed.. 8MB seems reasonable.
Developers don't generally think about their stack size, especially for single-threaded programs[1] so the defaults need to be a sweet spot of not unnecessarily big (such that you can catch unbounded recursion) but not so small that you'd segfault more than a very small fraction of all programs.
[0] https://git.kernel.org/pub/scm/linux/kernel/git/history/hist...
[1] which would be why e.g. OSX has a large main thread stack (8MB) and a relatively puny secondary thread stack (512k).
Unfortunately the work seems to have stalled out and never made it into the kernel. If that work actually makes it into the Linux kernel, then other languages like C++ and Rust that have more stringent runtime requirements could make uses of lightweight threading as well.
[0]: https://docs.microsoft.com/en-us/windows/win32/procthread/us...
I'm not sure why this leads you to the conclusion that "the whole POSIX threading model seems broken."
If you had lots of goroutines do lots of things that stalled on hung NFS mounts, you would build up a lot of OS threads (all sitting in system calls) and might run into limits there. But that's inevitable in any synchronous system call that can stall.
(I'm the author of the linked-to article.)
¹ It's possible the JVM, being the highly optimized workhorse VM it is, has specialized optimizations for I/O and does indeed skip over JNI and libc in these cases.
AFAIK Go does syscalls itself on any platform but Windows and macOS, this includes all BSDs. And even for macOS despite that having never been officially supported it took multiple breakages a few years back.
The first thread here mentions the issues that causes for openbsd.
There's been some discussion about using libuv but no consensus: https://gitlab.haskell.org/ghc/ghc/issues/8400. It's available as a library: https://haskell-stdio.github.io/stdio/
I didn't re-read the papers but IIRC GHC just spawns an extra OS thread any time a possibly-blocking function is called, as it doesn't follow a strict M:N model. There's a thread pool to reduce overhead but it's probably not as efficient as Go's method.
2. If M blocks in a syscall too long in the optimistic case:
2.a. is M unpinned from P but continues to block until the syscall returns?
2.b. is another thread from the pool used or new thread created, and pinned to P so that P can be used for other work? (I think this depends on configuration if there are fewer, same or more threads than processors.)
2.c. is there an upper limit on outstanding blocked syscall worker threads or will it simply be the last task any extra created threads beyond the normal limit would ever process?
(I believe the actual implementation treats Ms as a sort of secondary thing. For instance, I think that the local list of runnable goroutines is attached to the P, not to the M. At one level, the M is just a context for running things on Ps.)
In the optimistic case when the system call blocks for too long, the M is unpinned from the P it was using and continues to sit in the system call (the Go runtime doesn't attempt to interrupt the system call itself). If there is another runnable goroutine and there are no free M's, the Go scheduler will create another M to run the goroutine on the now-free P. I think that the runtime directly allocates the free P to the newly created M rather than letting the new M try to contend with other things for the P, but I'm not sure.
I don't think there's any limit on the number of Ms (OS threads) that the Go runtime will create, but I haven't checked the code carefully. Idle Ms are reclaimed under some circumstances.
(I'm the author of the linked-to article.)
Is there a way to check from the running process whether that is the case or not?
dump_vdso [1] will write the vdso to stdout, you can use binutils like objdump or nm to list the symbols present.
[1] https://kernel.googlesource.com/pub/scm/linux/kernel/git/lut...