Go 1.16 will make system calls through Libc on OpenBSD
utcc.utoronto.ca
utcc.utoronto.ca
Does this mean that the Go developers managed to work around this problem or that it was just a flimsy post-hoc justification for chronic NIH syndrome?
TFA itself links to another blog post discussing "Some reasons for Go to not make system calls through the standard C library"[1] but as far as I can tell it doesn't explain why it suddenly stopped being a problem on OpenBSD.
[1] https://utcc.utoronto.ca/~cks/space/blog/programming/GoCLibr...
OpenBSD is deciding to invent their own paintbrush without even looking at what other people are painting with, while Go took the closest bucket of paint and threw it on the floor, not realizing they were standing in a corner while doing so.
And I say that as a happy user of both.
No, there is almost certainly a bunch of headache having to do system calls through libc for Go. It didn't stop being a problem, there just isn't another option in OpenBSD's case. UNIXs are tricky for anyone who doesn't want to use libc since they typically define that as the official interface to the system. Linux is the outlier here by keeping its syscall ABI very stable.
After all as far I can tell it's not just an OpenBSD problem since famously they got breakage in MacOS as well.
I must admit that I haven't taken the time to analyze in depth the pros and cons here, but Go's history of NIH coupled with the fact that basically every other mainstream language manages to work by binding the libc leaves me very perplex.
In particular some of the points raised in the blogpost I linked seem fishy to me. For instance errno being a global: this is in no way a Go-specific problem (multithreaded C couldn't run concurrent syscalls if it was true). In practice errno is thread-local instead of being a true global. It's explicitly documented in the man page:
> errno is defined by the ISO C standard to be a modifiable lvalue of type int, and must not be explicitly declared; errno may be a macro. errno is thread-local; setting it in one thread does not affect its value in any other thread.
I can believe that there are other issues I fail to consider, but again it works for everybody else, what makes Go so special here?
Error doesn't work for everyone else. It's a hack, you wouldn't do it if it wasn't for backwards compatability.
What is the advantage in that? including that into a binary seems about as senseless to me as including the entire kernel in it.
Another example is Windows, the platform API does not provide a C library (even MSVC has its own). While there is an MSVCRT.DLL it is not recommended to link against it as it is there only because some other software relies on it and its semantics are for around Visual C++ 4 (IIRC).
https://devblogs.microsoft.com/cppblog/introducing-the-unive...
It is already next to impossible to write software that requires “Linux” and nothing more with all the kernel functionality that can be enabled or disabled.
Linux is a component of many different platforms, which indeed do provide different libcs, but also different t.l.s. libraries, different c.p.u. architectures, different Linux configurations and whatever else.
As far as I know with respect to Windows, it only provides stable interfaces viā C libraries, and does not have a stable binary interface to the kernel directly.
Because they really don’t want to, so they will avoid doing it until forced to, as previously happened with e.g. macOS.
> I must admit that I haven't taken the time to analyze in depth the pros and cons here, but Go's history of NIH coupled with the fact that basically every other mainstream language manages to work by binding the libc leaves me very perplex.
One of the issues Go has is it uses its own non-C stack. Libcs generally assume a C stack and don’t understand growable movable stacks so they can get very cross when called with an unexpectedly small stack (iirc Go defaults to 2k while the smallest C stack I know of is 64k on old OpenBSD, then macOS’s 512k for the non-main threads).
You're correct (AFAIK anyway) that stack pages are mapped lazily. That doesn't change the fact that C stacks are large allocations, while Go will only allocate a very small stack (2k last I checked) per goroutine.
Obviously the folks that created Go aren't stupid, far from it, so there must be a real and valid reason. But can't easily imagine what it is.
When you have 100,000 threads, big stacks add up.
Given how Go has no issue being prescriptive as hell on other things, I don't see why they couldn't just go "set vm. overcommit_memory to 1 and fuck off", that's exactly what e.g. Redis tells you.
And again it would hardly be the first time Go makes OS-specific decisions.
If you have 8MiB stacks, then a minimally allocated stack uses a 4KiB data page, but also 2MiB of address range uses up a full 4KiB bottom level page table, and 8MiB range takes up 4/512 x 4KB of 2nd level page table and so on. So you use about 8.03 KiB RAM if you never touch more than 4KB and your 8MB reservations are mostly grouped. Some architectures have bigger pages, increasing fanout but also the minimum allocation.
Contrast to 2KiB stacks without reservations/overcommit - you use about 2KiB of usable RAM + (1/2) x (1/512) of 4KiB 1st level page table + .. , assuming allocation are again mostly grouped. Hence, for up to 2KiB of stack memory you need about¹ 2.005 KiB of RAM. Works the same for 16 KiB and even 64 KiB page sizes.
100000 * (8MiB reserved, 1..4KiB used stack) needs ~784MiB RAM.
100000 * (2KiB reserved, 1..2KiB used stack) needs ~200MiB RAM.
Note that, if you actually touch your reserved stack, even once, your allocation can balloon to possibly tens or hundreds of GiB (100k * 8MiB = ~800GiB), unless you do complicated cleanup, while a segmented stack can keep the allocations within reasonable efficiency, freeing any excessive stack allocations in userspace.
¹ ignoring bookkeeping overhead in both cases, to keep the calculation clear. Hopefully it isn't more than a dozen bytes or so.
Yeah so nothing really relevant, that's 1/20th of a relatively basic dev laptop's memory for 100k threads.
And of course that's an insane worst-case scenario of 8MB stacks, which the linux devs picked because they wanted to add a limit to the stack but didn't really care for having one, Windows uses 1MB stacks and macOS uses 512k off-main stacks, so you don't need anywhere near 8MB to get C-compatible stack sizes.
With 8M/4M/2M/1M/512k stack size, 4K page size, you get about 800M/800M/800M/600M/500M RAM usage, of which only 400M is usable, rest is overhead. At 16K page size, it's ~1600M. Compared to 200M if using 2K side-by-side stacks, in any configuration.
Yes, probably not too excessive, even though it is noticeable when you spawn theads for anything and everything, and run more than a single application on a non-SV-developer PC.
I think main problems start when you actually touch more than the base allocation of the stack (or just use 16K pages). Maybe segmented stacks grow from 200M to 300M, 500M or whatever you actually use at a given time (with say 20-60% efficiency), but your C-compatible stacks might go from 500M to 3G¹ if you on average touch just 32K of stack per thread with some unlucky function, even though at any given time only the same ~100M of memory actually stores useful information.
¹ or more, no idea how high it typically goes, but that in itself is a nasty gotcha, and a likely reason you won't find "lightweight" threading in combination with native per-thread stacks
Simulating this (by creating 100k maps of the relevant sizes) there is no difference in RES between 8M pages and 512K pages: it was ~495M for both, only the VMEM varied (respectively ~780G and ~50G). Touching the second page increased the RES of both to ~816M, which is about what you'd expect.
This is on a more or less stock x64 Mint.
> I think main problems start when you actually touch more than the base allocation of the stack
The thing is you're unlikely to do that in all of your 100k routines, most of them will not grow beyond their first page, and maybe their second… at which point the routine's stack would have grown to 8k anyway.
> maybe segmented stacks grow from 200M to 300M, 500M or whatever you actually use at a given time (with say 20-60% efficiency)
Go has not used segmented stacks since 1.3.
> your C-compatible stacks might go from 500M to 3G¹ if you on average touch just 32K of stack per thread with some unlucky function
Your not-C-compatible stacks will do the exact same thing. Since 1.3, stacks are realloc'd and double in size on every overflow.
The Go runtime will also shrink stacks if able (halving them) during a GC run, but you can do essentially the same thing on your C stack using madvise(2), and without the need to copy stack data around.
So what gain there is, is really only for goroutines which never grow beyond their initial size (and only since the default stack was decreased from 8k to 2k in 1.4), at the cost of all the C incompatiblity mess. And it assumes these 2k stacks are allocated from a reusable pool (which is probably a fair assumption though I certainly did not check it) otherwise they'd be on different pages anyway.
Linux memory management has notoriously complicated reporting. Your program only has 495M of usable memory mapped (RSS; not sure where the extra 100M is coming from), but RSS does not count page tables.
You cannot actually use the 400M of sparsely allocated memory (4K at 2+MB intervals) without another 400M of page tables. I'd suggest you try allocating and using >50% of RAM, or just compare how much can you use sparsely vs. dense before your program hits OOM. Note, you may need to enable overcommit, increase maximum map count, if you are mapping region seperately and preferably disable swap to avoid thrashing.
sysctl vm.overcommit_memory=1
sysctl vm.max_map_count=10000000
swapoff -a
You can also watch /proc/memory PageTables total, which should show the difference.That being said, I don't use go, I was merely pointing out that virtual memory management is not magical, and it has real memory costs when doing sparse allocations (also has quite significant other costs).
> The Go runtime will also shrink stacks if able (halving them) during a GC run, but you can do essentially the same thing on your C stack using madvise(2), and without the need to copy stack data around.
While you could do that, I believe it requires a syscall per-stack, which might be more expensive than a bit of copying, and I'm not completely sure how easy it would be to determine whether the memory has in fact been allocated and needs freeing.
Also the default pthread stack size is (usually) 2MB which is indeed non negligible if you have a ton of threads, but with pthread_attr_setstacksize you can lower it. I can't seem to find the actual minimal size at the moment but I have a vague memory that you can reduce it to 16kB portably.
I'm currently working on a multithreaded Rust program with a bunch of threads and strong memory constraints and I use the runtime to reduce the stack to 64kB, it seems to work just fine.
And again, given that the language is garbage collected I can't help but find it a bit amusing that they're being stingy with a few MB of VMEM per thread. I guess that frees virtual memory for use in the balasts!
It matters for goroutines because that means you can't straight call into foreign code from a goroutine's stack, which is why cgo and friends are so problematic.
> Also the default pthread stack size is (usually) 2MB which is indeed non negligible if you have a ton of threads
That's not relevant because it's not vmem is the point. The actual resident size of a 2MB stack is 4k unless you start dirtying more pages. Unless you're running on 32b, VMEM doesn't really matter.
> I can believe that there are other issues I fail to consider, but again it works for everybody else, what makes Go so special here?
errno being thread-local doesn't really help all that much with a M:N threading model -- the runtime is going to have to be extremely careful to not stomp all over it. (Whenever calling into anything which uses errno. Of course system calls do that as well, but it's a much smaller surface area than "most of libc".)
The only counterargument I can think off would be to remove the libc footprint for ultra-small, go-only distributions but I highly doubt that it's significant on any modern system (even an embedded one) and if you care so much about reducing the runtime footprint why would you go with a GC language in the first place?
That's actually exactly the use case for a lot of containerized applications written in Go.
edit: I've worked on multiple applications like this and there's the odd service where you need some other dependency, and then that service has to start running on Alpine Linux or something. But if the majority of your many "microservices" are pulling in bits of userland from a distro, they stop being "micro" very quickly which is a big deal when there's a lot of them. It's far preferable in that architecture if the entirety of your userland is your own binary.
From their website:[0] >"Alpine Linux is built around musl libc and busybox"
I could be off-base here, but it seems to me the better analogy would be some sort of "Go-Linux" that's like a Docker linux OS written entirely in Go.
And ironically it was the BSDs and Solarises of the world that had more fully developed similar concepts long before they were mainstream on Linux. I'm not terribly surprised by OpenBSD's stance here, though, and even though I'm more in the Go / Linux ecosystem I can't really argue - it's an unfortunate collision of very different philosophies.
"The eventual goal would be to disallow system calls from anywhere but the region mapped for libc"
I've always considered claims like this to be bullshit. I've worked a lot with both libc level system calls and (for crash report generation) direct system calls bypassing libc. They're the same interface. There's nothing you can do with a raw system call interface that you can't do through raw libc --- with rare exceptions for things like libc not having system call wrappers for some new system call or legacy ill-advised emulation like in fallocate. None of these exceptions is a justification for bypassing libc for calling open(2) or write(2).
As far as I can tell, the real reason Go eschewed libc system call interfaces is that the Go developers want to think of themselves as "not C".
The bullshit lies in these FUDlike insinuations that using libc would limit Go in some way. These insinuations are never backed up with technical specifics. I don't care that famous names are involved with Go: the presence of these people doesn't make Go's behavior correct or necessary.
There is zero technical case for Go doing what it does on Linux. You can make a "fully static binary" (which is a terrible idea anyway) with libc. Nobody should be making static binaries.
https://github.com/golang/go/issues/16606
So now they're gradually backtracking on all platforms - macOS switched to libc a while ago.
Having a common entrypoint into the kernel we can hook into is fairly valuable IMO, especially when said interface is for the most part stable between all unx (and to a somewhat lesser extent even on Windows).
Having this lowest common denominator ABI makes interop relatively* painless. But clearly Go wants to eat the world and doesn't seem to care very deeply about this. I suppose it's a valid approach, even if I don't necessarily agree with it.
If it ain't broke...
user supplied read addressesgetpwent() - actually really the whole notion of users
file permissions
aynch storage event completion (all asynchrony is pretty broken)
system metadata control (sysctl, /proc, interface socket messages...)
signals
ttys
errno
affinity interfaces ...
getpwent(), the whole notion of users - we dont use computers the same way we used to in the 70s. talk, write, finger, wall - they aren't very fun anymore since its either just me on my laptop, or one of the 100s of virtual machines floating around. more importantly, the attempts to glue unix system user identity to distributed identities (PAM) have really turned out to be a mess
filesystem permissions - these are clearly insufficient given the number of system-specific addons here.
signals are so riddled with constraints and incompatibilities that they are basically useless - except you have to fiddle with them for things like SIGPIPE
ttys were already kind of broken when they were relevant,
errno is actually a property of the libc, but the status interface is pretty broken - have you even grepped the kernel code to find out what might be issuing an EINVAL?
why dont you mail me at yuri tenuki org? i love these kinds of chats
It doesn't. It's just that as previously with macOS, Illumos, or Windows, the Go project ended up with its back against the wall: in this case, ultimately only the OpenBSD libc will be allowed to make syscalls so their choices are "use libc" and "no Go on OpenBSD".
And in all honesty the followup https://utcc.utoronto.ca/~cks/space/blog/unix/CLibraryAPIReq... is much more convincing as to why it's a hassle to go through libc.
https://marc.info/?l=openbsd-ports-cvs&m=158083696719245&w=2
Big Ouch. I wonder how the Linux developers will approach this bug since they don't enforce syscalls to be done from glibc.
I don't think it'll be a bit problem, anyway. In my experience not very much calls syscalls directly. Go is a big exception, though...
What about musl? uClibc?
Linux is well known for the fact that it guarantees its syscall interface as primary contract to userspace.
In either case though there is a libc that most software will use, and the mitigations can be applied there. Even though direct use of syscalls is legal on Linux, the fact that it's stable is primarily relevant and interesting to said libc developers.
The fact that syscalls aren't guaranteed on other systems is usually of little consequence since the libc is developed in tandem with the kernels of those systems. Linux's situation as a fully decoupled kernel means it does things differently in that sense. The developers are fully separated, so there needs to be a strong "contract" that syscalls will be stable.
Doesn't mean (IMO) it's a good idea for end-users e.g. software developers to use syscalls except in exceptional circumstances. Which is usually the case!
If someone successfully injects code that perform a direct syscall they can successfully use this info leak despite a safe and patched (g)libc.
As Raymond Chen wrote, that's the other side of the airtight hatch (https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31...).
If someone is capable of directly running their own code that ignores libc (or other existing mitigations, such as may exist in javascript runtimes / go compiler) then cool, they can use spectre to perform a timing attack against their own code that they're running. Or they could just read their own memory.
Spectre's main risk was for reading other program's memory or for doing so remotely with javascript. If you can already make the process you're attacking run arbitrary syscalls instead of use glibc, then you've already won and no amount of protection will help.
You're basically arguing that in a post-spectre world, native processes can fundamentally never be a security boundary again, right?
I'm wondering if this is necessarily true. For the concrete example at hand, Linux could offer some opt-in mechanism, e.g. an argument to exec(), that restricts syscalls to glibc only. A sandboxing mechanism could then require all executed processes to go through glibc and instantiate them only using that option.
libc is not intended to be the official entry point in any way or form, and kernel vulnerabilities and workarounds are not meant to be handled by a libc implementation.
That other OSs make libc their official interface is primarily because it's the simplest thing to do when kernel, libc and the rest of userspace is co-developed, as it allows for breaking kernel changes and other fun things that are not allowed under Linux ABI guarantees anyways.
It is not because it is the most secure choice, or that dealing with syscalls is hard (syscalls are easy and safe to work with). It's just that stable ABIs are a lot of work to develop, and this structure is just the simplest for smaller OS communities to develop.
Depending on how you define "workarounds", glibc is full of those. For example stat(2) is very much not the same syscall now as it was back in the 1990s. In some cases glibc will do a runtime test to see which syscall variants are supported by the kernel and implement workarounds (I even saw a case where this caused a bug in some programs).
stat(2) is not a single syscall. The changes are exposed as new, isolated syscalls (sys_stat, sys_newstat, sys_stat64), with glibc switching internally between them as it sees fit, surprising developers in the process.
This makes stat a great example of the syscall being easier to work with, more stable and more reliable than the glibc wrapper.
Making libc the stable support boundary has all sorts of advantages to an operating system and basically zero downside. Only vanity argues for doing it the Linux way.
Also, Linux is not POSIX compliant and doesn't necessarily care about being. There are several important IO options on Linux that have nothing to do with POSIX.
Desktop Linux is hardly the majority of Linux installations. How many containers in the cloud are using something like Alpine Linux?
I'd dare say by and large glibc is in the minority.
Note that the mail is from 2020. So it probably already is fixed. It might be this one
https://patchwork.kernel.org/project/linux-arm-kernel/patch/...
vDSO is also a language independent construct, so there's no special treatment for any favorite language, be it C (OpenBSD), C++ (Windows) or Oberon OS (Oberon).
Seems like this is a good idea for multiple reasons. Another benefit is that it would seem to make something like "Wine in reverse" possible/easier to implement.
As far as I understand, Wine (a userspace-only Win32 emulator) is possible because Windows applications always go through the standard library for system calls, which therefore can be hooked without kernel support.
The same is not generally true for Linux binaries – these usually do go through glibc, but direct syscalls are possible (through raising the appropriate interrupt or instruction).
I don't think there's an easy way to trap these without kernel support (which is what WSL 1 has been doing).
OTOH Linux having a well defined set of syscalls you can very convincingly fake being a Linux kernel. That’s what wsl1 did. That’s also what SmartOS does, which allows it to mix native and Linux (LX) zones.
But I don't see why we can't have both (a stable syscall interface and requiring all syscalls to go through a standard library).
What's the point of having a stable syscall interface when you can require that only libc perform syscalls?
Thanks four your explanation. So is WSL2 different in this regards?
(Our company, Fly.io, runs container images for customers on Firecracker microVMs around the world, and I had to build a DNS-dependent service, in Go, that runs directly from our (Rust) init and can't assume a libc exists).
I wouldn't want to take the thread off on a huge tangent, it's just funny that this was just recently super relevant to me (it would have been problematic if DNS-dependent Go programs depended on libc, because right now I can't assume there's a libc binary to be dynamically linked to).
Bringing it back to Go and its (sometimes libc-dependent) DNS libraries: it is very annoying how fiddly it is to get a Go program to use an alternative DNS server.
The sum total of what I want to do: bring up loopback, bring up the one and only Ethernet interface, set up its IP and basic routing, run another program, and do some basic log reporting.
Golang's rules for what implementation to use are found here: https://golang.org/pkg/net/#hdr-Name_Resolution
A really solid alternative DNS client implementation can be found here: https://github.com/miekg/dns. Real easy to read and vet compared to a few other libraries I ran into when working on this problem.
That leaves users annoyed, because sending all DNS into a VPN tunnel is not always an option.
Most recently, this happened with Concourse and it's Fly binary: https://github.com/concourse/concourse/issues/3691
The developers don't want to use CGO, or when they do use CGO they disable the net part... and now stuff doesn't work.
Projects have to specifically build with cgo enabled on macOS, or else it fails.
Not using the libc was always a risky proposition on BSDs anyways. They don't have a stable kernel ABI the same way the Linux kernel does. From OpenBSD's perspective, the stable ABI is the libc, and anything using the kernel ABI directly is liable for breakage with each update.
Citation needed. glibc has a long history of security bugs.
https://www.cvedetails.com/vulnerability-list/vendor_id-72/p...
> especially if going through the lowest level function that directly wrap the syscall.
I think the argument is that security bugs are unlikely to occur in those low-level wrappers.
> Going through the libc is very unlikely to actually introduce vulnerabilities, especially if going through the lowest level function that directly wrap the syscall.
Sure, glibc has a bunch of bugs. But the lowest level of functions, that just wrap the syscalls, are very unlikely to have bugs. Here's the `read` implementation, for instance:
https://github.com/bminor/glibc/blob/21c3f4b5368686ade28d90d...
All it does is delegate to the low-level syscall, with some extra handling around to handle async calls (which can be removed when compiling glibc yourself, but you're probably not doing this).
Here's clone:
https://github.com/bminor/glibc/blob/21c3f4b5368686ade28d90d...
This one's written in asm, and you can't really simplify it all that much more.
All the functions that wrap the low-level syscalls are very hard to get wrong, really. Where the glibc bugs come from are the high-level functions, like pthread. But those can trivially be bypassed if necessary.
But since you're asking, I can think of two upsides:
1. It restricts the kernel attack surface available to only those syscalls that are exposed through the wrapper. (This is obviously dubious if the libc provides a generic syscall function, I don't know if OpenBSD libc has one).
2. It forces the ROP to either find the libc ASLR base, or to find a gadget that calls into the target libc function. This makes ROPs a lot harder to write.
As with any mitigations, they're mostly meant to make the attacker's life miserable. They're not full protections, and can often be bypassed. The point is to increase the cost of the attack.
Musl[1] for one.
[1] https://www.cvedetails.com/product/39652/Musl-libc-Musl.html...
I'm not sure what the situation for Go is. Assuming Go has proper support for dynamic loading and a stable ABI, it would be doable.
Usually the only time the SYSCALL ABI breaks, it's because kernel authors intentionally chose to do it for no apparent reason. For example, OpenBSD at some point changed how mmap() was defined so that it takes seven arguments:
void *sys_mmap(void *addr, size_t len, int prot,
int flags, int fd, long pad, off_t pos);
The sixth argument doesn't do anything. It just breaks binary compatibility. It's also noncompliant with the System V ABI specification, which says system calls have six arguments max.I work on a project called Cosmopolitan Libc which lets you create static binaries that just work on Linux + Mac + Windows + FreeBSD + OpenBSD. It was only possible to do this because Unix systems generally agree on definitions. I worked really hard to support OpenBSD since I believe in the project. I just hope they keep the ABI stable going forward.
For example, I'm really happy that this restriction only applies to dynamic binaries. It seems perfectly reasonable that they'd want to make the assumption that if a program chooses to link OpenBSD's Libc that it intends to use it. That's fine just so long as we continue having the ability to build static binaries with an alternative cross-platform Libc like Cosmopolitan.
Speaking of which, I think I might actually implement some of OpenBSD's ideas in Cosmopolitan. I could probably track down all the functions that need raw SYSCALL and use __section__ so they're all linked to the same part of the binary and then call the msyscall() function to limit it just to that page. That way as a guest libc author I'm upholding the spirit of the intent. When in Rome do as the Romans.
As far as I can tell it might be used to align the memory layout of the following 64 bit arguments to 64 bits. Or at least ensure that there is no auto generated padding that might contain random values.
Edit: The link provided by ainar-g 3 mentions that the way gcc handled padding of the offset field changed between gcc 1 and 2.
It seems like that was the case from day one, when a copy of the NetBSD code was imported to later become OpenBSD[1]. And if my reasoning is correct, saying that OpenBSD “changed it at some point” is not quite correct.
https://github.com/openbsd/src/blob/df930be708d50e9715f173ca...
return((caddr_t)(long)__syscall((quad_t)SYS_mmap, addr, len, prot,
flags, fd, 0, offset));
The first argument is the syscall number, the rest are
the seven arguments in question.That's a really weird way to phrase it. Apple says that interface is unstable. Go uses it anyway. The unstable interface turns out to be unstable.
It wasn't Apple that broke Go binaries. It was Go.
Windows never had a stable syscall ABI either.
Almost only Linux does it... because the kernel and libc are maintained as separate projects in the Linux world.
Practically as long as it's trivially callable from C I'm not bothered.
That's because Linux is just that: a kernel. And people are free to build their userspace on top of it.
Dynamically link against libc and statically link the rest for all I care, but there's no reason not to talk to libc.
Also: the vDSO is also a form of dynamic linking. Are you opposed to the vDSO?
I'm not opposed to vDSO but I disagree with how Linux maps it into memory by default. Linux should not be putting anything into the address space that the executable does not specify. MMUs are designed to give each process its own address space. Operating systems that violate that assumption are leaky abstractions imposing requirements where they shouldn't.
The main thing dynamic shared objects accomplish is giving upstream dependencies leverage over your software. They have a long history of being mandated by legal requirements such as LGPL and Microsoft EULAs. It's nice to have the freedom to not link the things.
Other people have made software for decades without writing program-specific libc instances. Tell me you at least started with something decent like musl instead of literally writing your own libc from printf on up.
> Linux should not be putting anything into the address space that the executable does not specify
Execution has to start somewhere, and kernels have often reserved parts of the system address space for themselves.
> The main thing dynamic shared objects accomplish is giving upstream dependencies leverage over your software.
Loose binding in interfaces allows systems on both sides of the interface to evolve. If you want 100% complete control over your system for some reason instead of writing programs that play well with others, just ship your thing as a VM image and be done with it.
Trapping (SYSCALL/INT) is a loose binding. The kernel can evolve all it wants internally. It can introduce new ABIs. Processes are also a loose binding. I think communicating with other tools via pipes and sockets is a fantastic model of cooperation. Same goes for vendoring with static linking. Does that mean I'm going to voluntarily load Linux distro DSOs into my address space? Never again. Programs that do that, won't have a future outside Docker containers.
Also, my executables are VM images. They can boot on metal too. Read https://justine.lol/ape.html and https://github.com/jart/cosmopolitan/issues/20#issuecomment-... Except unlike a Docker distro container, my exes are more on the order of 16kb in size. That's how fat an executable needs to be, in order to run on six different operating systems and boot from bios too.
Strong claim. Wrong, but strong claim.
The completely-statically-linked model you're proposing might be acceptable on servers, but on mobile and embedded devices like Android, it's a showstopper: without zygote pre-initialization and DSO page-sharing, Android apps would each be at least 3MB heavier than they are today and take about 1000ms longer to start --- and a typical Android device has a lot of these processes running.
More broadly, yes, in most contexts, I see a general trend away from elaborate code-sharing schemes and towards "island universe" programs that vendor everything. But these universes need to interact with their host system using a stable ABI somehow, I believe that SYSCALL is fundamentally the wrong layer for this interaction, as it's not flexible enough. For example, the Linux gettimeofday() optimization couldn't have been done without the ability to give Linux programs userspace code to run pre-kernel via the vDSO. How do you propose the kernel do things like vDSO gettimeofday optimizations?
Doesn't everything on Android start off with the JVM as a dependency? In that case the freedom to not use DSOs is something that Google has already taken away from you. That's not a platform I'd choose to develop for unless I was being paid to do it.
On x86 RDTSC returns invariant timestamps so you technically don't need shared memory to get nanosecond precision timestamps. XNU does the same thing and they don't call it a DSO. Because that's just shared memory. I have nothing against shared memory.
It sits so close to the kernel that the concerns of changing semantics apply as easily to the kernel as they do to the system call wrappers.
There is a way out. I've previously proposed on LKML that the Linux kernel team provide an official userspace system call library that sits below libc and that all libc implementations would share. We'd forbid new system calls being called except through this library. Optionally, we'd enforce this constraint on all system calls.
This is how Fuchsia works, by the way: all Fuchsia system calls must go through one giant vDSO.
Maybe my impression was wrong, but it was my understanding that Go originally preferred direct syscalls because of stack management headaches. You can't know how much stack space a libc syscall wrapper requires, which even for seemingly simple syscalls can be quite complex--e.g. glibc has to emulate POSIX thread semantics. OTOH, treating such libc wrappers like regular C FFI functions would obliterate the design and implementation assumptions around goroutine stack management. Considering that Linux was originally the first (and, let's be honest, only real target), it made perfect sense to rely on Linux syscall ABI promises.
Fast-forward a few years: 1) Go has a more mature binary format and dynamic linking capabilities, shrinking the gap between Go's internal ABI and the native libc ABI. 2) Goroutine stacks switched from split-stacks to movable stacks, and the minimum stack size became larger. 3) Demand and motivation for supporting libc wrappers (i.e. for Windows) grows. Result: Go surmounts one of its original simplifying design compromises. Though, I would assume that libc wrapper support still incurs ongoing maintenance costs on each platform; namely, managing the minimum stack requirement for each particular call, which could change overtime, while it's important not to be too pessimistic so that a syscall doesn't force an unnecessary stack resize.
Unix systems have never had that restriction. Because if you use the official libc and pass -static to gcc then that means the "stable api" creates a binary that depends on the kernel abi. If the kernel authors break the abi then it means you need to recompile all your software in order to upgrade.
Yes they do.
> Only Windows forbids developers from linking static binaries. That's because they change the SYSCALL ordinals every few months.
That Windows somewhat actively precludes raw syscalls doesn’t mean they are supported elsewhere. It’s always been bsd (and especially macOS) policy that syscalls are an implementation detail.
> Because if you use the official libc and pass -static to gcc then that means the "stable api" creates a binary that depends on the kernel abi.
Try doing that on macOS, you’ll find out that there is no static version of libSystem, or crt0.
The main advantage of not having a stable kernel ABI is that you have the freedom to change it for security/performance reasons. This means the OS can change how they implement things over time without breaking things as easily. MS is notorious for this and for using this fact to allow them to 'emulate' older ways of doing things even when the real implementation has long since moved on.
The main advantage of having a stable kernel ABI is that the userspace can be whatever the developer wants it to be realistically. If they want it to be just a single GO program and literally nothing else they can do that (routers are a good example of devices that do this).
Thanks
That means that programs that directly make system calls will keep working on newer OSes.
On many (¿most? https://unix.stackexchange.com/questions/473137/do-other-uni...) other operating systems, that’s not the case; the OS ships with a library that provides a stable interface, and system call numbers, arguments, or calling conventions can change (in theory, the interface could even change across reboots or process runs). On openBSD, that library is Libc (and, unfortunately, is a lot larger than just the OS interface. IMO, in an ideal world, it should be split in two parts, the OS interface and a C library)
Go wants to produce statically linked executables. It can’t do both that and link with LibC. It now changed to dynamically link with LibC, guaranteeing that what you compile today will run as well on next year’s OpenBSD as it does on today’s one.
On top of that, openBSD has a security feature where it verifies that system calls are made via LibC. That feature wasn’t implemented as thoroughly as possible because go made direct system calls. This change allows OpenBSD to tighten that feauture.
I've been developing with Go for a few years now but never strayed too low level so was unsure how this change in Go 1.16 would affect me. I doubt my code will run on OpenBSD in the near future however I'm happy to know that if it does it will be supported in the future.
Edit: It's just the go mantra of compile once run forever doesn't really work on OpenBSD. You do have to recompile under certain conditions.
Edit: All libc basically adhere to it. Except for some extentions that are not in POSIX but might be in glibc/musl/FreeBSDs libc or some extentions that are OpenBSD specific.
So basically executable are mostly only ever good for 1 stable release. So don't throw away your source code, you gonna need it in 6 month.
Go already uses libc by default on many platforms. But there are issues - sometimes the libc behaves differently than Go's documented APIs, but this is primarily a documentation issue. Contrariwise, sometimes, Go's native APIs don't behave on systems due to platform-specific implementation bugs.
I think this is a good move overall.
But also: while it's stable, it's not public. IIRC, Go developers were told about this, but decided to ignore.
And of course: stop breaking LD_PRELOAD hooks :P
Also, many things libc does (which go programs need to do themselves) is to wrap a lot of the calls to things like malloc/calloc/realloc to use internal buckets and as seldom as possible call out to the kernel to get one or ten pages of ram in one call, then hand out suballocations from them for each "obj = malloc(8);" so that you minimize the amount of syscalls made, regardless of if you do them directly or via libc.
Since syscalls always were expensive (you need to save registers, check userid permissions, flip to kernel mode, do the work asked for with or without SMP locking protections, give permissions to uid for the resource returned, perhaps check if its time to deliver signals or switch to another process and if not, flip out of kernel mode, restore registers and return to the process again) a lot of the calls done by a C program is kept in libc if possible.
For example, gettimeofday() springs to mind, where a lot of trickery is done to give all programs a readonly page with the current time mapped into your program space so the call doesn't have to go via the kernel but instead becomes a memory read, since programs tend to call this thousands of times.
The kernel ABI, however, does not. vDSO are shared objects exposing a C ABI, and oddball compilation settings have broken Go's vDSO calls in the past: https://marcan.st/2017/12/debugging-an-evil-go-runtime-bug/
You can go ahead and do the following on your Mac:
GOARCH=riscv GOOS=linux go build
to build a binary that just runs on Linux on RISC-V.These versions are not chosen at random: I ran in to this issue with people trying to run my binary on Ubuntu 16.04 (LTS release), which was solved by linking it statically.
Also, people may use musl libc, and while it has some compatibility with GNU libc this is far from complete.
So in short, linking it statically means it will work for the largest amount of people with a minimal of fuss for both the person building the binaries, and the people running them.
As people have mentioned, these issues are less present on non-Linux systems.
You could also have linked against glibc 2.23 or 2.22 or older instead of bloating your binary.
Adding an extra megabyte or so is a reasonable trade-off, with no real other downsides. It's not that large – smaller than many websites.
OpenBSD is an OS that people choose to use when they want security prioritized as a trade-off against other things such as performance and binary compatibility (there's always trade-offs). Many firewalls use it, for example.
If you value other things more than sheer data security, then there's other (beautiful) choices. (Analogously, there's no single best vehicle for everybody.)
Anyone with a modicum of actual OS kernel development experience would acknowledge these mere facts as universal truths, instead of "unpopular opinions."