OpenBSD System-Call Pinning
lwn.net
lwn.net
I know this annoys unix people. But I have to say I actually really like that Go shakes this up. I believe the C function monopoly just isn't healthy for things. You should be able to make a new completely unrelated language. The Go developers were the first in a long time to do this, not because they are stupid, but because they were ambitious.
As always, things are about API boundaries: what is internal to you, and what is exposed externally to your users.
With Linux, well, we all know the rms copypasta: "I'd just like to interject for a moment. What you're referring to as Linux, is in fact, GNU/Linux, or as I've recently taken to calling it, GNU plus Linux." This is a joke, but it also points to something real and serious: Linux, being just a kernel, means that they provide a stable kernel API. glibc, being written by a different organization, builds on top of that API and adds its own stable API.
By contrast, many other operating systems are developed as a full operating system. They don't produce a kernel as a standalone component. As such, they choose some other sort of boundary as the API for the OS. On unices, that's often the libc. On Windows, that's the Windows API, of which win32 is an example.
There are good reasons to not make the kernel your API boundary. Different systems make different choices, and that's a good thing, not a bad thing.
Good reasons for the developers, not so much for the end users.
But Linux has, to the best of my understanding said "yes, we are ok with users using syscalls". Linux doesn't think that glibc is the only project allowed to interface with the kernel.
But for other platforms like OpenBSD and windows they are quite simply relying on implementation details that the vendors consider to be a private and unsupported interface.
This whole thing is also separate from "is making libc the only caller of the syscalls instruction" a good and meaningful security improvement.
The library you have to link to access system services is not going to pollute your language environment with bad runtime.
DNS without CGO works perfectly. The vendor specific ad hoc mechanisms for extending DNS in a site local context are not well supported. If they were implemented more sensibly, then Go, or any other language, would have no problem taking advantage of them even without the "C Library Resolver."
Speaking of which, that "C Library Resolver," in my opinion, has one of the worst library interfaces in all of unix. It's not at all a hill worth new projects dying on.
It does not. I know this because it impacts my daily work and the work of others. Honestly if you could make my day and go figure out exactly what's going wrong with the pure go DNS implementation it would make my life alot simpler and I wouldn't have to maintain shell scripts that update etc/hosts to hard code in ipv4 addresses for the APIs I access with terraform.
https://github.com/hashicorp/terraform-provider-google/issue...
Should it be your DNS resolver library that takes stock of your OS environment and only make calls for A records and not AAAA records when it "detects" some configuration?
Shouldn't your application itself have an environment variable or command line option that allows you to specify that your dials should only be done using tcp4? Wouldn't this be immensely useful to have outside of "auto detection" in some library somewhere?
As a user, if ipv6 is flaky today, I want one central place to configure for ipv4 only, I don't want to go and change every application'settings only to revert that tomorrow.
DNS is one of those things that OS vendors think should be extendable and configurable. It allows VPN apps to redirect DNS only for certain subdomains, for example, which enables proper split-horizon DNS. I think this is totally reasonable behavior, and it’s undeniably useful. If a particular programming language reimplements DNS on its own, you lose guarantees that the OS is striving to provide to the user.
You can make the case that OS’s shouldn’t make these guarantees, and we’re free to disagree on that, but from a practical standpoint it is a very useful feature and it sucks that pure Go apps don’t work with it.
I am not super happy with systemd-resolved but it solves this particular issues. No requirement to use libc, but same (os configurable) behavior for all users.
Yes, if you're writing netstat or lsof or ps or something, you need tight coupling with the binary and the kernel, and you can argue Linux does that better, but most people aren't writing netstat or lsof or ps.
God I hope not
For information of general interest such a special kernel page could be mapped as read-only in the address space of all user processes.
Much of the information that is provided in the special file systems /proc and /sys could have been provided in some appropriate data structures in such read-only shared memory, for a faster access, by avoiding the overhead of file system calls and of file text parsing.
The aforementioned tools use these interfaces.
Emphasis on most of the sundry information for the live kernel now comes from sysctl, I note the (root only) mem/kmem interface for completeness and rare utilities (eg btsockstat) use it.
Going way back, this is how it all used to work, the more structured interfaces were a 90s thing. https://github.com/v7unix/v7unix/blob/master/v7/usr/src/cmd/... Even early Linux used kmem for ps. Not ideal https://cdn.kernel.org/pub/linux/kernel/Historic/old-version... Also why the package is still called procps, for a while it coexisted with kmem-ps.
Arguably, using the same ABI for userspace C functions and for communicating with the kernel reduces the amount of work required of completely new languages, because they are likely to need C interop support anyway.
Also, if you're gonna add new ELF sections anyway, why not do syscall "relocation" directly (with a similar randomization like ASLR) instead of going through stubs? This "relocation" doesn't actually need to change any memory or use any offset tables since the syscall numbers are a farce under the new system anyway. Just make it a map of the location of syscall instructions to implied syscall numbers. Once you've phased out the old syscall model, you can even repurpose RAX as an additional syscall parameter.
I think they should split that part of, too. Such a split would better reflect the split of responsibilities between kernel maintainers and libc maintainers.
At the same the effective difference for people using the code is minute.
Any decent linker will strip out unused functions in libraries you link with, so if you currently only use the part of libc that would become “libsyscall”, the end effect would be the same as when “libsyscall” existed.
> avoiding [...] errno
Note that errno can be optimized by a sufficiently smart compiler, simply by treating it as a register, then annotating various functions with whether it is preserved, clobbered, or conditionally used for return.
The libc boundary is quite annoying though; for this case in particular, that it doesn't expose the fact that errno is at a fixed offset from the TLS register. And generally, libc is vehemently opposed to the existence of smart compilers, since all libc calls are expected to be treated as opaque barriers.
Yes, as far as I'm aware none of the major libcs on Linux support LTO
One can only hope that more will eventually find its way into Linux, like with how paid employees at Google have been spending the past ~6 months cloning mimmutable (which HN characters decried as "useless") to make mseal() for defending Chrome.
Really? And how many systems have made any effort to use it, on OpenBSD most of a programs static address space is now automatically immutable (main program .text, ld.so .text, .bss, main stack, dymamically-loaded shared libraries, and dlopen()'d libraries mapped w/ RTLD_NODELETE.
Nobody else have done the work on a complete operating system.
You're basically going "nobody else did this properly" because others did a different implementation. In other operating systems at least they go "oh we saw a chain that targeted xyz structure in this page and modified it so we are going to make sure it is really immutable". How did OpenBSD arrive at the conclusion that what other people are doing doesn't actually confer the full security benefit?
https://news.ycombinator.com/item?id=38579913
I think? this is the same work, just now merged into the kernel?
> Security researchers have expressed doubt about how useful this check is at preventing compromises.
Doesn't cite the HN thread, but two other cases.
> I think? this is the same work, just now merged into the kernel?
That is my understanding, yes. From the article:
> In December, De Raadt sent a patch to the OpenBSD mailing list expanding OpenBSD's restrictions on the locations from which a process can make system calls. ... Now that patch has been merged, finishing a process which De Raadt said has taken five years.
Part of why it is hard to take its reputation as as security focused OS seriously.
Failed experiments are also a good part of research.
> Only two remote holes in the default install, in a heck of a long time!
I'm assuming not, but I could always be mistaken.
this is on top of a lot of very careful programming and interesting security research, and this post isn't meant to take anything away from the OpenBSD devs.
Probably this issue has been hashed out many times over the decades, but arguably the security gain isn't a fortunate or incidental benefit of minimizing default enabled services, nor a cheat like weighted dice, it's a very real benefit resulting from an effective, intentional technique. Maybe other OSes should do the same, and then everyone would have that benefit.
The other OSes have other priorities, and that's fine. Embrace that. Yes, most users (and developers) don't want to deal with the compatibility issues. But when you say OpenBSD has few default security holes because they have few default services, that's a complement.
Also, doesn't SEL4 have widespread, practical application? IIRC as the microkernal (maybe under Minix?) on the baseband hardware on cell phones? Maybe I'm confusing it with something else.
As an example, CompCERT is a formally verified C compiler, and it's had a couple bugs as a result of their specification of the underlying hardware being wrong.
Secure in the absence of an attacker?
This part:
xx: b8 05 00 00 00 mov $0x5,%eax
xx: 0f 05 syscall
This means "perform operation #5, which is open(2)"
Inside the kernel, we know the system call # and the address of
the syscall instruction.
Why can't the attacker just jump to an existing syscall instruction? Maybe one followed soon by a return. 4) in libc.so's table, all the system call stub "syscall instructions"
are registered.
I don't know a lot about binary exploitation techniques. Is all this entirely reliant on layout randomization?A more rigorous analysis of the security environment as a whole would be useful.
Well... if they are able to craft a call to a syscall from some random piece of memory you are already f-d and this little hurdle is at most going to be an annoyance.
How are/were static executables handled? I’m a little fuzzy on how execve notices that a given executable is dynamic (and needs ld.so to run it) or static (in which case.. it just jumps to the _start symbol?). If the dynamic linker isn’t involved, what calls msyscall/pinsyscall?
Is it materially different in openbsd?