File Descriptor Limits
0pointer.net
0pointer.net
I've personally had to debug and fix crashes due to this assumption at multiple jobs.
Back at Hostway decades ago we had Apache+SSL instances randomly crashing which turned out to be due to something opening /dev/random and using select() on the returned file descriptor in a .so. This only happened when the machine had enough virtual hosts to exceed FD_SETSIZE on the spurious open of /dev/random.
At Bizanga we had random crashes in their "carrier grade mail server" software, which would link in proprietary modules from vendors like Symantec for realtime virus scanning before even accepting messages for delivery. This stuff ran in-process for performance reasons, just in separate threads. Being "carrier grade" meant thousands of concurrent connections -> thousands of open files. Those vendor-provided plugins would phone home behind the scenes for virus db updates. Turned out they were using select() on the socket connected for the update, if there were enough active sessions at the time, the socket's value would exceed FD_SETSIZE and boom: segfault.
People tend to forget that it's the value of the fds you pass to select() which can't exceed FD_SETSIZE, not just the quantity of file descriptors. It could be just one file descriptor, but if it's a high enough value exceeding FD_SETSIZE, it will segfault. The value is used as an index into a fixed size array.
I ended up making an execution wrapper @ Bizanga to launch everything with >FD_SETSIZE file descriptors already opened once I saw the vendor modules were linking select(). We found numerous crashes this way, without waiting for them to occur in production.
Perhaps it might be time to start treating select() like gets() was treated: the compiler should emit a warning by default whenever it's used, and after perhaps a decade, the API should no longer be available for newly compiled code. While not as bad as gets(), which according to Rusty's classification of API quality (https://ozlabs.org/~rusty/index.cgi/tech/2008-03-30.html and https://ozlabs.org/~rusty/index.cgi/tech/2008-04-01.html) gets a -10 ("It's impossible to get right"), select() is around a -5 ("Do it right and it will sometimes break at runtime"), or perhaps even a -7 ("The obvious use is wrong").
And break perfectly working code?
I'm with you here. I'm also guilty of using it still for some CLI tools or other small programs. Just to be clear you don't need to assume all fd values will stay within FD_SETSIZE. Just that the fds you will use with select will stay in that range. Even if there will be lots of fds (more than FD_SETSIZE) for other resources, it works for example to select on stdin if you're not doing funny business reopening it.
There's no much of an excuse to keep using it though other than laziness.
Edit: I'm biased by Python which has `select.poll` with API similar to epoll and it's really hard to use it instead of select.
For many years pselect on macOS didn't actually fix the race condition, notwithstanding its UNIX03 certification. pselect was just a naive wrapper around select that didn't atomically unmask signals. So I ended up having to create a kqueue-based implementation for use on macOS and stragglers like OpenBSD which at that time lacked pselect entirely. See https://github.com/wahern/cqueues/blob/87f14f9/src/cqueues.c... I can't remember if pselect has been fixed on macOS.
The same approach could be used to emulate ppoll on macOS. Though epoll, kqueue, and Solaris ports are so similar that you can easily wrap all of them behind a thin shim. See https://github.com/wahern/cqueues/blob/87f14f9/src/lib/kpoll... (Caution: it's an incomplete refactor of a similar shim I used in some non-open source code. The original code is even simpler and has seen extensive use. The refactored shim should work but is untested.)
AIX is really the only mainstream Unix lacking an API w/ equivalent semantics. AIX has a better-than-poll interface, but it doesn't work anything like the others. Notably, epoll, kqueue, and Solaris ports control descriptors are all pollable in their own right, permitting you to nest different [readiness polling-based] event loop frameworks.
The self-pipe trick is where you create a pipe(), add the read side to select(), poll() or other syscall, and have your signal handlers write a byte to the write side. When reading the pipe to flush this state, remember to read until it's empty because there may have been multiple signals. Reading 2 bytes and looping if you get 2 will get this done usually in one syscall.
It's best in principle to make the pipe non-blocking so the writes can never block if the pipe fills. This is basically the portable equivalent to Linux's eventfd. (If you're worried about out-of-memory preventing non-blocking write of a single byte, you can close the pipe instead, but this causes several complications because it's not safe to close the same fd twice, and you may have signals interrupt other signal handlers, so race hazards around closing the fd twice must be protected against by blocking signals inside the signal handler and using more flags.)
You can optimise self-pipe by setting a write-enable flag just before select() etc and clearing it afterwards, with the signal handlers only writing a byte (or closing the pipe) when the flag is set, and setting a separate flag to say that the signal occurred, which is checked between setting the write-enable flag and before the select() etc syscall. Check the signal-happened flag after select() etc and if it's set, read the pipe to flush even if it doesn't show as readable in the fd set, because the signal could have happened later. A further optimisation is for the signal handler to clear write-enable when writing the pipe, so subsequent signals don't do it again. Don't assume this means there can only be one byte in the pipe, because nested signals introduce a race hazard to writing and clearing write-enable atomically.
The sigsetjmp/siglongjmp trick (setjmp/longjmp on platforms without the "sig" versions) is to sigsetjmp() before select(), poll() or other syscall, then set a jump-enable flag (or just jmp_buf pointer) telling your signal handlers to siglongjmp(), then clear the flag after the select() etc syscall returns. This is more complicated than self-pipe in a number of ways (and a bit slower), because you have to be careful to prevent nested signal handlers or deal with them appropriately, and it's doesn't work reliably on all platforms. So I would stick with the self-pipe trick.
What should we use instead? poll()?
So as long as you use a Linux kernel made this century, you will be fine.
Is that true? I have a memory of working on systems where you can pass larger arrays in if you want, you just have to allocate your own bitvectors instead of using the convenient struct. A quick look at the linux kernel makes me think that this is true for Linux; the kernel doesn't appear to reference FD_SETSIZE in its implementation.
"The Linux kernel imposes no fixed limit, but the glibc implementation makes fd_set a fixed-size type, with FD_SETSIZE defined as 1024"
Everyone using select, with very few exceptions, is using glibc.
That limitation does appear to exist in glibc if you look at the source code. The source code requires you to pass a glibc defined 'fd_set' type https://man7.org/linux/man-pages/man2/select.2.html
That struct is defined here: https://github.com/bminor/glibc/blob/595c22ecd8e87a27fd19270...
It does seem to have the limits described by the blogpost.
> The default size of FD_SETSIZE is currently 1024. In order to accommodate programs which might potentially use a larger number of open files with select(), it is possible to increase this size by having the program define FD_SETSIZE before the inclusion of any header which includes <sys/types.h>.
Although, if you link to a library which uses select and doesn't check that fds it uses are valid with the FD_SETSIZE it was compiled with, you're still in for a bad time, as suggested in the article. And still, probably use something better than select.
glibc: https://github.com/bminor/glibc/blob/master/sysdeps/unix/sys...
musl: https://git.musl-libc.org/cgit/musl/tree/src/select/select.c
*BSD allows you to define your own FD_SETSIZE before including system headers.
Some really old platforms with select() don't even define fd_set or FD_SETSIZE, and you're on your own :-)
Linux added the poll() syscall in 2.1.23, and Linux select() had a kernel-side limit of 256 fds up to 2.1.26, when it raised the kernel-side limit to 5120/10752 depending on architecture. It was 2.5 years until the kernel-side limit was removed, by which point everyone would have had poll() in Glibc.
So if you are going to the effort of writing portable select() except using the trick of defining your own, larger arrays, and using poll() when it's available, on Linux don't bother. You'll hit a kernel limit anyway on those very old systems which don't have poll().
Can somebody please identify the POSIX method by which one can wait on a set of descriptors, that does not involve spinning in a polling loop and gobbling up a CPU thread? The last time I had a need for such a thing, select() was the preferred way to do that.
Aside from being a lot more recent (posix 2001), poll was also broken on macOS until 10.9. Though you can always use kqueue instead that remains a portability annoyance: Linux still doesn’t have kqueue.
And with io_uring around it never will unless someone does a pure userspace implementation. Not certain if possible but it might be.
If you want both portability and performance (as in, O(1) instead of O(N) as poll() and select() ), there's no portable standard. In that case you're better of with a wrapper library like libuv or livev.
That's what most of the article is about.
It’s faster than poll, and if you are accept()ing a lot, faster than epoll.
Multiple threads can do better, but only by working together. SO_REUSEPORT+select is faster if you can get away with it. Iouring looks promising though.
> Using it assumes the execution environment keeps all file descriptor values within FD_SETSIZE.
That’s what the rlimit is for. If you call a program that uses select with 1025 open FD you get what you deserve.
I thought that FD_SETSIZE was a compile-time constant that limited what values for file descriptors could be passed to select.
IOW, it doesn't matter if you have only two files open; if one of them has an fd with a value of 1026, then select will do the wrong thing.
The Linux syscall does not have a fixed FD_SETSIZE so glibc could increase the size.
This is exactly correct.
> The Linux syscall does not have a fixed FD_SETSIZE so glibc could increase the size.
The application writer can define FD_SETSIZE before including <sys/socket.h> and set it to whatever they want on BSD (including OSX). glibc could* adopt this trick, but my feeling is you probably don't want to stack-allocate this stuff anyway.
Dealing with a large number of very-active connections concurrently efficiently requires more effort than just this though.
A compatibility hack for some applications reliant on badly-done select() might be to set RLIMIT_NOFILE to FD_SETSIZE.
Of course you can.
dup2(0,1026)It's more than promising. It's an unbridled joy to work with (compared to things like epoll) and performs very well already.
The reason it's not generally done is because the select API has been effectively replaced with better things like kqueue and epoll and poll and ppoll. If you actually want to do something like select and you expect to have a lot of FDs, using something newer than select is a better choice.
Edit: the other thing is everytime you call select, you've got to pass that N-bit set value, 1024 with the normal value of FD_SETSIZE. I've run on systems with 4M fds (limit set to 8M), that's a huge value to pass and iterate through for every select.
The performance difference between select() and poll() isn't measurable for small sets of descriptors, and for epoll and friends, fixed setup costs are substantially higher. It's therefore not correct to suggest select() has been replaced with better interfaces, in the case of poll() it is a more complex interface and in the case of epoll (or io_uring or ..) it is a much more expensive interface.
typedef struct {
int fds_bits[howmany(FD_SETSIZE, sizeof (int) * CHAR_BIT)]
} fd_set;
<sys/select.h> also provides a set of macros for initializing the structure (FD_ZERO), and testing (FD_ISSET) and setting (FD_SET, FD_CLR) descriptor numbers. But these macros have no way of reporting if a specified descriptor overflows the bit array. In all the implementations I'm familiar with, including Linux/glibc and Linux/musl, they don't even check, and will happily overwrite the array if given a descriptor greater than or equal to FD_SETSIZE.But the maximum descriptor number in the bit field is reported to the kernel as an argument to select. And most kernels, including Linux, obey this parameter[1], and thus impose no intrinsic limit. If userspace provides a validly allocated and initialized buffer all will work just fine.
But how to allocate and use a properly sized buffer? Some systems, like OpenBSD, allow you to define FD_SETSIZE (or a similar constant) before including <sys/select.h>. On other systems, like modern Linux/glibc or Linux/musl, you have to get clever, like doing something like
#define MYFD_SETSIZE 100000
#define MYFD_SETSIZE_OVERFLOW (MYFD_SETSIZE - FD_SETSIZE)
union {
fd_set fdset;
char extra[sizeof (fd_set) + howmany(MYFD_SETSIZE_OVERFLOW, CHAR_BIT)];
} u;
but then be careful not to rely on FD_ZERO, which won't know to zero out the high bits.[1] The same is true of (struct sockaddr_un).sun_path. On most systems you can pass Unix domain socket paths longer than 108 or whatever the local constant is, and the kernel will happily use the full path, though there may be another internal limit, like 256 or 1024. It can do this because syscalls like bind(2) and connect(2) take a separate addrlen parameter. One possibly shocking consequence to many people is that types like `struct sockaddr_storage` are not sufficient to represent every possible socket address. And unless you check the output of the addrlen pointer parameter to syscalls like accept(2), you may have received a truncated address without knowing it.
https://idea.popcount.org/2016-11-01-a-brief-history-of-sele...
(for example: did you know early BSD had a limit of 20(!) file descriptors?)
This sounds like a great change in that at least software can opt-in without having to even edit their systemd unit.
But with the systemd change he mentioned (which for Ubuntu, exists in 20.04 but not 18.04) - that the HARD limit is now 512k - applications can chose to adjust the ulimit from inside the application that is not running as root. So all modern software can just, if it knows its OK, make the call during startup to increase the limit to whatever it thinks is needed. Then we dont have to override.conf systemd services every time you hit a scaling limit or think about it.
Servers, sure; each connection can easily consume 2-3 fds, that’s 300-500 simultaneous connections - that’s nothing.
Systemd and other service supervisors, sure, with pidfds, timerfds, sockets, proc file sockets and whatever else.
But for your average image editing softwares, email clients, CLIs, etc; id think it’s probably better to keep the limit? That way you are saving yourself from badly written software draining too much resources.
I go into common ways to configure the NOFILE/ulimit here:
https://wakatime.com/blog/47-maximize-your-concurrent-web-se...
Mainly, keep checking the ulimit of your daemon process while configuring each layer until the ulimit bubbles down and shows up on the process.
The issue is included in the main description and bugs section.
Sure, Lennart is perhaps the Antichrist of Linux, but I'm pretty positive he's the Antichrist we need.
The long-lived fear of breaking existing code is certainly valid, but at the same time, why not try and see what will break before we write something off. It's often not much. It's practically the only way for us to shed the crufty bits, and boy are there a lot of those.
Pedantic clarification, but it's worse than that. It doesn't work on any fd with a value higher than FD_SETSIZE (default is 1024) - 1. You could be working with single fd with a value of 1024 and it would not work.
On the bright side you can extend this limit pretty easily, but it's certainly a good reason to avoid select whenever possible.
...which unfortunately due to the Linux ecosystem, where people apparently like to inject code into running processes (https://skarnet.org/software/nsss/nsswitch.html) (can't imagine this is brittle at all), can cause your program to break if someone injects a module that doesn't ALSO follow this best practice into your program.
I wonder if something like the eBPF compiler could be repurposed here, to generate lightweight "binaries" for a configuration that can be linked without needing to dynamically link their own dependencies at runtime.
IMO PAM should be replaced with a daemon. This way all the code that deals with authentication can be isolated into its own process, and my screensaver doesn't need root access to check my password against /etc/shadow.
It'll also simplify SELinux settings.
I'm not a Linux person, but select requiring help to work beyond FD 1023 isn't new. It also happens on FreeBSD and MacOS, so either you recompile your software (including glibc on linux, apparently) with a bigger FD_SETSIZE, or you use one of the better than select interfaces that doesn't care. If you tuned your FD limit higher, you're already on the hook for knowing about it.
Keeping the default soft limit of 1024 means that programs that don't fiddle with the soft limit can have a working select, and programs that do fiddle with the soft limit should know better. Auto tuning the hard limit to something appropriately sized seems reasonable, but fiddling with the soft limit may end up with this select problem, so it's nice that they didn't fiddle with it.
Also, I'm deeply surprised to be defending Mr. Poettering, but in this case, he's right.