> why not just use poll(2) instead?
Because as I mentioned later poll basically is the same as select but requires more memory to be copied to/from kernel and was slower in some test cases although I can come up with cases where poll will be faster than select. Networking libraries like libev and others allocate fd_set the same way.
>> To find out which sockets have events some loop through data structures and check each socket they contain. But a much better way is to iterate through fd_set as array of long and check 32-64 bits at a time.
> you're better off switching to a socket API that scales well (epoll or kevent) and does this filtering for you. Or like another commenter suggested, using a library that abstracts this functionality.
That's how kernel, libev and others work with fd_set.
>> The correct way is to maintain a descriptor set all the time and create a copy of it and pass the copy to select.
> Again -- if you've reached the point where this tradeoff matters, just go directly to epoll/kevent.
Again -- libev and others do this
> Maybe your program deals with non-socket fds, and the set of socket fds is fairly sparse. Using a map is actually pretty reasonable even if fds are dense.
Yes, there might be cases where you have non-sockets and you can't use array indexed by socket. But it's great in most of the cases and kernel will keep it as dense as possible. It might be that you have 10,001 connections and then 10,000 are closed and highest socket is still in use and array memory is wasted. But it will not require more memory than during peaks.
libev use of select: http://cvs.schmorp.de/libev/ev_select.c?revision=1.56