When you're dealing with shared resources (like a listening socket), and you are using threads, and you are trying to maximize performance with networking, confusing and painful is kind of the nature of the problem.
When you're dealing with shared resources (like a listening socket), and you are using threads, and you are trying to maximize performance with networking, confusing and painful is kind of the nature of the problem.
>accept() on a fd that has already been closed.
Why are people doing that?Source: at the time I worked on high-performance server products and had discussions with various people active on the kernel side of the problem and encountered much head-shaking and muttering about patents, when I mentioned implementing IOCP.
fwiw there was an IOCP-like capability added to AIX, again supposedly because IBM did not have fears about patent infringement (likely due to cross-licensing arrangements).
What is so disturbing is this has been known forever that epoll should be drown in a bathtub but no one does, it's just "la dee da I can't here you! I can't hear you!" and Linux trudges on with yet another limp.
As he also says "kqueue and event ports are T-Ball, compared to ZFS and Dtrace and jails are a lot harder... if you got the little stuff wrong..."
It's just frustrating. It's broken. You know it's broken. FIX IT!
While I trust the author to do this (thankfully, as he's my coworker) there is a lot of Linux software that doesn't, even assuming it was updated in the last year and you're running something vaguely bleeding edge (not Debian).
Some people using a completely different kind of OS insist in continuing to use a version from 2009, instead of its more recent version. I guarantee you that this version from 2009 also has its share of dark APIs, some of which have been fixed in the modern version...
Yet I've yet to find articles titled like if all of them are equally broken, while the content would precisely describe the caveats to do not-broken things on the capable recent releases.
when you deal with multi threads, not only epoll will cause some problems, but also global variable, memory, etc. global variable solved by mutex, but epoll solved by avoid using it to epoll_wait fds in multi threads.
I wonder if we could put a sane fix into libc to fix these problems.
Everyone's approach towards these mechanisms is broken. Just don't treat it as a reliable notification mechanism, do your own scheduling using information from epoll/kqueue only as hints and everything will be fine.
It's also hard for me to imagine a scenario where accept() takes longer than servicing a request and becomes a bottleneck. That is, why would you need multiple threads accepting on the same socket?
Second issue is cache locality. If you do accept() in one thread only, then you will need to move the new accept-ed client socket to another worker thread. Depending on details this might not be efficient - aRFS comes into mind. (but frankly, epoll alone won't help here, you need SO_REUSEPORT with SO_INCOMING_CPU).
Even if you don't agree that scaling out accept is a real concern - that's missing the point. The point is: the epoll() model should take this into account and at least support this problem. Or loudly say that scaling out accept() with epoll is not possible. But neither things happened. Up till kernel 4.5 it was impossible to do it correctly, but undocumented, from 4.5 you can use the EPOLLEXCLUSIVE flag, which I feel is a hack.
Your writing will be more convincing if you profile.