Dyad: Minimal, portable async networking library for C
github.com
github.com
The problem with poll/select is that you need to pass (transfer) to the kernel the entire list of events you're interested in, just to wait for a single event. Each time you re-enter your event loop you need to do this, and if you have thousands of connections, it will break down, because poll is O(number of connections).
epoll, for example, works around this by having an "epoll FD", which on the kernel side contains all the events you're interested in. You wait by just waiting on that epoll fd with epoll_wait, but you don't specify the events you're interested in when you wait: you do that ahead of time, with other system calls. This allows you to change the event list only when you need to, which is much less frequent than some data arrived from somewhere. The API is supposed to be O(1), instead of O(number of connections) per wait call.
My understanding is the kqueue works similarly, but I'm a Linux guy, so I can't really tell you.
select also has other problems w.r.t. FDs with high numbers.
create a socket, bind, listen, connect, send etc.
Why do you need select?
What's wrong with tying up one thread per socket?
http://blog.tsunanet.net/2010/11/how-long-does-it-take-to-ma...
On 32-bit systems, you can also easily run out of address space for your thread stacks.
Also in case of an IO bound thread, they are not just spinning on a futex aimlessly. There is a different mechanism, so should really benchmark with a more characteristic workload.
Speaking of characteristic workload, they should have probably also measured on a tickless kernel since I saw they complained about time quanta and HZ=100. Well recent kernels are tickless so they'll behave differently. (Might be worse even).
> On 32-bit systems, you can also easily run out of address space for your thread stacks.
Well don't run large servers with so many threads on 32 bit systems ;-) Many database vendors don't even package for or support 32 bit versions of Linux.
Sorry, I haven't bought into the whole "async is always better" trend. Some (ex?) Senior Google engineer (Paul Tyma) agrees with me:
http://www.mailinator.com/tymaPaulMultithreaded.pdf
Async / select pattern is usually good where there is very little business logic. Like a router, proxy or simple web server and so on. In a large application having a giant dispatch call at the center of it, with callbacks branching out is not a healthy pattern.
When you have 10,000 tasks and about 8 cores (give or take a few) the number of context switches is very large. Switching in the kernel will happen mostly in the system call boundary of blocking IOs and require the scheduler to make a decision on what thread to wake up next and then change the running process.
This can be seen in function context_switch inhttps://github.com/torvalds/linux/blob/master/kernel/sched/c... without the arch dependent components and can hardly be compared in complexity and effort to switching between 4 and 8 registers in user-space.
The above still doesn't include any changes to the TLB and memory protection tables as I assume the OS optimized those away when it switched between two threads of the same program. An optimization I'm not sure that happens normally.
I personally prefer the old-school async approach, because there you are forced to explicitly manage your connections' state, and the application/process-wide data access is inherently race-condition free. I'd use this as far as possible.
If you let your OS schedule threads, obviously you have to be careful that shared data is correctly locked/only atomically changed, but you get parallelism (especially for CPU heavy tasks) for free. If you are used to do these chores (I'm not), perfect! And your connections state (or the state of required computations) can be arbitrarily complex (ugly?) and still quite elegantly hidden in your threads's stack.
So, I don't see that one approach is better than the other. For me the extremes are probably clear in favor of one or the other, with a large grey area in between.
* closing the socket (it would be from another, monitoring thread in this case)
* setsockopt(fd, SOL_SOCKET, SO_RCVTIMEO)
* Would that be different than blocking on select with a an infinite timeout. How do you cancel that? Or are you relying on other sockets getting constant stream of data to wake you up?
* do something with an ALARM signal
A timeout will work fine, but now you're polling, meaning you have an unpleasant tradeoff between efficiency and how long it takes for your thread to notice that it's dead.
Canceling select or any other multi-fd call is really easy. Create a pipe and add it to your fd set. Any time you want the thread to wake up (e.g. because you need to tell it that you're canceling something) you just write to the pipe.
Signals have a similar race condition as closing the socket. If the signal is delivered after you check for cancellation but before you enter the system call, you'll hang.
That is true. To go more in-depth, you'd do shutdown first. But I think you have to be connected for that.
> Canceling select or any other multi-fd call is really easy. Create a pipe and add it to your fd set. Any time you want the thread to wake up (e.g. because you need to tell it that you're canceling something) you just write to the pipe.
That a good way, agree. But I would still use a select with 2 file descriptors per thread. One fd for the pipe and one for socket itself. Each thread handles its own request and processing as needed without having one global dispatch in the whole application. Pipe is exposed to the outside in case shutdown needs to be triggered (from another thread).
That's one of the big reasons to use a lib for this.. so you get the best performance, without having to change your code to get it (or bother detecting which is best, etc).
"portable" -- portable to what?
comparisons -- vs. libevent, libuv, etc.
in the sample program -- when will dyad_getStreamCount() drop to zero? will it ever? what exactly is a "stream" anyway and how does it relate to individual connections?
http://www.wangafu.net/~nickm/libevent-book/Ref8_listener.ht...
It looks like this library has only 2 weeks worth of commits so I'll reserve any criticism. My vote is still with libevent though!
So for 'r', why bother using the FILE* formatted IO if all you're going to do is pull single characters at a time? You can just use file descriptors.
I'm not saying there's not a need for libraries like libevent, but IMO it's not needed for most applications. I might be biased because I want complete control and know what happens, I think that's important. I don't want to use a library before I know what happens in the "background". When I know that and know what I need, only then it might be appropriate to use a library.
I could do it myself, but it would just end up looking more and more like libevent each time I did it.
It's a very neat concept that allows for cleaner code (sequential rather than callback-hell) and use one or a few threads for many actions but without the overhead of thread-per-socket.
There are also some hidden gains in terms of TLB caches and other costs of kernel threads switching.
An additional advantage is that between user-space-threads you have fewer locking problem since they implicitly lock out each other between context switch points so you only need locks when you need to protect an area across several context switch points.
Once the stack allocated for a thread, context switches are almost as cheap as a function call.
RiNOO has an event driven scheduler, based on epoll, which resume/release these user-space threads (that I call tasks) according to pending IOs. The library provides with IO functions (read, write...) which use the RiNOO scheduler.
As a bonus, real threading is quite easy: just need to run a scheduler per thread (see examples with multi-threading).
I wrote my own user-space-threading library called libwire. Mostly for not liking the malloc-everywhere approach so common everywhere. The tradeoffs are different and the code is more verbose at times but I like the fact that there are no mallocs in code, at least not very explicit ones. I do provided a memory pool so it allocates memory but it is bounded in size.
libwire: https://github.com/baruch/libwire
list of coroutine/user-space-threading libraries: https://github.com/baruch/libwire/wiki/Other-coroutine-libra...