Show HN: Chan – pure C implementation of Go channels
github.com
github.com
chan_t* chan_init(int capacity);
should be "size_t capacity", since negative capacity values are invalid. This also eliminated the need for the first check in the function. chan->buffered
is duplicate in meaning to chan->queue being non-NULL.Why chan's mutexes and pcond are malloc'ed and not included in chan's struct? They are also not freed when channel is disposed, unless I'm missing something. Mutexes are also not destroyed when buffered chan is disposed, but they are created for it.
(Edit)
Also, my biggest nitpick would be that chan_t structure should really not be in chan.h header since it's meant to be transparent to the app and the API operates exclusively with pointers to chan_t. Then, if it is made private (read - internal to the implementation), then you can divorce buffered and unbuffered channels into two separate structs, allocate required one and then multiplex between them in calls that treat buffered and unbuffered channels differently. I.e.
struct chan_t
{
// Shared properties
pthread_mutex_t m_mu;
pthread_cond_t m_cond;
int closed;
// Type
int buffered;
};
struct chan_unbuffered_t
{
struct chan_t base;
pthread_mutex_t r_mu;
pthread_mutex_t w_mu;
int readers;
blocking_pipe_t* pipe;
};
...
or better yet, just include a handful of function pointers into chan_t for each of the API methods that need more than the shared part of chan_t struct, initialize them on allocation and put type-specific code in these functions. It will make for a more compact, concise and localized code.(Edit 2)
Here's what I mean with function pointers - https://gist.github.com/anonymous/5f97d8db71776b188820 - basically poor man's inheritance and virtual methods :)
struct chan_t { pthread_mutex_t* m_mu;... }
And initialize it like so: pthread_mutex_t* mu =
(pthread_mutex_t*) malloc(sizeof(pthread_mutex_t));
pthread_mutex_init(mu, NULL);
chan->m_mu = mu;
eps's suggestion is to put the mutex directly in the struct: struct chan_t { pthread_mutex_t m_mu;... }
and initialize it by passing its address to pthread_mutex_init: pthread_mutex_init(&chan->m_mu, NULL);
chan->m_mu = mu;
that saves some mallocs and dereferences. If you are new to C, the notion of taking the address of fields in a struct may be unnerving, but you get used to it :)Valgrind is your friend :)
http://valgrind.org/docs/manual/quick-start.html#quick-start...
struct foo
{
pthread_mutex_t bar;
...
};Just allocate a pointer on the stack, read into it, and then assign the value into the provided space:
void *msg_ptr;
read(..., &msg_ptr, sizeof msg_ptr);
*data = msg_ptr;Also worth noting, he really shouldn't be using '_t' since it's reserved by the standard.
Not to say that it wouldn't work, just that it isn't really necessary - casting a pointer to a structure to a pointer to its first member (and back) is perfectly portable. Perhaps even idiomatic. (Though the casts in that gist seem to be backwards?)
That said, it is 'portable' since in general no compilers reorder the struct members, but container_of is guaranteed to work, works for multiple members in a struct exactly the same way, and will reduce to a simple cast when your offsetof is zero so there isn't any performance penalty in situations like this. IMO there's no reason not to use container_of in this case.
To quote the relevant section (C99 6.7.2.1.13, same wording present in C89):
Within a structure object, the non-bit-field members and the units in which bit-fields reside have addresses that increase in the order in which they are declared. A pointer to a structure object, suitably converted, points to its initial member (or if that member is a bit-field, then to the unit in which it resides), and vice versa. There may be unnamed padding within a structure object, but not at its beginning.
I'll repeat that I consider direct casts to be somewhat idiomatic. Macros are more traditional for offset members.
Or put differently, direct casting comes with an assumption, while container_of doesn't. The fewer obscure assumptions are there in the code, the better. I think we can all agree on that :)
ucontext is unbearably slow, to the point where it's just as fast to use real threads and mutexes, even on a single-core system. There is nothing lightweight about it. (The technical reason is because they call into kernel functions to perform their magic.)
The actual logic of saving/restoring registers and the stack frame requires about 5-15 instructions per platform, and is very easy to write in assembly. And for the platforms you don't do this for, there's a non-standard trick to modifying jmpbuf which works on x86/amd64, ppc32/64, arm, mips, sparc, etc. These techniques are literally hundreds of times faster.
Try libco instead. It's the bare minimum four functions needed for cooperative threading, which lets you easily build up all the other stuff these libraries provide, if and only if you want them. And if you benchmark libco against all of the other stack-backed coroutine libraries, I'm sure you will be stunned at just how bad ucontext really is. That people keep using ucontext in their libraries tells me that they have never actually used cooperative threading for any serious workloads.
That said, your library is quite nice in that it just provides the co-routine wrappers and nothing else, which I appreciate. Also, your switch code has less instructions (are there cases it doesn't handle?), but if you look at libtask's code, it probably not going to be 100x faster. :-)
static void
contextswitch(Context *from, Context *to)
{
if(swapcontext(&from->uc, &to->uc) < 0){
fprint(2, "swapcontext failed: %r\n");
assert(0);
}
}
But it looks like you're right, context.c is actually defining its own version of swapcontext to replace the system function that can use other backends. I'm sorry for the mistake, thanks for correcting me.> are there cases it doesn't handle?
No, it conforms to the platform ABIs. I even back up the xmm6-15 registers that Microsoft made non-volatile on their amd64 ABI (compare to how much lighter the SystemV amd64 ABI's switch routine is; the pre-assembled versions are in the doc/ folder.)
Their own fibers library even has a switch to choose whether to back those up or not, which I think is quite dangerous.
Still, nothing can top SPARC's register windows for being outright hostile to the idea of context switching.
...
Oh, and it also supports thread-local storage for safe use with multiple real threads.
> it probably not going to be 100x faster.
No, definitely not compared to his ASM versions. Even with a less efficient swap, once you get past the syscall, most of the overhead is simply in the cache impact of swapping out the stack pointer, so his will likely be very close. In fact, even I had to sacrifice a tiny bit of speed to wrap the fastcall calling convention and execute non-inline assembly code (to avoid dependencies on specific compilers/assemblers.)
I do still strongly favor jmpbuf over ucontext for the final fallback, but with x86, ppc, arm, mips and sparc on Windows, OS X, Linux and BSD, you've pretty much covered 99.999% of hardware you'd ever use this on. That and libtask lacks Windows and OS X support.
Although, as far as I recall Go does map a bundle of them to a thread and can thus handle parallelism as well.
Just to clarify, this is not a complaint, I am trying to get a better idea of these abstractions.
@tylertreat
> I believe you're confusing goroutines with channels. Goroutines are lightweight threads of execution. Channels are the pipes that connect goroutines.
That is indeed correct. I was thinking more in the line of fibres that meld the concepts of a coroutine and a channel into a single abstraction. So the idea is Chan connects threads rather than coroutines.
The feature is called "dthread" or "duda-thread" which aims to expose Channels and generic co-routines functionalities, you can check the following relevant links:
a. The co-routine implementation:
https://github.com/monkey/duda/blob/master/src/duda_dthread....
it does not need mutex or anything similar, short code but a lot of thinking behind it.
b. Example: web service to calculate Fibonacci using channels:
https://github.com/monkey/duda-examples/blob/master/080_dthr...
the whole work will be available on our next stable release, if someone want to give it a try, just drop us a line.
[0] http://blog-swpd.rhcloud.com
[2] http://www.google-melange.com
[3] http://duda.io
See example on https://github.com/jorisvink/kore/tree/master/examples/tasks
pthread_cond_destroy(&chan->m_cond);
After https://github.com/tylertreat/chan/blob/master/src/chan.c#L1...That is not "pure C". "pthread.h" (POSIX threads) is not part of the C language standard, it is a system-specific header commonly found on UNIX systems. Windows has a different threading system for example.