Show HN: uThreads – Concurrent User Threads in C and C++
samanbarghi.com
samanbarghi.com
Unsolicited suggestion - while benchmarks and the motivation are important for a threading library, a code snippet of a simple parallel program on the home page would be something that I'd love to see.
Great job, though!
There is no fork-and-join in uThreads yet, and I created it using create and join. The interface will improve in the future :)
uThreads is still a work in progress and I have specific plans for it in the future that might differ of what Qthreads is trying to accomplish. For now my focus is more on providing auto tuning of Clusters based on the workload. I also will try to explore the pros and cons of uThread migration based on Cluster placement (NUMA and cache locality), and from the 2008 paper it seems that it is what you are trying to study as well. I am open to collaboration if there is an ongoing study around this topic.
Based on your auto-tuning discussion...RaftLib aims to do something similar, but for big-data stream processing workloads. Here's my 2016 IJHPCA paper: http://hpc.sagepub.com/content/early/2016/10/18/109434201667...
It looks like it's behind a paywall so if you don't have access I'll update my website with the "author archive" copy sometime today...will be at ( http://jonathanbeard.io/media ) once I update it. Bottom line if there's intersected interest, definitely open to collaboration :).
Sorry if this sounds snarky, and I do not blame you in particular, but it is just a theme I have encountered here the last years that I find toxic as it portrays the GPL (and FSF) as some evil organization that would restrict "developer's rights".
"Hi, I'd like to use your library, but your license is incompatible with the other <N> licenses in my codebase. Consider BSD please?"
JFK
GPL is by definition more restrictive than Apache and BSD, since there are more requirements to use the code, so asking if the developer is willing to consider a less restrictive (not "better") license shouldn't be met with hostility.
My motivation for asking is simply because the project looks interesting and potentially useful; however, the proprietary nature of my current work means that GPL licensed code isn't really an option for me. I absolutely wouldn't hold it against the author if they chose GPL for this or any other reason.
I don't GPL license a lot of my software (I used BSD/MIT/Apache most of the time, and my company revolves around and writes BSD-licensed code), but IMO, having seen this tired argument over and over again, I'm pretty sure 90% of all GPL complaints come down to this, even if the people don't come out and say it: "I can't make money off of your free work as easily, and that isn't fair to me. Please reconsider." Given this is Hacker News where half of everyone is in a rat-race to make money, I speculate this is a strong part of it.
People just like to wrap it in words like "restrictive" and "viral" to make themselves sound more palatable and reasonable.
And of course, there are also many reasonable alternatives to this library, many which are BSD or permissively licensed, which these people could also use instead -- but that won't stop them from complaining that GPL is unfair or "bad", of course, even though they could pick from a dozen alternatives...
They do the same shit with the GPLv3 as well. Piko Interactive recently backed out of a licensing agreement with me, and just released my GPLv3 emulator in their Steam application without even telling me. When someone called them on it, they said to e-mail them for a link to the code (not enough to take my work for free, they have to play games with their obligations under the GPL.) Which by itself is useless, as it's just a UI modification. The value is the ROM image they don't include in their source, which gets you into a nasty gray area of the GPL.
It's part of the Faustian bargain all open source devs have to make: if you add a non-commercial clause, FOSS proponents will label your work "non-free" and "not open source", and you'll be banished to obscure disabled-by-default nonfree repositories on Linux distros. Which I was until I caved and moved to GPL to avoid punishing my users.
If you don't add that clause, you'll get taken advantage of. We have to rely on people being fair and sharing their profits off our work, and very often, they don't.
If someone's licensing an entire application that's usable in its own right, like an SQL server or emulator or game or kernel, then it totally makes sense that anything you build around it should have its code be accessible.
But if it's some smaller part of your entire codebase, I feel it's a little less reasonable for people who aren't working on something that's already *GPL licensed. Especially when you have middlewares or other proprietary pieces of code linked with yours. That, plainly and simply, will prevent you from using the library. (Except if it's the LGPL, which I've never had problems with. I personally like it the most, and I'd use it if I really cared about people upstreaming their changes.)
I'm not going to say "this library is bad because it's GPL'd" but it does mean I would avoid baking it into a programming language runtime, for example.
And it really bothers me that you have such a low opinion of people who comment here, and of people who disagree with you. Maybe it's just an exaggeration for argument's sake, but I feel like there are plenty of reasonable arguments against use of the GPL.
And speaking as someone who writes software for a non-profit, we figure we should just give away the code we write rather than put our own political license on it, especially considering we don't pay taxes in order to make it easier for us to benefit "the people".
http://libdill.org/tutorial.html
See, in particular, Step 6 of the tutorial.
Using uThreads you can decide what part of the code should be executed over each kernel thread and, if taskset is used, which core to execute your code. You can do this by creating Clusters and using migration or uThread creating in runtime. Thus you can decide which thread is used to multiplex connections and which one is used to run CPU bound code in addition to having a thread pool for example to do the disk IO asynchronously. Ultimately, one can create a SEDA[1] like architecture using uTrheads. Also you can always use uThread as a single threaded application.
------------------------------------------
[1] https://en.wikipedia.org/wiki/Staged_event-driven_architectu...
Any thoughts on Welsh's Retrospective on SEDA?
http://matt-welsh.blogspot.com/2010/07/retrospective-on-seda...
Does anyone know how it compares?
For sure at some point the poller thread will be saturated and the program will not scale past a certain number of threads. I used to have a poller thread per cluster for better scalability, but that would add overhead for migrations between clusters, thus I had to remove it for now until I can somehow find a low overhead solution. uThreads is a work in progress and all these need to be carefully considered in the future :) Thanks for your feedback
This sounds like the Linux kernel. I'm curious to understand why copying this logic into user space is worth while?
Run queues can be part of any scheduler, they are queues with runnable tasks.
But as why the approach is worth while in user space, it has to do with low cost of operations (context switches) in user space, and also using cooperative scheduling instead of preemptive. Cooperative scheduling provides more control to the user over the tasks, and also has lower overhead since there is no need to manage a quantum for each taks (thread).
8KiB stacks are a bit on the small side though for production usage. Go gets away with this because they're stacks act more like Vectors then flat arrays.
Why did you decided to roll your own stack swapping software instead of using say using `boost::context`?
---------------------------------
[1]https://github.com/samanbarghi/uThreads#what-are-uthreads
For all N:1 mappings, since there is only a single process, there is no need for synchronization, also there is no need for any scheduler as a simple queue suffice. But as soon as multiple threads are introduced, there is a need for orchestration among threads and it also changes all other aspects of the code. Of course, I could develop on top of an existing codebase, but I suspect I had to change so much that it is better to start from scratch anyway.
------------------------------------
Wouldn't an M:N model simply amount to work stealing among N kernel-thread-local queues? This seems like it should be a pretty straightforward extension to one of the user-level C thread packages.
Or are you doing something more elaborate for your research?
Usually these libraries either provide their own queue or rely on underlying event system e.g. epoll/kqueue to manage the state of the light weight threads. Also, the other main part is IO multiplexing, which can be done using select/poll/epoll/kqueue ...
Lets say if they are using some sort of queue, since there is no other thread in the system, thus no synchronization required and its as simple as pushing to the tail of queue and pulling from the head. Now, if I add more threads, now I have to think how I want to synchronize among threads. The straight forward way of doing this is using the same queue and multiplex among kernel threads and use Mutex and CV. However, Mutex has high overhead and is not very scalable (pthread_mutex does not scale well under contention in Linux). What about N-kernel-thread-local queues along with Mutex? Well, if you see the documentation this approach does better but has high overhead as well. What is next? remove the queue and add some sort of lock-free queue, either MPMC, MPSC, or SPSC. This requires some work to determine which one has lower overhead, in my case I have not tested SPSC, and for now settled with MPSC. So the queue part is totally gone, since they probably did not care about all this and used a simple queue.
Next, comes the IO multiplexing. Relying on an epoll instance per kernel-thread is absolutely fine specially for cases where uThreads stick to the same kernel thread, but as soon as I introduce the notion of migration then moving from on kernel-thread to another kenrel-thread means issuing at least two system calls (in case of epoll, deregister from current epoll instance and register to the new one). This has high overhead, so I need to provide some sort of common poller that do multiplexing over multiple kernel threads, but due to having more than one kernel thread, it means connections should be properly synchronized as more than one thread might try to access a connection. Also, it has to be done in a way with low overhead as migrations for my research require to have very low overhead. Thus, the IO multiplexing part should be replaced.
So the main parts are required to be changed, and I believe the effort required to make fundamental changes to an existing system might be more than the efforts required for writing it from scratch. Also for each part implemented, I did performance optimisations and build on top of that which helped to keep the performance at an acceptable level, it would be hard to do the same with an existing system as it requires to isolate various parts and optimise each part, which requires additional effort.
I hope it makes it more clear :)
I was also assuming a single kernel thread performs I/O via epoll/kqueue/etc. and either has its own queue from which other threads steal, or simply pushes results onto a random queue when requested I/O are complete. This accomplishes the migration I believe you were describing.
When I/O needs to be done, you enqueue the uthread on the I/O queue and invoke a reserved file descriptor to notify the I/O thread, which then reshuffles its file descriptors and again calls epoll/kqueue.
I'm not sure whether this would scale as well as what you're doing since you mentioned an epoll-per-kernel thread, but I wouldn't be surprised if it got close since it's so I/O-bound.
> When I/O needs to be done, you enqueue the uthread on the I/O queue and invoke a reserved file descriptor to notify the I/O thread, which then reshuffles its file descriptors and again calls epoll/kqueue.
If I remember correctly, what you described is similar to how golang perform I/O polling since they are using work stealing, except there is no dedicated poller thread, and threads take turn to do the polling whenever they become idle.
Also, with epoll/kqueue there is no need to enqueue uThreads and you can simply store a pointer(a struct that has reader/writer pointers to uThreads) with the epoll event that you create and recover it later from the notification, this way you can set the pointer when need to do I/O.
I agree a single poller thread does not scale as well as per kthread epoll instance, and that requires some meditation to provide a low overhead solution that does not sacrifice performance over scalability.