Most developers cannot do multithreading correctly, and unless you're particularly good about it it's just going to introduce not only lots of bugs but also performance problems.
The only folks in that space that seem to do it well are ScyllaDB.
Most developers cannot do multithreading correctly, and unless you're particularly good about it it's just going to introduce not only lots of bugs but also performance problems.
The only folks in that space that seem to do it well are ScyllaDB.
All it takes is one critical section to not be protected (i.e. locked) to cause a bug. A series of tests can run hundreds of times correctly without detecting the problem. It is only when a context switch happens at a certain microsecond that the error is exposed.
I am a true believer in multithreading as my own code can see tremendous performance gains using it on the latest multi-core CPUs; but tread very carefully when programming in this manner.
And sorry, but that is multithreading, there are several cores.
Native multi-threading is used when you have functionality that already works on threads and you don't want to port it.
Multi-thread is not used in the hot path.
A single data-part/shard is served by a single thread.
but how does this kind of multithreading (one thread per core) is better than proper multithreading (many threads per core)?
I’m not sure it’s valid to say that only SMT is “proper multithreading”, especially since multithreading as a concept predates it by quite a way.
SMT has a quite a few performance issues since resources such as the L1, L2, and branch predictor are shared between the threads, which can lead to contention that hurts the performance of all the SMT threads sharing a physical core.
SMP is no less “proper”, and as core counts have increased significantly on commodity CPUs, the use of spinning threads bound to a single core each has become a common paradigm.
Oversubscription without SMT (i.e. many threads per core) is possible, but unless you have a workload where each thread is I/O bound with a substantial amount of time spent blocking, the overhead of scheduling and context switching means throughput will likely decrease.
Of course it increases latency, since those resources are not fully exclusive to a particular thread anymore.
Whether or not it's a good thing depends on what you care about. You could also argue that a good program would be able to saturate a single superscalar core with a single thread and thus wouldn't benefit from SMT at all, but I think that would be hard to guarantee in practice.
Why is it an anti pattern, this is news to me?
And setting up, say, one thread per HTTP request will likely be negligible because blocking I/O is where time is spent anyways..
And we have had non-blocking I/O for quite some time now.
Your I/O should only be done synchronously if it's non-blocking.
Now for disk I/O, it's a more muddy thing, it's actually quite different from networking since it's more transparently managed by the operating system.
Userland threads (or fibers, or stackful coroutines) do scale better though.