I'm not sure if this is justified (e.g. concurrency is inherently too hard to be viable), or due to the lack of tooling/conventions/education.
I'm not sure if this is justified (e.g. concurrency is inherently too hard to be viable), or due to the lack of tooling/conventions/education.
Is it? There isn't going to be an official declaration from the Masters of Computer Science that "2019 was the year concurrency ceased being Too Hard." or anything.
My perception is that it is steadily becoming less and less notable for a program to be "concurrent". Tooling is becoming better. Common practices are becoming better. (In fact, you could arguably take common practices back to the 1990s and even with the tools of the day, tame concurrency. While the tooling had its issues too, I would assert the problem was more the practices than the tooling.) Understanding of how to use it reasonably safely is steadily spreading, and runtimes that make it safer yet are starting to get attention.
I'm not sure I've seen a case where there was a problem that ought to be using concurrency, but nobody involved could figure out any way to deal with it or was too afraid to open that door in a long time. There's still plenty of cases where it doesn't matter even now, of course, because one core is a lot of computing power on its own. But it seems to be that for everyone out there who would benefit from concurrency, they're mostly capable of using it nowadays. Not necessarily without issue, but that's an unfair bar; you can still get yourself in concurrency trouble in Haskell or Erlang, but it's a lot easier than it used to be to get it right.
The question is: by whom?
High performance software, like game engines, DAWs or video editors, has been heavily multithreaded for a while now.
Maybe it's consumer or business software that could profit from more multithreading? I don't know, because I don't work in those areas.
I do feel similarly, even though I wouldn't classify myself as a great engineer.
I've been writing concurrent software in managed languages, such as Java and C#[0], from the very beginning of my career, up until today. The level of multithreading has varied, but it's always been there. For anything beyond basic CRUD it pretty much becomes a requirement, both on desktop and on the web[1].
That doesn't mean I've never had a tricky race condition to debug (and, yes, they're hard to debug) during development, but I've never shipped a concurrency related bug to production[2].
The canonical examples of concurrency gone wrong are things like giving somebody a deadly radiation dose from a medical device but, in terms of serious software bugs, I do wonder how common concurrency bugs are relative to other types of bug, and whether they're really more serious in aggregate than those other types of bug.
[0] Admittedly these languages make it a lot easier to avoid shooting yourself in the foot than C and C++ do.
[1] Also worth bearing in mind that an inherent property of most, if not all, distributed software is that it's also concurrent: the moment you have multiple processes running indepedently or interdependently you also usually have a concurrent system, with the potential for "distributed race conditions" to occur. I.e., if you have a SPA that also has a non-trivial back-end, you have a concurrent system - just spread across different processes, generally on different machines.
[2] In the context of in-process concurrency. Distributed concurrency is a different matter.
For example: https://github.com/David-Haim/concurrencpp
I think it's the tooling. Rust's modelling of concurrency using it's type system (the Send and Sync traits) make concurrency pretty straightforward for most use cases. You still have to be super-careful when creating the core abstractions using unsafe code, but once you have them they can easily be shared as libraries and it's a compile error to violate the invariants. And this means that most projects will never have to write the hard parts themselves and get concurrency for close to free.
The benefits are that no other synchronization is needed than the data sent between processes, and race conditions are ruled out as long as only one process is allowed to process a data item at a time (this is the rule in FBP).
The main blockers I think is that it requires quite a rethink of the architecture of software. I see this rethink happening in larger, especially distributed systems, which are modeled a lot around these principles already, using systems such as Kafka and message queues to communicate, which more or less forces people to model computations around the data flow.
I think the same could happen inside monolithic applications too, with the right tooling. The concurrency primitives in Go are superbly suited to this in my experience, given that you work with the right paradigm, which I've been writing about before [1, 2], and started making a micro-unframework for [3] (though the latter one will be possible to make so much nicer after we get generics in Go).
But then, I also think there are some lessons to be learned about the right granularity for processes and data in the pipeline. Due to the overhead of message passing, it will not make sense performance-wise to use dataflow for the very finest-grain data.
Perhaps this in a sense parallels what we see with distributed computing, where there is a certain breaking point before which it isn't really worth it to go with distributed computing, because of all the overhead, both performance-wise and complexity-wise.
[1] https://blog.gopheracademy.com/composable-pipelines-pattern/
[2] https://blog.gopheracademy.com/advent-2015/composable-pipeli...
As to OP, well it better be viable, because we certainly need to deal with it. So better tooling and conventions encapsulated in expert developed libraries. The education level required will naturally fall into the categories for those who will develop the tools/libraries, and those that will use them.
Don’t OSs expose that in the sense you can pin a threads to closest cores according to the memory access you need?
Most of the computers in the world are either dedicated embedded controllers or end user devices. Concurrency in embedded controllers is pretty much an ordinary thing and has been since the days of the 6502/Z80/8080. For end user devices the kind of concurrency that matters to the end user is also not extraordinary, plenty of things happen in the background when one is browsing, word processing, listening to music, etc.
So that leaves concurrency inside applications and that just isn't something that affects most of the end users. There really isn't much for a word processor to actually do while the user is thinking about which key to press so it can do those few things that there was not time for during the keypress.
Mostly what is needed is more efficient code. Niklaus Wirth was complaining that code was getting slower more quickly than hardware was getting faster forty years in 1995 and it seems that he is still right.
See https://blog.frantovo.cz/s/1576/Niklaus%20Wirth%20-%20A%20Pl...
This obviously depends on the application. Whenever you need to wait for an app to finish an operation that is not related to I/O, there is some potential for improvement. If a CPU-bound operation makes you wait for more than, say, a minute, it's almost definitely a candidate for optimization. Whether multithreading is a good solution or not depends on each case - when you need to communicate/lock a lot, it might not make sense. A good part of the solution is figuring out if it makes sense, how to partition the work etc.; the other part of this hard work is implementing and debugging it.
- Current abstractions are mostly based on POSIX threads and C++/Java memory models. I think they are poorly representing what is actually happening in hardware. For example Acquire barrier in C++ makes it really hard to understand that in hardware it equals to a flush of invalidation queue of the core, to sanity check your understanding try answering the question "do you need memory barriers if you have multithreaded application (2+ threads) running a lock-free algorithm on a single core?", correct answer is no, because same core always sees it's own writes as they would happen in program order, even if OOO pipeline would reorder them. Or threads, they seem to be an entity that can either run or stop, but in hardware there are no threads, CPU just jumps to a different point in memory (albeit through rings 3->0->3). Heck even whole memory allocation story, we have generations of developers thinking about memory allocation and it's safety, yet hardware doesn't have that concept at all, memory range mapping concept would be the closest to what MMU actually does. Hence the impedance mismatch between hardware and current low level abstractions created a lot of developers who "know" how all of this works but doesn't actually know, and a bit afraid to crush through layers. I want more engineers to not be afraid and be comfortable with all low level bits even if they would never touch them, because one day you will find a bug like broken compare-exchange implementation in LLVM or similar.
- Tooling is way off, the main problem with multithreading is that it's all "in runtime" and dynamic, for example if I'm making a lockfree hashmap, the only way for me to get into the edgecases of my algorithm (like two threads trying to acquire same token or something) is to run a bruteforce test between multiple threads and wait until it actually happens. Bruteforce-test-development scales very poorly, and testing something like consensus algorithms for hundreds of threads is just a nightmare of complexity of test fixtures involved. Then you get into ok, so how much testing is enough? How do you measure coverage? Lines of code? Branches? Threads-per-line? When are you sure that your algorithm is correct? Don't get me wrong, I've seen it multiple times, simple 100 lines of code passing tons of reviews only for me to find a race condition (algorithmical one) half a year later, and now it's deployed everywhere and very costly to fix. Another way would be to skip all of that and start modeling your algorithms first, TLA+ is one of the better tools for that out there, prove that your model is correct, and then implement it. Using something like TLA+ can make your multithreading coding a breeze in any language.
And probably absence of transactional memory also contributes greatly, passing 8 or 16 bytes around atomically is easy, but try 24 or 32? Now you need to build out an insanely complicated algorithm that involves a lot of mathematics just to prove that it's correct.