With modern computers I wonder if explicit parallelism is more fundamental to what our computer science will be than is in vogue in textbooks. Perhaps we should always be writing explicitly parallel code at this point.
With modern computers I wonder if explicit parallelism is more fundamental to what our computer science will be than is in vogue in textbooks. Perhaps we should always be writing explicitly parallel code at this point.
eg `for` loops are being replaced by `foreach` loops,`map` and `filter` operations, etc. These tell the compiler/interpreter that you want to do some operation to all the items in your datastructure, leaving it up to the compiler/runtime whether and how to parallelize the work for you.
I've thought this way ever since MacOS added the Grand Central Dispatch [1]. Of course, I thought industry would follow quickly and that tooling would coalesce around this concept pretty quickly. Seems the industry wants to take its sweet time.
For basic parallelism, nothing beats OpenMP for ease of adapting existing code (often a single "#pragma omp parallel for" directive is enough). Even for more complex parallelism, particularly where per-thread resources need to be managed, OpenMP still provides a much simpler programming model than the alternatives.
I'd also say that most languages have something similar to OpenMP, parallel for loops, etc. Great if you have some read only data in arrays and wish to process it.
However, in my opinion, it doesn't really matter how convenient a parallel / async programming model is to use as the real work is ensuring that there isn't any shared mutable state being updated in parallel. The other issue is, once you have formulated / re-formulated a particular problem to this model, ensuring that it remains this way is pretty challenging on larger teams. Someone can easily unknowingly commit something that breaks such assumptions.
Parallelism introduces an additional class of bugs, but they are fundamentally addressed the same way as any other class of bugs - e.g. testing, tools, and code review. If some_one_ can unknowingly break a system, that means the tools and processes weren't good enough.
Also, the code introducing a race condition may get lucky when your CI system runs the tests and still make it into your main branch.
I agree that tooling (like static analysis, Rust's borrow-checker, etc) can play a big role here though.
Particularly for dynamic analysis, you need to have test cases that usefully cover the design behavior. E.g. if you design a component to be safely shared, you need tests that exercise that sharing where the static/dynamic analyzer(s) will identify unsafe sharing. Likewise, if you know something is unsafe, you should probably have tests that demonstrate that the static/dynamic analyzer(s) do detect the unsafe usage.
My HPC programming lectures were done on PVM, and I bet only grey beards know what it stands for.
We're inching towards an in vogue way to do what erlang had figured out in the 80's. We'll pick up the pace any day now. Surely.
There's difference between doing it in order 1, 2, 3 and 3, 1, 2.
foreach will not be replaced behind the scenes into multithreaded version since it changes behaviour.
for is replaced with foreach because usually you dont need index and foreach is just handier and safer, that's it.
.NET's std lib has Parallel.ForEach for such a thing.
We really don't need magic to write multithreaded code. All we need is just really, really well designed APIs and primitives.
It only (meaningfully) changes behavior if you're both iterating over an odered datastructure and the body of your loop has direct or indirect side-effects. (like printing, writing to a file, making network requests, etc)
So like... huge % of the real world code bases
Right, and not always even then, because that depends on what the consumer is concerned with as well. But the fact that it can means it's not a safe automatic substitution.
> Humans are bad at reasoning about multiple threads simultaneously
I am not so sure this is true, I do believe that people are poorly practiced. My experiences have led me to believe Universities silo explicit parallel programming too much. It's generally it's own non-compulsory subject in a Comp-Sci major.
How this has played out in my life gives me caution about making this standard in computing.
Only now I can enjoy in modern hardware what I had to imagine when reading papers about Star Lisp and the Connection Machine, alongside other similar approaches.
Every statement in a HDL language runs in parallel but you can still write implicitly sequential code in VHDL processes.
Humans are bad at reasoning about way too many things. I think mostly because many are lazy and do not want to learn. The ones who do have little problems. I do not find thread management particularly hard for the most parts (there are some exceptions but those are very uncommon).
> Reordering of reads and writes can be done both by the compiler and by the processor. Compilers and processors have done this reordering for years, but on single-processor machines it was less of an issue.
https://learn.microsoft.com/en-us/windows/win32/dxtecharts/l...
This is why we have things like WaitForSingleObject and many other that deal properly with the chance of reordering and other concurrency related issues. All is fine with the reasoning on CPU, OS, Compiler and my own level. One just have to understand what is going on and know the tools. Those who are setting boolean flag to indicate the data is ready should not be programming for modern CPU's and have a basic knowledge first.
But ensuring some parallelism while maintaining thread safety is straightforward in many contexts - an uncontended mutex is close to zero overhead. Languages like Rust make it harder to struggle with some of the thornier issues (data races which cannot be simulated by stepping threads).
E.g. in Java a typical system looks like a thread per request with some shared underlying data structures like caches or connection pools. Relatively easy to use these safely, or to guard some shared object with a synchronized.
Likewise with parallelism - a lot of problems just boil down to ‘do a few map reduces’ and the parallelism is pretty trivial.
Obviously, concurrent systems are fiendish to reason through - but there are a lot of cases where the complexity can be side stepped. Doesn’t seem to stop people writing scary code on the daily though.
For things like running a web service, requests are fast enough, and the real win from parallelism is in handling lots of requests side-by-side. This is where No-GIL comes in.
Within handling a single request, if there are a lot of sub-requests, that's usually handled by async code, but not so much for the async performance win as much as spinning up threads is either expensive or thread pools are a hassle. Remember that async is better for throughput, but worse for latency, and if you're parallelizing a service request, you're probably more worried about latency. Async won mostly on ergonomics.
The other place you see parallelism is large offline jobs. Things like Map-reduce and Presto. Those tend to look like divide-and-conquer problems. GPU model training looks something like this.
What never happened is local, highly parallel algorithms. For a web service, data size is too small to see a latency win, they're complicated, and coordination between threads become costly. The small exceptions are vectorized algorithms, but these run one one core, so there isn't coordination overhead, and online inference, but again, this is heavily vectorized.
GPUs maybe? Also, excellent answer.
And there are things even further afield like what I work on [1][2][3].
I noticed the Regent code is inside the Legion repo. Is Legion the system, and Regent the language?
Can Legion be used without Regent, or vice versa?
Regent is a programming language. The compiler for Regent generates Legion code. Semantically, Regent is mostly a simplification of Legion. There are fewer moving pieces, so fewer things you need to worry about. Many of the "gotchas" that exist in Legion are taken care of by the language/compiler so idiomatic code usually "just works". It also does GPU code generation for you so you don't need to hand-write CUDA/HIP/etc. The tradeoff is that you're using a new programming language, so you have to be willing to take that risk.
"Safely" for a certain kind of definition of safety: https://github.com/search?q=repo%3Arayon-rs%2Frayon+unsafe&t...
When people say Rust is better for memory safety than e.g. C++, it's not because you can't write a provably-safe library with C++—you can do that with some effort. It's because the Rust language differentiates compiler-asserted safety from programmer-asserted safety via the `unsafe` keyword, an opt-in.
[1]: https://rust-lang.github.io/unsafe-code-guidelines/glossary....
(edit: Sniped, but I believe I've expanded upon the sibling comment.)
> ... would be disingenuous, or at least very ignorant.
That's quite literally what it means - there is a "safe" Rust code that never uses the "unsafe" but the example given by the parent comment is just not the example of that. I am not sure why would my comment come out as ignorant - it's factual state of things.
> it's not because you can't write a provably-safe library with C++
Hm, you really can't? Neither you can with safe or unsafe Rust. None of the compilers for those languages are formally verified.
It's an implementation detail, and if we can abstract it away to make it easier to utilize, we should.
The great thing about the disruptor it is that multiple threads can receive the same message without much more effort.
https://github.com/LMAX-Exchange/disruptor/wiki/Performance-...
https://gist.github.com/rmacy/2879257
I am dreaming of language that is similar to Smalltalk that stays single threaded until it makes sense to parallise.
I am looking for problems for parallelism that are not big data. Parallelism is like adding more cars to the road rather than increasing the speed of the car. But what does a desktop or mobile user need to do locally that could take advantage of the mathematical power of a computer? I'm still searching.
I am thoughtful of the Itanium and VLIW architecture for parallelism ideas.
The things we currently let servers do but it would mean we can keep user data local and not hand it over to service providers. I believe that is a worthy end goal.
Multithreading doesn’t always have to be around increasing speed, it can also reduce power
If you can find enough independent sequential problems in your programs then you can easily fill up cores, mostly because we don't have that many. I only have eight.
The problem is that this requires additional graph processing and there it makes your programs slower, which kind of defeats the point. The goal is to find the right tradeoff.
There's interesting stuff going on in the VHLL world with languages like Futhark, Jax, Mojo, etc that would be a better peer group for Python and its high level of abstraction.