Making a parallel Rust workload 10x faster with (or without) Rayon
gendignoux.com
gendignoux.com
If your algorithm is the equivalent of a couple of nested iterations, you have essentially three options: parallelize outer, inner, or both. In the vast majority of the cases I've run into, you want thread/task level parallelism on the outer loop (only), and if required, data/simd parallelism on the inner loop(s).
It's a rule of thumb, but it biases towards batches of work assigned to CPUs for a decent amount of time, allowing cache locality and pipelining to kick in. That's even before SIMD.
Valid concern but I don't think this was the OP case though?
From my understanding, author gained the most benefits by dumbing down the generic rayon implementation to the same kind (thread-pool with task queues) but with different work-stealing algorithm.
> Rayon is not going to magically distribute your work perfectly, though it very often does a decent job.
Work-stealing by definition kinda makes distributing the work "correctly" a difficult task, doesn't it?
If every CPU is 100% utilized without needing context switch (and running the right number of worker threads without switching those), then work stealing is not required.
But my comment is solidly "Rule of thumb". I claim no theoretical basis other than "Giving fewer longer tasks to fewer threads, (still >= number of worker threads), is better than giving more shorter ones"
Works on Windows, has a great GUI, shows pretty flame charts, really pleasant to use. Costs money but worth it.
> So far, we’ve only used perf to record stack traces at a regular time interval. This is useful, but only scratching the surface.
For cache hits and other counters, you're gonna have to go deeper than just sampling.
Never heard of it and I do a lot of performance related work. Is Superluminal really that popular?
The "before" graph shows a ~300ms wall-clock time which drops to ~150ms parallelised.
The "after" graph shows a ~27ms wall-clock time which drops to ~17ms when parallelised.
Isn't all the improvement still outside of the parallelisation then? There's an awful lot of discussion about parallelisation and the behaviour of schedulers given there wasn't any improvement of how parallelisable the end result was?
is it? I've heard that about Go, not so much Rust. cargo doesn't statically link system dependencies by default, only other rust dependencies.
Go binaries are statically linked by default if they do not call out to any C functions.
Go will automatically switch to dynamic linking if you need to call out to C functions. There are some parts of the standard library that do this internally, and so many folks assume that go is dynamically linked by default.
Maybe it's been fixed since, but I wouldn't assume Go doesn't have any C dependencies.
> Now, you may wonder why profiling Rust code within Docker, an engine to run containerized applications. The main advantage of containers is that when configured properly, they provide isolation from the rest of your system, allowing to restrict access to resources (files, networking, memory, etc.).
https://gendignoux.com/blog/2019/11/09/profiling-rust-docker...
But a container does not provide a security boundary in the same sense as the security boundary between kernel mode and user mode.
You certainly can build a statically linked binary with musl libc (in many circumstances, at least), but it's not the default.
The default is a binary that statically links all the Rust crates you pull in, but dynamically links glibc, the vdso, and several other common dependencies. It's also, IIRC, the default to dynamically link many other common dependencies like openssl if you pull them in, although I think it's common for the crates that wrap them to offer static linking as an option.
I ran into this last week — I wanted to install the neovim text editor from the official binary distribution, but because the project’s CI distro has a more recent glibc installed, the executable would fail to run on Ubuntu 18.04.
There are plenty of great details spelled out here, in case you aren’t familiar with this kind of thing: https://stackoverflow.com/questions/57749127/how-can-i-speci...