What scientists must know about hardware to write fast code (2020)
biojulia.net
biojulia.net
http://web.archive.org/web/20221208105716/https://biojulia.n...
That said, my own experience of noticing what people's machines are wasting time on around me is very different.
For scientists at firms (outside of academia) a major bottleneck is downloading, decompressing and shuffling data around. Unfortunately the lessons in the piece are in fact detrimental to this workload: 1. decompression is usually not faster than 500MB/s, which is much slower than SSD or network, but I still see compression everywhere with the idea that "cpu is faster". 2. Standard tools for downloading data, like awscli, are crazy slow compared to the actual ability of the network cards that are attached to the machines we (in big firms) use in the cloud. Specifically awscli caps at about 250MB/s no matter the configuration or network bandwidth, and people think it's normal because "network io is slow". (this of course doesn't apply to scientists who work out of the cloud, with much lower download speeds)
None of this contradicts anything in the piece. It's that the units of work for different sectors are different. Whereas a geoscientist working on modeling the ocean has bottlenecks in computation, scientists at firms often have them in data.
There are some, but you have to select the right one. https://github.com/inikep/lzbench#benchmarks
Systems programming involves multiple systems with many components. It's not uncommon to find you simply can't make a solution work any better than it's designed, and you need to develop brand new solutions to problems (which makes the system even bigger).
System programming is simpler because it's one system and one or two codebases. It can still be quite complex, but you can often just optimize the one solution rather than having to come up with new ones.
Anecdotally when working with biological signals at my day job, compression is a massive win and an absolute no brainer tradeoff when shuffling data across the network or even just storing in memory.
That said, I think your first point is still reasonable for certain types of data and compression algorithms.
[1] https://github.com/lemire/SIMDCompressionAndIntersection
It starts kind of slow, but then somehow morphs into effortlessly explaining advanced topics. I knew bits and pieces of this before, but I found information that was genuinely new to me, like this gem:
”Going back to the original example, that is why the perfectly predicted copy_odds_branches! performs better than code_odds_branchless!. Even though the latter has fewer instructions, it has a memory dependency: The index of dst where the odd number gets stored to depends on the last loop iteration. So fewer instructions can be executed at a time compared to the former function, where the branch predictor predicts several iterations ahead and allow for the parallel computation of multiple iterations.”
Also, this document doesn’t load whatsoever when Javascript is disabled.
Such a great and condensed write up... and then at the last moment it looks like one is finally going to learn about (dedicated) GPU programming, and... nope ! :P
> 0.001321 seconds
> about 13 miliseconds
Uh no, that's about 1.3 milliseconds.
Hmmm? This part stands out. It sounds like the author is here. Can you help me understand this advice? It feels like it should have a million caveats and not be general advice. Often times memory mapping the file is going to be better. Unless you’re an expert you’re probably not going to do better than the page cache.
If you have a relatively simple access pattern, and most applications do, then you don't need to be a database kernel engineer to design a trivial manual buffering system that can significantly outperform mmap(). It is a handful of lines of code in most languages.
The caveat is that even though the code to do it is simple, you still have to understand how Linux storage and APIs work. As you correctly observe, naive attempts that don't understand the fundamentals can be significantly worse than just mmap()-ing things, like never actually removing the page cache from your I/O path (common mistake).
Because a page cache is expected to support all possible processes, it can neither overfit simple access pattern cases, because then it would offer poor performance for common cases that are more complex, and it will never have enough context to offer excellent performance for complex access patterns because doing so would reduce performance for many common simple cases. Real page caches try to straddle both worlds; it is better to be mediocre for simple and complex cases than to be excellent at only simple or only complex cases. Bypassing the page cache allows processes to design their own page caches that are highly optimized for their specific requirements, should it be worthwhile to do so.
A reason people might not want to do this is that the APIs are operating system specific.
What I mean is that, although the operating system caches files in RAM, and some (but not all) languages like Python and Julia implements a cache in its abstraction over file reading, these caches are not as fast as simply loading a bunch of bytes into an array and reading from the array, because they are hidden behind an abstracted IO, and the abstraction is not zero-overhead.
The difference is especially noticable when you read byte-by-byte, as in the issue I linked to in the text. For example, if you try to implement parsing with a finite state machine reading one byte at a time, you basically need to access memory directly without any overhead.
It's possible there are IO abstractions where reading from a "file" in RAM is just as fast as loading a byte from memory, but I don't know of any such abstractions.
So, if I want to multiply a bunch of integers together, does this mean it's faster for me to represent them as floats (even though I don't need the additional precision)?
Even if your CPU can do some instructions faster with floats it is generally not the kind of optimization you would want to make. You would end up making things more complicated for little practical benefit.
Funny thing by the way, integer arithmetic units are generally cheaper to make than float units, but most multiply/division heavy code is floating point, so CPU makers prioritise the floating point units.
https://gist.github.com/jboner/2841832
EDIT: That was an uncharitable take and I apologize to the author.
You know, 80% of the content of my blog post?