Source: was involved in development of multiple HPC frameworks, attended multiple HPC conferences, despite wailing and gnashing of teeth all actual software continues to be written the same way it's always been.
Source: was involved in development of multiple HPC frameworks, attended multiple HPC conferences, despite wailing and gnashing of teeth all actual software continues to be written the same way it's always been.
Hats off for making a good shot. But don't be surprised to see reluctance to move.
(See e.g. the Mario 64 decompilation — https://github.com/n64decomp/sm64 — , whose reverse-engineering process at all times kept a source tree that could build the original byte-identical ROMs of the games, despite being increasingly [manually] rewritten in an HLL.)
C lets you control memory alignment, in a way that very few other languages do. Until the competitors in HPC catch up to that, C is going to continue to outrun those other languages.
(Now, for doing the tweaking to get the most out of the algorithm, do I want to do that in C? No. I want to do it in something that makes me care about fewer details while I experiment.)
C isn't the Lingua Franca of fast software anymore. The reason why C programs can sometimes be faster, today, however is not an intrinsic property of the language but rather its inability to allow incompetent programmers to hide their bad data structures.
C also encourages programming styles that are harder to optimize so some loop optimizations are no longer possible.
Strong disagree. I don't see any reason why a competent programmer would hand-code a linked list, which is more of a hassle to make than a simple array (e.g. using buffer pointer + capacity and/or length field), unless the linked list makes actual sense for performance.
Try writing a type that can automatically switch layout from AOS to SOA in C, you just can't do it.
(you could use a macro to encapsulate the layout though, or in C++ could code a function that returns a reference)
In what situations are linked lists the path of least friction??
C is a bad programming language in 2022
Anything fast is in C++ or C++ like languages.
The only language I've actually done it in, is D. It's probably doable in many other nu-C languages these days, but D at very least can make it basically seamless as long as you do some try-and-break-shit testing to make sure nothing is relying on saving pointers when they shouldn't. This obviously constrains the definition of automatic ;)
I don't have my implementation to hand because it grew out of patch that failed due to aforementioned pointer-saving in code that I'm not paid enough to refactor (), but here's one someone else made https://github.com/nordlow/phobos-next/blob/master/src/nxt/s... there's another one in that repository too. I've never used those particular implementations but they're both by people I know so hopefully they're not too bad.
A more subtle thing, which I haven't used in anger, but would like to try at some point is to use programmer annotations (probably in the form of user defined attributes) to try and group things so things are stored such that temporal locality <=> spacial locality, but I've never bothered to actually do it.
There are some arrays of structs in an old bit of the D compiler that are roughly the size of a cacheline, and aren't accessed particularly uniformly. I profiled this and found that something like 75% of all LLC misses (hitting DRAM) were due to 2 particularly miserable lines... inside an O(n^2) algorithm.
struct Point { float x, y, z; };
soa<8> Point pts[...];
for (int i = 0; i < 8; ++i) {
pts[i].x + pts[i].y
}
behaves as if originally written like struct Point { float x[8], y[8], z[8]; };
Point pts;
for (int i = 0; i < 8; ++i) {
pts.x[i] + pts.y[i];
}
See https://ispc.github.io/ispc.html#structure-of-array-typesI've not yet had an excuse to use ISPC myself, unfortunately :(
https://zig.news/kristoff/struct-of-arrays-soa-in-zig-easy-i...
What is it, then?
> The reason why C programs can sometimes be faster, today, however is not an intrinsic property of the language but rather its inability to allow incompetent programmers to hide their bad data structures.
I'm not sure I manage to parse that properly. What's wrong about not being able to hide bad data structures? And how is this making C faster?
Just a random example, I wrote a two lines Python function using Numba jit, because I know in my mind that some Python behavior is going to make it less performant and it is a high performance kernel so we need that fast (and also lower memory footprint.) Compiling using Numba jit is no brainer because I just add one more line there with like 10 seconds effort.
But the code base I am merging to has a policy against Numba (reasonably as we are targeting HPC platform where Numba has some performance problem related to oversubscribing if not setting up carefully.)
So I end up rewrote that in C++ and wrap it with pybind11. And the result is that it is faster by around 30%.
Since the algorithm is entirely trivial, the only explanation I have is exactly memory alignment, where I can control that in C++, but in Numba jit there's no way (both to guarantee the array allocated is aligned, or tell the compiler to assume that.)
(The 30% number is also reasonable in textbook examples.)
P.S. But it has a cost:
> Showing 14 changed files with 230 additions and 36 deletions. > from https://github.com/hpc4cmb/toast/commit/a38d1d6dbcc97001a1ad...
where the 36 deletions are mostly just documentations. So it's an ~200 lines effort comparing to a 4 line effort...
But you're right: today, MPI is dominant. I suspect this will change, if only because HPC and data center computing is converging, and the (much larger) data center side of things is very much not based on MPI (e.g., see TensorFlow). Personally, I find it reassuring that many of the dominant solutions from the data center side have more in common with the upstart HPC solutions, than they do with the HPC status quo. I'd like to think that at some point we'll be able to bring around the rest of the HPC community too.
Using Python on HPC system, depending on who you ask, is kind of cheating as the main work is always C/C++ (or Fortran). So it is still really just C/C++ and Fortran.
Julia is really promising as it is the first new language that can efficiently run on HPC and can use other form of message passing.
Also we should not forget UPC, a special kind of C that has its own parallel programming model (not MPI.) It has been used on HPC system as well but perhaps only from a selected few. (This probably reveal where I come from.)
Real work was always C++ and Fortran.
In terms of scripting, a professor once showed us a workflow script from his old days written entirely in bash and is very very long and very complex. He basically said if your script gets that complicated it should be written in Python nowadays.
But Python’s HPC presence is deeper than scripting in that sense. Machine learning frameworks definitely is a contributing factor, but even more traditional HPC workflows are benefited when stuffs have Python API, especially with mpi4py.
I once used something called mpi bash that I need to compile it myself on an HPC platform. It is much less well known and used, perhaps because of what the professor said—if it needs MPI, probably even for scripting it should be scripted in Python instead.
I even suppose that the orchestration layer can be in basically any convenient language, even in Python, while the underlying compute modules may be written in highly-optimized compiled languages; see Thensorflow and PyTorch for hihgly successful examples.
Obviously it's not very performant, but the lessons learned on how to program with MPI in mind is where the value lies.