I thought the README was long enough that I wasn't going to go into those details, especially as 1) I am not an expert on the topic, 2) the choice of curve depends on your application, and 3) there are a lot of resources on the topic. I listed a few.
If you want to learn more, see "Space-Filling Curves An Introduction with Applications in Scientific Computing", listed as one of my sources, eg, "Algorithm 13.3: Matrix-vector multiplication based on Hilbert order matrix traversal" and "Matrix Multiplication Using Peano Curves".
One of the sources I cited is "Improved Data Locality Using Morton-order Curve on the Example of LU Decomposition" at https://ieeexplore.ieee.org/document/9378385 :
> Here, we look at the block factorization algorithm, where the LU decomposition performance depends on the performance of the matrix multiplication. In both cases, the LU decomposition and the matrix multiplication, such a matrix is traversed by three nested loops. In this paper, we propose to traverse such loops in an order defined by a space-filling curve. This traversal dramatically improves data locality and offers effective exploitation of the memory hierarchy. Besides the canonical (or line-by-line) access pattern, we demonstrate the traversal in Hilbert-, Peano and Morton order. Our extensive experiments show that the Morton order (or Z -order) and the inverse Morton order (or И-order) have a better runtime performance compared to the others.
I first learned about this issue from "The Anatomy of High-Performance 2D Similarity Calculations" at https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4839782/ :
> Certain self-similar space-filling curves provide a natural hierarchy and iteration order to optimize locality in a cache-oblivious algorithm. In particular, Morton ordering, based on the hierarchical Z-order curve, has been previously used to optimize cache locality in matrix multiplication,22 suggesting that it may also be useful for similarity matrices. .... Other space-filling curves, such as the Hilbert curve, also have excellent cache-locality properties; however, the calculations required to implement the Z-order curve are particularly simple.
I want to reproduce their result, then see how other traversal orders affect the timings. Careful reading of that paper shows they only used sizes which are a power of two, while real-world data sets aren't so convenient.
I couldn't find an off-the-shelf library with non-powers-of-two/rectangular generalizations, so I wrote one.