I wish they could have modeled memory cache movements as well, somehow. It would have made the analysis more difficult but often these large matrix algorithms live or die by how cache-friendly the memory access patterns are.
Splitting into 4x4 blocks is typically very nice, though. Maybe it doesn’t matter so much to practical runtime.