Visualizing 6D Mesh Parallelism
main-horse.github.io
main-horse.github.io
In other words, i would like to explore a way to approach a problem of statically scheduling (heterogenous) computations and data transfers between multuple tiers of caches (disk, ram, vram) in OR fashion (e.g. formulating MILP).
What isn’t clear is how completely any of the COIN-OR projects address your rather interesting question; I’m commenting partly to share if the resource is helpful and partly as a reminder to come back to see what other smart folks here may contribute :-)
I am working on something similar, but for the Cerebras CS-3 engine, visualizing stuff like this really helps when designing algos for these architectures.