- Halide
- Tiramisu
- Ngraph
- Weld
- PyTorch JIT & PyTorch Glow and Tensor Comprehensions
- TVM
- Futhark
- A toy MLIR language build on the linalg dialect. I'm sure Google would love to help them.
to name the main ones.
Everyone is moving towards a compiler approach.
I'm actively researching the domain and I think Halide is the most promising: by separating the algorithm from the schedule you can leave the algorithm side to the researchers and optimization is done by HPC devs or automated methods without having to rewrite in C/Fortran.
Much more attempts in the past:
- https://github.com/mratsim/Arraymancer/issues/347#issuecomme...
- My own domain specific language and compiler: https://github.com/numforge/laser/tree/master/laser/lux_comp...
there are a low_level_design_considerations.md and challenges.md that should be interesting as well.
Unfortunately I didn't go beyond the proof of concept as I need to revamp the Nim multithreading runtime to improve composition of parallel linear algebra kernels.One thing for sure is that OpenMP is today something that needs to be workaround due to it's lack of composability and GCC implementation is a contention point due to using a single queue protected by a lock. While traditional linear algebra algorithms can deal with static distribution and no load balancing, some algorithms would hugely benefits from a more flexible task system, for example Beam Search.