How is this different from superoptimisation?
Also, how do you ensure that newly generated kernels are correct w.r.t. the original naive kernel that you use as specification?
Also, how do you ensure that newly generated kernels are correct w.r.t. the original naive kernel that you use as specification?
the search space is designed to remain logically equivalent at all times, by virtue of how its built (applying rewrite rules we know dont change the logical equivalence).
all the optimizations for matmul so far have been straightforward trajectories from naive (tiling, smem caching, tensor core offload, etc.)
https://cacm.acm.org/research/stochastic-program-optimizatio...