The area of e-graph optimizers seems well-suited to this, btw. It's not really deployed outside of some niche tooling though, as it's a big paradigm shift in optimizer pass handling (e.g., doesn't work well with chairs classic call graphs, so control flow needs to be massively revamped to deploy e-graphs outside/across basic blocks and for loops (break and return not supported!)).
I would like to understand why you say e-graph would need control-flow to be revamped.
Do you have anything I could read on it ?
Happy to chat more btw, feel free to hit me up on discord.
I'm not sure what the state of the art in compiler optimisation is with regard to data positioning and targeting maximum processor usage
There was a video on optimisation a while back that showed small optimisations caused increases in speed that were insignificant when compared to the speed variance induced by the memory layout that the optimisation (or even a random change) caused.
While that talk was more focused on getting a signal past the noise. That noise itself is an artifact of compilers being not particularly good at handling a much simpler form of the problem you describe.
CPU and memory architectures are complex when caches and access patterns impact upon speed.
When you add in GPU architectures to the mix I think you might be in fairly uncharted territory.
Maybe one day.
Of course since we are in the field of AI there is also the question of could a sufficiently smart AI do this. It depends on the value of sufficient.
I would like to think that an extremely high level test for an AI model could be to give it something like micrograd and tell it to produce something with the same interface that outperforms torch.
We're not even in the ballpark of being able to do that yet, but it will be interesting when and if that happens.
> Seems like TVM
Fair enough, though technically they are still about different things but it's indeed very close, but
> and tinygrad
?????? what gives you this impression?
[0] so you can avoid materializing intermediate matrices and still being able to compute in blocks.