ThunderKittens: Simple, fast, and adorable AI kernels
hazyresearch.stanford.edu
hazyresearch.stanford.edu
torch.compile is a pt2.0 feature and has nothing to do with handwritten cuda kernels
> How easy is it to run on older GPUs
this is a torch cpp extension
https://github.com/HazyResearch/ThunderKittens/blob/8daffc9c...
so you're going to have the same exact issue (whatever issue you're having)
Any update on this?
It would be wild if some of these time-efficiency boosts you're getting with TK turned out to be energy-efficiency boosts too!
https://discord.com/login?redirect_to=%2Fchannels%2F11894982...