Very sad, shouldve used an agnostic framework instead of CUDA
Although I wonder if it would work well with GCC PTX OMP offloading.
"Currently, I am working on [...] direct CUDA implementation, which will be significantly faster and probably come close to PyTorch."
The most interesting one IMO is OLMo from AI2, which is truly open. You can read their blog post about it (https://blog.allenai.org/hello-olmo-a-truly-open-llm-43f7e73...) but basically it is open everything - they released everything you need to reproduce their weights (training data, training code, evaluation code, and weights) with a friendly (Apache) license.