And to be more precise, they still use some core libraries from ROCm stack[3], they just don't use all these fancy multi-gigabyte[4] hardware-limited rocBLAS/hipBLASlt/rocWMMA/rocRAND/etc. libraries.
[1] https://tinygrad.org/#tinybox
[2] https://github.com/pytorch/pytorch/issues/119081
[3] https://github.com/tinygrad/tinygrad/blob/v0.9.0/tinygrad/ru...
There are multiple bounties just for it in https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic...
It would be nice to see less whining and blaming AMD (PyTorch and llm.c actually work on 7900 XTX, and blow tiny grad out of the water in terms of perf!), and more just getting stuff to work.
[1] https://github.com/anthonix/llm.c [2] https://github.com/tinygrad/tinygrad/issues/4301
I have seen last month getting a lot of work done in improving performance (it's in the release announcement as well), but of course I still don't think it can compete with that number...still, a new comparision would be cool.
And still no comment on the issue, will re-run if there is any comment.
But this is interesting and probably strong evidence that the CUDA API isn't the moat people thought it was. CUDA multiplies matricies and that is close to a commodity operation. The moat actually seems to be Nvidia's higher generic software engineering standards, the difficulty in writing job scheduling/memory management infrastructure and possibly the fact that closed firmware is the norm.
Sorry for no direct link, but he has so many and very long videos that it is hard to find the exact spot.