Tinygrad 0.9.0
github.com
github.com
But this is interesting and probably strong evidence that the CUDA API isn't the moat people thought it was. CUDA multiplies matricies and that is close to a commodity operation. The moat actually seems to be Nvidia's higher generic software engineering standards, the difficulty in writing job scheduling/memory management infrastructure and possibly the fact that closed firmware is the norm.
Sorry for no direct link, but he has so many and very long videos that it is hard to find the exact spot.
[1] https://github.com/anthonix/llm.c [2] https://github.com/tinygrad/tinygrad/issues/4301
I have seen last month getting a lot of work done in improving performance (it's in the release announcement as well), but of course I still don't think it can compete with that number...still, a new comparision would be cool.
And still no comment on the issue, will re-run if there is any comment.
And to be more precise, they still use some core libraries from ROCm stack[3], they just don't use all these fancy multi-gigabyte[4] hardware-limited rocBLAS/hipBLASlt/rocWMMA/rocRAND/etc. libraries.
[1] https://tinygrad.org/#tinybox
[2] https://github.com/pytorch/pytorch/issues/119081
[3] https://github.com/tinygrad/tinygrad/blob/v0.9.0/tinygrad/ru...
There are multiple bounties just for it in https://docs.google.com/spreadsheets/d/1WKHbT-7KOgjEawq5h5Ic...
It would be nice to see less whining and blaming AMD (PyTorch and llm.c actually work on 7900 XTX, and blow tiny grad out of the water in terms of perf!), and more just getting stuff to work.
https://github.com/tinygrad/tinygrad
Another good alternative would be:
* https://hn.algolia.com/?q=https%3A%2F%2Fgithub.com%2Ftinygra...
But, suit yourself. I just think you'll get higher engagement if you put your best food forward by helping newcomers understand what this thing is. But hey, maybe I'm wrong.
I'm biased, because I already know what tinygrad is, but there's a link to the main github page for the repo at the top of the page. There's "get higher engagement" but there's also "readers here aren't drooling morons and know what a github is and can click on the link to the repo and find the README.md". But hey, maybe I'm wrong.
That's pretty wild
https://code.dlang.org/packages/tiny-autodiff
Unlike other autograd libraries it utilized native D Mir GLAS linear algebra library [1]:
[1] Numeric age for D: Mir GLAS is faster than OpenBLAS and Eigen
http://blog.mir.dlang.io/glas/benchmark/openblas/2016/09/23/...
cache_key = (device, st, dtype, op, arg, tuple(ref(x) for x in srcs)) if base is None else (st, ref(base))The line count probably does still act as a limit on complexity overall but perhaps less than hoped for.
The README no longer mentions the limit and it looks like they just raise it whenever needed. Three months ago it was bumped to 6500 LOC. One month ago it was bumped to 8000 lines.
PyTorch does more than tinygrad, but does it really do 343x more things?
The Python source distribution has long maintained the philosophy of “batteries included” – having a rich and versatile standard library which is immediately available, without making the user download separate packages.
https://peps.python.org/pep-0206/
OTOH:
Simple is better than complex.
Complex is better than complicated.
https://peps.python.org/pep-0020/LOC limits have to be one of the worst incentives you can give programmers.
However, that doesn't solve everything. Many things it does not accurately measure, e.g. complexity, number of stuff in one line, program speed, memory usage, etc. Those are other things to measure, and it can be helpful to reduce memory usage etc, but that is not the number of lines of codes.