exllama is really memory efficient and really fast
[0] https://huggingface.co/docs/transformers/main/model_doc/llam...
[1] https://github.com/turboderp/exllama
[2] https://github.com/Lightning-AI/lit-llama
EDIT: Or do you mean cuda? Because yeah, it's such a shame AMD's Rocm is so bad even geohot gave up. it's examples don't even run without crashing.
https://github.com/RadeonOpenCompute/ROCm/issues/2198#issuec...
edit: Note that this is my project.
AMD gave him a binary blob driver and that fixed his problem. Also, tinygrad is the only Python framework I know that has full OpenCL acceleration.
So far it doesn’t look that AMD is fully on board with Tiny Corp, but they are talking…
a) CUDA won in a free market because NVidia showed they cared about it
b) Llama has support for OpenCL (via CLBlast) and Apple Metal
The OpenCL support already has a custom kernel for token generation.
No open source though.