George Hotz already implemented LLaMA 7B and 15B on Twitch yesterday on GPU in Tunygrad llama branch:
https://github.com/geohot/tinygrad/tree/llama
The only problem is that it's swapping on 16GB Macbook, so you need at least 24GB in practice.
https://github.com/geohot/tinygrad/tree/llama
The only problem is that it's swapping on 16GB Macbook, so you need at least 24GB in practice.
George Hotz | Programming | can we fit a LLaMA inside a tinygrad? https://www.youtube.com/watch?v=0kRDs9BW2NU
George Hotz | Programming | ChatLLaMA: get in losers we're building a chatbot https://www.youtube.com/watch?v=nctqc8FBJ2U
That's double of what Tinygrad uses
although, there is a VOD channel on YT that might be better.