1. Obtain the Llama models. Apparently you have to sign up for access and get a download link? I don't want to do that, found some public download links instead. Ok, now I have a few hundred GBs of model files.
2. Compile llama.cpp. Missing dependencies, took an hour to figure out how to resolve.
3. Quantize the model? What does that even mean?
4. Install pytorch. Run the command in the README. Python exception.
5. Install NVIDIA helper libraries. Doesn't work. Try installing the AMD helpers to run on CPU instead. No instructions for how to do this. Eventually figured it out.
6. Try running pytorch again. Same exception.
I gave up after a full day of trying to make this work. The Docker image is for people like me.