https://github.com/ggerganov/llama.cpp/pull/3362
https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF/tree/ma...
Though I can't figure out that prompt and with LLama2's template it's... weird. Responds half in Korean and does unnecessary numbering of paragraphs.
Just one big sigh towards those supposed efforts on prompt template standardization. Every single model just has to do something unique that breaks all compatibility but has never resulted in any performance gain.
MODEL=./models/mistral-7b-v0.1.Q5_K_M.gguf N_THREAD=16 ./examples/chat-13B.sh
So great performance on a cheap CPU from 2 years ago which costs, what $130 or so?
I tried Llama.65B on the same hardware and it was way slower, but it worked fine. Took about 10 minutes to output some cooking recipe.
I think people way overestimate the need for expensive GPUs to run these models at home.
I haven't tried fine tuning, but I suspect instead of 30 hours on high end GPUs you can probably get away with fine tuning in what, about a week? two weeks? just on a comparable CPU. Has anybody actually run that experiment?
Basically any kid with an old rig can roll their own customized model given a bit of time. So much for alignment.
Mistral AI's github page has more information on their sliding window attention method to achieve this performance: https://github.com/mistralai/mistral-src
If Mistral 7b lives up to the claims, I expect these techniques will make their way into llama.cpp. But I would be surprised if the required updates were quick or easy.