These days I use FastChat: https://github.com/lm-sys/FastChat
It’s not based on llama.cpp but huggingface transformers but can also run on CPU.
It works well, can be distributed and very conveniently provide the same REST API than OpenAI GPT.
It’s not based on llama.cpp but huggingface transformers but can also run on CPU.
It works well, can be distributed and very conveniently provide the same REST API than OpenAI GPT.
In terms of speed, yes running fp16 will indeed be faster with vanilla gpu setup. However most people are running 4bit quantized versions, and the GPU quantization landscape as been a mess (GPTQ-for-llama project). llama.cpp has taken a totally different approach, and it looks like they are currently able to match native GPU perf via cuBLAS with much less effort and brittleness.