I am looking for an open source models to do text summarization. Open AI is too expensive for my use case because I need to pass lots of tokens.
I am looking for an open source models to do text summarization. Open AI is too expensive for my use case because I need to pass lots of tokens.
It’s not based on llama.cpp but huggingface transformers but can also run on CPU.
It works well, can be distributed and very conveniently provide the same REST API than OpenAI GPT.
In terms of speed, yes running fp16 will indeed be faster with vanilla gpu setup. However most people are running 4bit quantized versions, and the GPU quantization landscape as been a mess (GPTQ-for-llama project). llama.cpp has taken a totally different approach, and it looks like they are currently able to match native GPU perf via cuBLAS with much less effort and brittleness.
You can run open-source models, but the software itself is closed-source and free for non-commercial use.