If you're looking to do the same with open source code, you could likely run Ollama and a UI.
https://github.com/jmorganca/ollama + https://github.com/ollama-webui/ollama-webui
If you're looking to do the same with open source code, you could likely run Ollama and a UI.
https://github.com/jmorganca/ollama + https://github.com/ollama-webui/ollama-webui
It looks like a lot of the tooling is heavily engineered for a set of modern popular LLM-esque models. And looks like llama.cpp also supports LoRA models, so I'd assume there is a way to engineer a pipeline from LoRA to llama.cpp deployments, which probably covers quite a broad set of possibilities.
Beyond llama.cpp, can someone point me to what the broader community uses for general PyTorch model deployments?
I haven't quite ever self-hosted models, and am really keen to do one. Ideally, I am looking for something that stays close to the PyTorch core, and therefore allows me the flexibility to take any nn.Module to production.
[1]: https://github.com/jmorganca/ollama/blob/main/docs/import.md.
As far as I know, ollama doesn’t support exllama, qlora fine tuning, multi-GPU, etc. Text-generation-webui might seem like a science project, but it’s leagues ahead (Like 2-4x faster inference with the right plugins) of everything else. Also has a nice openai mock API that works great.
https://github.com/oobabooga/text-generation-webui/issues/41...
ooba's worth keeping an eye on, but koboldcpp is more stable, almost as versatile, and way less frustrating. It also still supports GGML.