ollama run mixtral
That's it. You're running a local LLM. I have no clue how to run llama.cpp
I got Stable Diffusion running and I wish there was something like ollama for it. It was painful.
ollama run mixtral
That's it. You're running a local LLM. I have no clue how to run llama.cpp
I got Stable Diffusion running and I wish there was something like ollama for it. It was painful.
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
make
wget https://huggingface.co/TheBloke/Mixtral-8x7B-v0.1-GGUF/resolve/main/mixtral-8x7b-v0.1.Q4_K_M.gguf?download=true
./main -m ./mixtral-8x7b-v0.1.Q4_K_M.gguf -n 128It's probably a simple build if everything is how it wants it, but it wasn't in my machine, while running ollama was.
I only need to know the model name and then run a single command
I use Docker Compose locally, Kubernetes in the cloud
I run in hot-reload locally, I build for production
I often nuke my database locally, but I run it HA in production
It is very rare to use the same technology locally (or the same way) as in production
Look, I’m not arguing that a prebuilt binary that handles model downloading has no value over a source build and manually pulling down gguf files. I just want to dispel some of the mystery.
Local LLM execution doesn’t require some mysterious voodoo that can only be done by installing and running a server runtime. It’s just something you can do by running code that loads a model file into memory and feeds tokens to it.
More programmers should be looking at llama.cpp language bindings than at Ollama’s implementation of the openAI api.
Ollama makes that super easy. I tried llama.cpp first and hit build issues. Ollama worked out of the box
Just be aware that there’s a lot of expressive difference between building on top of an HTTP API vs on top of a direct interface to the token sampler and model state.
Python seems to be the way to go deeper though. Is there a good reason I should be aware of to pick llama.cpp over python?
Most of AI has centralized around Python, I see more of my code moving that way, like how I'm using LlamaIndex as my primary interface now, which supports ollama and many more model loaders / APIs
Finding the correct model weights is also a challenge in my experience, there are a lot of alternatives and it is often difficult to figure out what the differences are and whether they matter.
The README is clear that I'm probably about to lose an hour debugging if I follow it. It might be one of those rare cases where it works first time but that is the exception not the rule.
On multiple occasions I've been modifying llama.cpp code directly and recompiling for my own purposes. If you're using ollama on the command line, I'd say having the option to easily do that is much more useful than saving a couple commands upon installation.
I stopped using C++ when Go came out, no interest in ever having to write it again.
For non technical people there is a possibility their os don't have git, wget and c++ compiler (especially in windows)
This is just like dropbox case years ago.
Kudos if Ollama has this sorted out.