Ollama is now available as an official Docker image
ollama.ai
ollama.ai
1. Obtain the Llama models. Apparently you have to sign up for access and get a download link? I don't want to do that, found some public download links instead. Ok, now I have a few hundred GBs of model files.
2. Compile llama.cpp. Missing dependencies, took an hour to figure out how to resolve.
3. Quantize the model? What does that even mean?
4. Install pytorch. Run the command in the README. Python exception.
5. Install NVIDIA helper libraries. Doesn't work. Try installing the AMD helpers to run on CPU instead. No instructions for how to do this. Eventually figured it out.
6. Try running pytorch again. Same exception.
I gave up after a full day of trying to make this work. The Docker image is for people like me.
helps the command line actually work
docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
docker exec -it ollama ollama run llama2
And that's all! brew install ollama
brew services start ollama
ollama run llama2
Works great on my M1 MacBook Air, although it's definitely not as good as ChatGPT. docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
You'll need to have the NVidia container toolkit installed.For docker-compose, do
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
Docs: https://docs.docker.com/compose/gpu-support/Why is everyone so afraid to get their hands dirty?
This is a fast moving space and can definitely be confusing.
But rather than pushing people to improve themselves or LEARN we cater to spoiled lazy child syndrome.
Put in the work, reap the rewards. No shortcuts!
The person you responded outlined their process (to a level you can look at and see they clearly made an effort) and stated they put a full day in. This is not "spoiled lazy child syndrome", although there are for sure those out there who fit that mold.
I asked GPT 4, Claude 1 & 2, and Bard questions about a single documentation page and they all gave both incorrect and hallucinated answers. For example, I asked them to list all functions that accepted integers based on provided documentation. Their responses included functions that didn't accept integers and functions never mentioned in the documentation.
> Get up and running with large language models locally.
> To run and chat with Llama 2:
ollama run llama2What are the system requirements for getting a decent model running in background?
Is it something I can run on a laptop with 8-12GB of RAM and not a huge harddrive?
StableLM 3B ~4 GiB
You could go even lower with smaller quantization if necessary. I personally wouldn't use anything smaller than 7B and Mistral already pushing it in coherence. Overall it depends on your use case, not everyone needs smart models, or large context that sometimes takes half of required memory.
Codellama is also surprisingly good even for non-coding tasks
version: '3'
services:
ollama:
image: ollama/ollama
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama-volume:/root/.ollamaI still don't see any official documentation showing how to get Ollama running on Windows?