Continue with LocalAI: An alternative to GitHub's Copilot that runs locally
old.reddit.com
old.reddit.com
Llama.cpp (to serve the model) + the Continue VS Code extension are enough.
The rough list of steps to do so are:
Part A: Install llama.cpp and get it to serve the model:
--------------------------------------------------------
1. Install the llama.cpp repo and run make.
2. Download the relevant model (e.g. wizardcoder-python-34b-v1.0.Q4_K_S.gguf).
3. Run the llama.cpp server (e.g., ./server -t 8 -m models/wizardcoder-python-34b-v1.0.Q4_K_S.gguf -c 16384 --mlock).
4. Run the OpenAI like API server [also included in llama.cpp] (e.g., python ./examples/server/api_like_OAI.py).
Part B: Install Continue and connect it to llama.cpp's OpenAI like API:
-----------------------------------------------------------------------
5. Install the Continue extension in VS Code.
6. In the Continue extension's sidebar, click through the tutorial and then type /config to access the configuration.
7. In the Continue configuration, add "from continuedev.src.continuedev.libs.llm.ggml import GGML" at the top of the file.
8. In the Continue configuration, replace lines 57 to 62 (or around) with:
models=Models(
default=GGML(
max_context_length=16384,
server_url="http://localhost:8081"
)
),
9. Restart VS Code, and enjoy!
You can access your local coding LLM through the Continue sidebar now.Is it possible to use GPU for this? With R9 7900x and 32GB RAM it takes 15-30sec to generate response. I have a 6900XT which might be more suited for this.
./server -t 8 -m models/wizardcoder-python-34b-v1.0.Q4_K_S.gguf -c 16384 --mlock -ngl 60
(You might need to play around with the number of layers.)[Edit: make sure to compile llama.cpp with GPU support first, e.g., "make clean && LLAMA_CUBLAS=1 make -j"]
Like I can't find simple straight foward solutions or content that isn't tied back to a company.
Thanks.
https://huggingface.co/models?sort=trending&search=wizardcod...
But I don't know if there's enough good public data for open source models to get there.
Continue also works with various backends and fine-tuned versions of Code Llama. E.g. for a local experience with GPU acceleration on macOS, continue can be used with Ollama (https://github.com/jmorganca/ollama):
ollama pull codellama
from continuedev.src.continuedev.libs.llm.ollama import Ollama
config = ContinueConfig(
models=Models(
default=Ollama(model="wizardcoder:34b-python")
)
)Almost there!
Edit: when I bought my Macbook in 2021, I was like "Ok, I'll just take the base model and add another 16 GB of RAM. That should future proof it for at least another half-decade." Famous last words.
(written from a 64GB M1 Pro Max)
If it's running slow, make sure metal is actually being used[0]. You can get as much as a 50-100% boost in tokens/s, if by chance it's not enabled.
I'm averaging 7 to 8 tokens/s on an M1 Max 10 core (24 GPU cores).
[0] if using llama-cpp-python (or text-generation-webui, ollama, etc) try:
`pip uninstall llama-cpp-python && CMAKE_ARGS="-DLLAMA_METAL=on" FORCE_CMAKE=1 pip install llama-cpp-python`
However, when I run the LLM, OSX becomes sluggish. I assume this is because the GPU's utilized to the point where hardware-based rendering slows down due to insufficient resources.
I wonder if there's a way to avoid that slowdown?
n_gpu_layers should also be set to anything other than 0 (default). I don't think the exact number matters for metal, but I use 128.
* llama.cpp - https://github.com/ggerganov/llama.cpp
* KoboldCpp - https://github.com/LostRuins/koboldcpp
* GPT4All - https://gpt4all.io/index.html
llama.ccp will run LLMs that have been ported to the gguf format. If you have enough RAM, you can even run the big 70 billion parameter models. If you have a CUDA GPU, you can even offload part of the model onto the GPU and have the CPU do the rest, so you can get some partial performance benefit.The issue is that the big models run too slowly on a CPU to feel interactive. Without a GPU, you'll get much more reasonable performance running a smaller 7 billion parameter model instead. The responses won't be as good as the larger models, but they may still be good enough to be worthwhile.
Also, development in this space is still coming extremely rapidly, especially for specialized models like ones tuned for coding.