LocalPilot: Open-source GitHub Copilot on your MacBook
github.com
github.com
1) CodeLlama 13B or another FIM model https://huggingface.co/codellama/CodeLlama-13b-hf. You want "Fill in Middle" models because you're looking at context on both sides of your cursor.
2) HuggingFace llm-ls https://github.com/huggingface/llm-ls A large language mode Language Server (is this making sense yet)
3) HuggingFace inference framework. https://github.com/huggingface/text-generation-inference At least when I tested you couldn't use something like llama.cpp or exllama with the llm-ls, so you need to break out the heavy duty badboy HuggingFace inference server. Just config and run. Now config and run llm-ls.
4) Okay, I mean you need an editor. I just tried nvim, and this was a few weeks ago, so there may be better support. My expereicen was that is was full honest to god copilot. The CodeLlama models are known to be quite good for its size. The FIM part is great. Boilerplace works so much easier with the surrounding context. I'd like to see more models released that can work this way.
e.g. with mitmproxy and llama-cpp-python server
python -m llama_cpp.server --n_ctx 4096 --n_gpu_layers 1 --model ./path/to/..gguf
and then with mitmproxy in another terminal mitmproxy -p 5001 --mode reverse:http://127.0.0.1:8000
and then set this in your vscode settings.json (the same as for localpilot): "github.copilot.advanced": {
"debug.testOverrideProxyUrl": "http://localhost:5001",
"debug.overrideProxyUrl": "http://localhost:5001"
}
works way better for me than localpilotI tried pretty much all of them with Continue in VSCode, and it's a bit hit and miss, but the main difference is the way the workflows work (Copilot is mostly line completion, Continue is mostly chat or patches). So the main value add here for me would be a more Copilot-like workflow (which seems to align better with the day-to-day experience I has so far).
Tab Completion seems fine for cases where I add one new line and the next two lines are just incredibly obvious. I am going to experiment with writing a little comment first to see if it primes the tab completion to do something non-obvious.
I'm often surprised that how quickly the model discovers reasonable patterns (even across files as well, that's often necessary to be correct).
With the diffs I find that by the time I describe things in a way it has a chance of working, I might have just written the whole thing myself. Especially as the diffs are often need correctly. With tabs that correction is part of a fast feedback look, with diffs it's so far a slower loop and just more awkward.
Of course, with changes to workflows one or the other can shine and if there's an interface that's faster than typing the instructions out that might just supercharge things for the diff/chat type too.
The autocomplete + chat only seem to work on their own model, all the other models currently one or the other.
Nonetheles it does look interesting, cheers for the suggestion!
But yes, confusing notheless.
Normally these projects hijack the word 'local' by sending all of your data to a 3rd party API and that is confusing... but we finally get one that runs the model locally, does what it says on the tin, and some people still find a way to paint it as deception?
So the confusion is because what it says it is explicitly different to what it actually is, which understandably will be confusing.
It also prevent other editors' users from building Copilot plugins. For example, there won't be a Copilot plugin for Emacs that can be accepted by Emacs's official repostiory.
> Visual Studio being closed source
If VScode isn't the de facto universal editor accepted by every programming language's community (notice that even this particular thread is about a VSCode plugin!), I won't be so worried.
I want Cody to work for me, but right now it doesn't. I really, really want whole project awareness (maybe even to leverage a concurrently running language server?) for my completions.
Do you guys have a usage tutorial or a video somewhere? Are you flexible with how your UI is being implemented (ie, can I pitch you ideas)?
We love feedback and ideas as well, and like I said are constantly iterating on the UI to improve it. I'm actually wrapping up a blog post on how to better leverage Cody w/ VS Studio, that'll be out either later today or sometime tomorrow. As far as feedback though: https://github.com/sourcegraph/cody/discussions/new?category... would be the place to share ideas :)
On Sourcegraph side, we do collect some telemetry to improve our products, and for enterprise use cases we can def work with you on what data Sourcegraph collects/stores and how it interfaces. For example, we recently added support for AWS Bedrock so you can run your own instance of the LLM and connect it. So we def have options we can explore with you.
- the pycharm plugin says I don’t have embeddings but the native app claims otherwise
- when indexing, it complains it cannot find the repo (i assume it is trying to fetch from remote, which is a private github, and not local disk) - i worked around this by removing the remote entirely from git but that is only a temporary solution
- i cannot choose the branch to index (i work on feature branches)
They have a great list of supported editors:
- Android Studio - Chrome (Colab, Jupyter, Databricks and Deepnote, JSFiddle, Codepen, Codeshare, and StackBlitz) - CLion - Databricks - Deepnote - Eclipse - Emacs - GoLand - Google Colab - IntelliJ - JetBrains - Jupyter Notebook - Neovim - PhpStorm - PyCharm - Sublime Text - Vim - Visual Studio - Visual Studio Code - WebStorm - Xcode
I have found that the completions are decent enough. I do find that sometimes the completion suggestions are too aggressive and try to complete more than I want so I end up leaving it off until I feel like I could use it.
It isn’t the code completion assistant like the one you mentioned above, and it probably never will be. I see it more as a perfect coding companion, that is always under your fingertips and relieves you of googling most of the times.
Yet it’s tied with OpenAI, and you have to pay it by yourself, but the former should be changed rather sooner than later.
Bonus: in develop branch there is some-kind-of release candidate that a way more robust that the current release is.
[1]: https://github.com/yaroslavyaroslav/OpenAI-sublime-text
Using a high quantised larger model gives you an unrealistic impression that smaller models and larger models are roughly equivalently capable… but it’s a trade off. The larger codellama model is categorically better, if you don’t lobotomise it.
It’d be better if instead of making opinionated choices (which aren’t great) it guided you on how to select an appropriate model…
Potato / potato
Also, the M1 Max has more bandwidth than an epyc Milan to actually feed all of that. It’s about the same bandwidth as a PS5, but in a mobile package, with none of the latency of GDDR6. Much more powerful than a standard dual-channel consumer cpu.
I’m not sure how possible that is to do, but I hope we can get there at some point.
But it might be useful if, say, you have a local GPU-powered machine on your LAN. I just wish they weren't using the advanced settings in the CoPilot extension and were using, say, one of the many OpenAI-powered alternatives (like Genie) -- would feel less like a hack and more like an alternative.