And instructions on how to change the provider to use Ollama w/ whatever model you want:
Install and run Ollama - Put ollama in your $PATH. E.g. ln -s ./ollama /usr/local/bin/ollama.
- Download Code Llama 70b: ollama pull codellama:70b
- Update Cody's VS Code settings to use the unstable-ollama autocomplete provider.
- Confirm Cody uses Ollama by looking at the Cody output channel or the autocomplete trace view (in the command palette).
- Update the cody settings to use "codellama:70b" as the ollama model
One issue, though: I took a look at the Cody website and it looks like one can't have unlimited completions even when self-hosting a LLM.
I understand you guys have a business model and need to make money out of it. I'm just asking because I work as a teacher and I have students who can't pay an extra subscription and/or students who want to hack into stuff.
A pull/merge request is being worked on: https://github.com/continuedev/continue/pull/758
ugh, not so easy.
Was testing with Codellama-70b this morning and it’s clearly a step up from other OS models
If those rent-seeking bastards at NVidia hadn't killed NVL on the 4090, you could do it on two linked 4090s for only $4k, but we have to live under the thumb of monopolists until such time as AMD 1. catches up on hardware and 2. fixes their software support.
I don’t _believe_ that either of these lets you bypass that restriction (although I’d love to be proven wrong), so if you don’t want to sign up for a subscription you’ll need to use something like Continue.
With careful prompt engineering, you can get a lot out of free Bard except when its censored.
But even if you don't have a faster rig, you can still leverage it for slower tasks to generate docs or tests.
Twinny should really be more popular, didn't find a more powerful no-bullshit plugin for VSCode.