arguably you can reduce even more latency by keeping the model on-device as well, but that would mean revealing the weights of the fine-tuned model.
If the user preferred reduced latency and had the RAM, is that an option?
If the user preferred reduced latency and had the RAM, is that an option?
I just used Cody with Ollama for local inference on a flight where the wifi was broken, and it never fails to blow my mind: https://x.com/sqs/status/1803269013310759236.
https://sourcegraph.com/blog/local-code-completion-with-olla...