I wrote a similar post some time ago just used ollama and opencode https://blog.kulman.sk/running-local-llm-coding-server/
As for oprncode, doesn't the system prompt eat too much of the context? Local models are really constraint in regards contex, and opencode AFAIR uses a 10k of it or some thing close.
I wish people would stop wasting their outrage budget on things like this and pay more attention to politics.
Technical people are rather good at learning new things, and ollama situation is a good learning experience.
llama.cpp gets you more tokens/s even if you ignore ollama team bad behavior.
Politicians count on your apathy so they can get away with their horseshit. You paying attention is kryptonite to crooked politicians.
FWIW, I took the dive on Pi today and I’m really happy with my decision so far