My suggestions if you want to further experiment with local models are to use llama.cpp instead of ollama [1], learn a little about the parameters that tune how much VRAM is used [2], look online for jinja template fixes for the model you're testing [3], and choose a model that was designed to do the task you want to achieve, with as high quantization as you can fit. The maximum model size you can run is VRAM + RAM, although you want as little of the model to be in system RAM as possible.
I'm running North Mini Code IQ3_XXS with some tuned parameters to fit my current tasks, and while it is not perfect for everything, it has not failed any tool calls I've asked it to make, or that it figured it should make on its own.
[1]: https://sleepingrobots.com/dreams/stop-using-ollama/
[2]: https://github.com/ggml-org/llama.cpp/blob/master/tools/serv...
[3]: https://gist.github.com/jscott3201/e4b155885cc68c038d6ac8909...