Local models help remove token cost uncertainty, but they shift the problem to infrastructure and ops. GPUs, scaling, maintenance, and latency can add up quickly depending on the workload. For many builders it ends up being a tradeoff between predictable infra cost and flexible API usage.
That isn't true, if you run local models you'll also need to have to spend on operations.
Maybe focus first on providing value and later you can optimize this setup.