The solution to 'GPT-4 sometimes breaks' isn't to use something that never works...
Finetuned local models can work just as well or better than gpt-4 in many use cases
Have you never used the open source models? They are getting really good - better than 3.5 for sure not as good as 4 except when domain trained in my opinion