Specifically, I fine-tuned gpt-3.5 and llama-2-7b on some real-world usage data, and I can't tell a single difference between the quality of these outputs compared to "base" gpt-3.5 with the same requests made to it. Moreover, if I attempt to remove a lot of the "static" sections of a prompt (after all, there's 2k+ lines of data it was fine-tuned on where this is all present), both models just go completely off the rails.
I'd love to get into a world where I can fine-tune a bunch of different models. But it's so, so much harder than just calling OpenAI's API and getting really good results with that and some prompting work. If you're able to help crack that nut then there's a lot of people like me who would pay money to have their problems solved.