> Given how simple it has become to get the correct response shape, do you feel there's still a point to it?
I am personally working on systems which don't do mrequire uch structured data output; my work is more around structure, style, etc. Fine-tuning is more effective and more consistent for me in those contexts.
One thing to remember is that the big models, like 4o and Claude, both are fine-tuned for chat interactions. This tuning can often make them dumber, more verbose, etc. because that's what the human feedback testing likes. You can look at some of the rating tests on chatbot arena as an example where two give the same answer, but people have voted more favorably for the one that delivers it with "more personality."
There are SLMs that are better tuned for specific use cases and outperform the big models as a result of this, and the examples OpenAI showed in the article make it clear the big models can benefit from this as well.
> From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
This is partially correct, fine-tuning can be a step of this process. The full pipeline looks more like:
1. Create a prompt which gets a large model to output (mostly) correct responses.
2. Build a dataset of those inputs/outputs, probably with human or LLM curation/judging in the loop since some will still be wrong.
3. Fine tune a small model on those inputs/outputs.
Now you have a smaller model which behaves more like the large model you were able to prompt engineer into instruction following.