So except for cases where you are really trying to perfectly match a particular writing style and want to fine-tune on loads of some author's text, I think prompt caching mostly dominates fine-tuning for most real-world workloads.
I'd still fine-tune if I wanted my model to have a really particular "voice" though.
From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
I am personally working on systems which don't do mrequire uch structured data output; my work is more around structure, style, etc. Fine-tuning is more effective and more consistent for me in those contexts.
One thing to remember is that the big models, like 4o and Claude, both are fine-tuned for chat interactions. This tuning can often make them dumber, more verbose, etc. because that's what the human feedback testing likes. You can look at some of the rating tests on chatbot arena as an example where two give the same answer, but people have voted more favorably for the one that delivers it with "more personality."
There are SLMs that are better tuned for specific use cases and outperform the big models as a result of this, and the examples OpenAI showed in the article make it clear the big models can benefit from this as well.
> From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
This is partially correct, fine-tuning can be a step of this process. The full pipeline looks more like:
1. Create a prompt which gets a large model to output (mostly) correct responses.
2. Build a dataset of those inputs/outputs, probably with human or LLM curation/judging in the loop since some will still be wrong.
3. Fine tune a small model on those inputs/outputs.
Now you have a smaller model which behaves more like the large model you were able to prompt engineer into instruction following.
When you have a question/answer tuple in your training data, you are also teaching the model that every other way of answering the question is wrong.
So while the LLM would probably be capable of generating maybe 100 answers to the question that would be equally useful (just using different phrasing, different choice of words, etc.) come on, you are forcing it to update its parameters to suppress all of these, except the one specific form that you selected.
So you’re not really adding knowledge. Instead, you’re chiseling knowledge away.
For example, create an "idea", generate thousands of Q&A pairs about that idea from different angles, and generate conversations about that idea, then train the model on it. This is essentially the Phi process, but with a single concept instead of "everything."
My guess is that fine-tuning cannot add the concept to the model without suffering from catastrophic loss everywhere else. However if we added that same data to the pre-training dataset and retrain the model, the model would express the idea correctly.