Fine-tuning now available for GPT-4o
openai.com
openai.com
You just cache a prompt with a ton of examples that you would have otherwise fine-tuned on. You can trivially update that prompt whenever, no asynchronous fine-tuning job needed.
Wonder if OpenAI will stick with fine-tuning or go towards prompt caching (or both). Fine-tuning has uses, but prompt caching gets you 99% of the benefits for 1% the effort.
From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
I am personally working on systems which don't do mrequire uch structured data output; my work is more around structure, style, etc. Fine-tuning is more effective and more consistent for me in those contexts.
One thing to remember is that the big models, like 4o and Claude, both are fine-tuned for chat interactions. This tuning can often make them dumber, more verbose, etc. because that's what the human feedback testing likes. You can look at some of the rating tests on chatbot arena as an example where two give the same answer, but people have voted more favorably for the one that delivers it with "more personality."
There are SLMs that are better tuned for specific use cases and outperform the big models as a result of this, and the examples OpenAI showed in the article make it clear the big models can benefit from this as well.
> From my understanding, fine-tuning allows for reducing model size, meaning less latency and cost. That seems like it would be the biggest advantage, no?
This is partially correct, fine-tuning can be a step of this process. The full pipeline looks more like:
1. Create a prompt which gets a large model to output (mostly) correct responses.
2. Build a dataset of those inputs/outputs, probably with human or LLM curation/judging in the loop since some will still be wrong.
3. Fine tune a small model on those inputs/outputs.
Now you have a smaller model which behaves more like the large model you were able to prompt engineer into instruction following.
So except for cases where you are really trying to perfectly match a particular writing style and want to fine-tune on loads of some author's text, I think prompt caching mostly dominates fine-tuning for most real-world workloads.
I'd still fine-tune if I wanted my model to have a really particular "voice" though.
When you have a question/answer tuple in your training data, you are also teaching the model that every other way of answering the question is wrong.
So while the LLM would probably be capable of generating maybe 100 answers to the question that would be equally useful (just using different phrasing, different choice of words, etc.) come on, you are forcing it to update its parameters to suppress all of these, except the one specific form that you selected.
So you’re not really adding knowledge. Instead, you’re chiseling knowledge away.
For example, create an "idea", generate thousands of Q&A pairs about that idea from different angles, and generate conversations about that idea, then train the model on it. This is essentially the Phi process, but with a single concept instead of "everything."
My guess is that fine-tuning cannot add the concept to the model without suffering from catastrophic loss everywhere else. However if we added that same data to the pre-training dataset and retrain the model, the model would express the idea correctly.
The only vendor I've seen with "automatic" prompt caching so far is DeepSeek https://platform.deepseek.com/api-docs/news/news0802/
> Maximize shared prompt prefix, by putting dynamic portions (e.g. RAG results, history, etc) later in the prompt. This makes your request more KV cache-friendly (which most LLM providers use) and means fewer input tokens are processed on each request.
source: https://platform.openai.com/docs/guides/latency-optimization
Claude's prompt caching gives a very material performance improvement - they claim up to a 4x performance boost in https://www.anthropic.com/news/prompt-caching - so if OpenAI have similar techniques, even undocumented, they should become visible through sending the same prompt a bunch of times and measuring the latency.
Without caching it's too expensive.
OpenAI's examples here are artificial, and there's no mechanism to try out the results: https://platform.openai.com/docs/guides/fine-tuning/fine-tun...
Fine-tuning is expensive in terms of both time and money (more so time these days, a lot of the vendors have free trials now). Before I put that work in I want to get a much better idea of the kind of results I can expect.
Prompt engineering, RAG, etc. provide so much more uplift per unit time invested. Maybe fine tuning makes sense if you have a 9 figure budget and thousands of humans to throw at it. I don't feel this works in a typical startup ecosystem. I couldn't move the needle with the amount of data I had on hand.
How might performance compare between:
1.) Using a model like Gemini to load it all into context at once
2.) Using one of the various summarization systems/embedding/RAG etc.
3.) Fine tuning the whole code base into Gpt-4o
There’s not been much opportunity previously to easily fine tune a Gpt 4 class model. I haven’t seen anything written up on this being tried.
I really like the control you get with aider over the LLM context. You can /add or /drop source code, markdown notes. You can /clear the chat. /tokens shows you the context and the cost, you can see what each prompt will cost you.
I find aider best used in conjunction with a git diff view in VSCode, I run aider with --no-auto-commits and then manually review each time in VSCode.
I'm keen to learn any AI coding workflows if anyone has any links. I've benefitted greatly from tips such as using type hints and documentation for the LLM's benefit.
Multi-file code suggestions with an intimate understanding of the entire code base.
https://x.com/aantix/status/1819794837375263228
P.S. Or Plandex, https://plandex.ai/
I saw this this morning - it really made me think about what the future of 1-man non technical founder startups will look like. https://x.com/0xluffyb/status/1825854097481736479?s=42&t=7-X...
Also on the very next paragraph:
> We’ve also implemented layered safety mitigations for fine-tuned models to ensure they aren’t being misused.
Well done, Sam
Am I wrong to split projects by env? Am I expected to run fine-tunes separately per env (surely not)? Am I missing an option to share fine-tunes across projects?
What do you mean by an "env" here?
I have a staging deployment and a production deployment. Ideally anything that I roll out to production, I can try on staging first — including switching from gpt-4o to a fine-tuned gpt-4o. I don't want the production API key to be accessible by my staging app, so I have two separate projects within the OpenAI dashboard. One is called my-app-prod, and the other is my-app-staging.
To illustrate the problem further, I also have infrastructure to eval models before changing what production is running. The eval infrastructure also has its own OpenAI project, so that I can set separate billing limits on running evals. Any fine-tuned model needs to be benchmarked first, but again, I'm not sure how to make the same fine-tune available to both the eval project, and the production app project.
Until this happens, and until we have flexibility of fine-tuning techniques, I will push investors still towards PyReft, Lora, related, and the open source models where all of these techniques work.
OpenAI's documentation has good examples of how the data should be formatted (and when it's appropriate to fine-tune): https://platform.openai.com/docs/guides/fine-tuning/fine-tun...
Fine tunes take that and change what gets dropped or compressed. Rather than somewhat evenly forgetting things, it forgets very little of the knowledge they’re targeting (like coding) and in exchange drops much more of “everything else”.
I don’t believe it works this way, but in a sense, it’s like GPT4o is a 400B parameter model but only devotes 20B parameters to coding because the rest are taken up by Wikipedia and knowing French and what not. A 70B coding fine-tune might be able to devote more than 20B parameters to coding, in exchange for only speaking English, having very little encyclopedic knowledge, etc.
It’s kind of like CPUs vs ASICs. CPUs (GPT4o) will perform better on average at random tasks, while ASICs (finetunes) do dramatically better on the task they were built for, even if they have a lower transistor count.
Wikipedia says that as of 2023, each expert is ~10 billion parameters, so they’re still much smaller than something like Deepseek Coder 70B.
I’m not sure why they’re so small though. I don’t know whether there’s some kind of architectural issue or super linear scaling somewhere preventing them from growing, or if they’re just trying to keep inference costs down.
It's possible OpenAI did some coding fine-tuning themselves; Meta's Llama 3 paper [0], section 4.3.1 mentions what sort of work is needed. However, anything OpenAI did is based on their own tooling and set of assumptions - e.g. how is the existing code input into the LLM, what set of actions can the LLM take (e.g. look up documentation), what language is the output code being written in, etc. Cosine's LLM framework may do things differently and have different features, so you'd need to fine-tune the LLM to take maximum advantage of the framework.
It's like dropping the LLM down in front of Vim when it had only ever used or even heard of notepad (or even emacs); there needs to be some training to make it work well with the new tools it has.
But the SWE-bench leaderboard (linked to in the post) doesn't show Cosine Genie at all, instead showing Amazon at the top with 38.8% accuracy.
If true, seems wild that Genie, a 10 person startup with $2.5m in funding can actually achieve SOTA results over Google, Amazon, Microsoft, Anthropic, OpenAI etc. It's not like the big players are overlooking this problem of using LLMs to automate software engineering. Anyone have more color on this? I see some speculation online that the training data can easily get contaminated with benchmark questions, but not much careful evidence.
Note SWE-Bench has recently modified their submission requirements, now asking for the full working process of our AI model in addition to the final results -their condition to have us appear on the offical leaderboard. This change poses a significant challenge for us, as our proprietary methodology is evident in these internal processes. Publicly sharing this information would essentially open-source our approach, undermining the competitive advantage we’ve worked hard to develop. For now, we’ve decided to keep our model’s internal workings confidential. However we’ve made the model’s final outputs publicly available on GitHub for independent verification. These outputs clearly demonstrate our model’s 30% success rate on the SWE-Bench tasks.
- frameworks
- hardware
- foundation models (OS vs 3rd party)
- evaluation & training data
Did the performance gains justify the investment for fine-tuning?
An example of this is DeepSeek-Coder, which can essentially be considered a fine-tune of a fine-tune of a Mixtral model. It performs very similarly to Claude 3.5 Sonnet, which is pretty damn impressive, but it does it at less than 1/10th the cost.
What I don't understand though is why anyone would even remotely consider fine tuning a GPT-4o model that they will never fully own, when they could spend the same resources on fine tuning a Llama3.1 model that they will own. And even if you absolutely don't care about ownership (???), why not do a fine tune of an Anthropic model which is already significantly better than GPT-4o. At this point, with the laggard performance of OpenAI and their shameless attempts at regulatory hostility to competitors, I can't imagine ever giving them any of my money, let alone owning my derivative work.
I've not heard that anywhere else. My impression from https://huggingface.co/deepseek-ai/deepseek-coder-33b-base was that DeepSeek-Coder was trained from scratch.
I’m now getting conflicting information about the origin of the deepseek MOE framework, so I may be wrong about it starting with a Mixtral model.
Why not give us nice things for integrating with knowledge graphs and rules engines pretty please?
[0] https://www.youtube.com/watch?v=1-hk3JaGlSU
I tried it for advent of code 2023, and it was pretty helpless.
Typing-in-code is only 10% of the work that I do, but this is still a very meaningful improvement for me.
I've written a bunch more about my own experiences here: https://simonwillison.net/series/using-llms/ and here: https://simonwillison.net/tags/ai-assisted-programming/
I recently made this blog post showing how y=mx+b is very error prone in GPT4o, and pretty accurate (to a point) in Claude 3.5.
I haven't gone down the rabbit hole yet, but I was wondering if fine tuning could fix math errors in LLMs. My initial hunch and understanding is it will not. I'll have to give your links a read/watch.
This is an existing phrase with an explicit meaning that you are clashing with.