LLMs Unleashed: The Power of Fine-Tuning
lucaspauker.com
lucaspauker.com
The article refers to the BERT and GPT papers as the source of the fine-tuning idea. However, we actually first demonstrated it for universal models in 2017 and published the ULMFiT (Howard and Ruder) paper in early 2018. Prior to that, Dai and Le demonstrated the technique for in-corpus datasets. So it would be more accurate to say the approach can be traced back to those two papers, rather than to BERT and GPT.
BERT and GPT showed the effectiveness of scaling up the amount of data and compute, and switching the model architecture to Transformers (amongst other things).
I think it's more accurate the say that GPT and then BERT massively popularized and simplified the idea/approach. Prior to ULMFiT/GPT/BERT, fine-tuning usually meant freezing most of a model and tuning a small layer on top of it. (ELMo also fits somewhere in here, being a kind of halfway step). ULMFiT was a relatively lesser known work but to my knowledge one of the first to do 1) LM pretraining and 2) fine-tuning all the layers, albeit with some complexity (gradual unfreezing/different learning rates for layers).
GPT simplified this massively by simply tuning all weights: no nuance about it. (Also added a classifier on top). BERT took this a step further, benefiting from the larger size of BERT-large and bidirectional attention, which works really well on the NLU datasets of the time along with a classifier head. (It's not until T5 that seq2seq for general tasks became prominent again.) The idea that you tune all the weights of what was then considered a massive model was considered somewhat excessive at the time, but BERT so massively dominated every other architecture and approach at the time that everyone switched over to it. (Adapters (alongside PAL), one of the first parameter-efficient tuning methods in the Transformer era of NLP, came out shortly after.)
I'm pretty sure the earliest deep learning work from Bengio and Hinton (maybe even the unpublished circulating TRs from 2007) were basically:
* Let's pretrain autoencoders, layer by layer, to a large depth.
* Let's learn task specific fine-tuning through full backprop.
* (Oh btw this is better than learning a task-specific head on the whole network (And this is briefly noted since it's obvious to us since we work with this stuff so much).)
It was just that the entire deep learning approach was so radical overall that whole the fine-tuning vs. frozen features thing was buried in a ton a other new insights. And I mean radical. NeurIPS (then NIPS) wouldn't give Hinton, Bengion, and LeCun a workshop, so they organized a fucking satellite conference co-located at the same hotel as NIPS during the same time, on the topic of deep learning. And managed to dazzle everyone and steal the whole show.
With that said, I agree with your assessment that these 2018 works very squarely and in a laser-focused way put the idea of fine-tuning. There is totally huge value and impact for correctly packaging and rigorously assessing a specific approach known to be good but that was sort of an after-thought in published work. by expert practitioners.
With that said, I'll self-cite here and acknowledge (take the blame for ?) turning the NLP community on to using large-pretrained models and NOT fine-tuning them at all in 2010. Since the existing practice of designing features using expert knowledge hadn't gotten the community very far, I noticed Collobert + Weston doing large-scale pretraining of features in their amazing early NLP work that everyone slept on, and was like: "Hmmm, maybe NLP people would actually try deep nets if they could just use the features on their existing models and NOT worry about learning how to fine tune a net."
Anyway, greetings from Berlin.
We have some 100k context models too that can ingest entire documents.
So right now, I would say fine-tuning is probably only useful for a very narrow set of use cases.
What if I want an AI assistant that is specifically trained on a large codebase? Or all my product’s docs (which might easily exceed 100k characters for a big project). Or one that knows the exact details of the entire tax code? Or one that knows every line of Dostoyevsky’s novels so I can have angsty existential conversations with it? Or that can fully remember conversations with me that stretch on for years?
It seems like you’d need fine tuning for these kinds of use cases? Or am I missing something?
I wouldn't look forward to a PR for a significant feature from a human developer that just skimmed through all of a project's files and only read a handful of them in depth, so I guess I'm also skeptical this would lead to good results from an LLM.
Additionally fine tuning isn't a great fit for mutable knowledge like a code base.
Getting into the nitty gritty of files and functions would be a step further, but it seems to me that just having this general knowledge of how a big project is structured would vastly improve any output related to that project compared to just loading in comparatively tiny pieces of it. Even if those tiny pieces are the most relevant to the prompt, it will be missing all the background context that it has for projects in its training set.
Some of us hope that just-in-time retrieval of data from an external source (usually a vector db) is going to work. Like, say, "retrieve chapter 3 of Dostoyevsky's CP and make an interesting comment about it". For the record, you can have plenty of interesting angsty existential conversations with both GPTs about D right now if you so desire. Just preface your dialogue with something slightly more sophisticated than "let us have an angsty existential conversation about Dostoyevsky".
This JIT retrieval might work for some cases, but my guess is it won't for large holistic pieces of work where you have to integrate and "understand" the entire edifice at once. I'm not completely sure if codebases merit such a distinction though, but you can imagine a domain like "law" might. We can push the context limit and this might bring some temporary relief, but I'm not convinced. The recent pushes to 100K seem to rely on "dropping" attention in an intelligent way. It sounds like cheating and while it might work OK for some cases, I think it'll drop the ball eventually.
The integrated understanding that results from, somehow, internalizing the relation between all the hundreds of thousands of seemingly unrelated datapoints is what makes us interesting. That's also what makes an LLM interesting. It's just that an LLM only has access to that level of development when it's training.
If I were to guess I think we somehow need to enable "always-training" and I'm not sure anyone has the faintest idea how. You can only go so far if you step out of school and never learn anything ever again, no matter how brilliant you are. GPT4 is quite the scholar, but we're stretching it.
To ask the naive question, why can't we just keep continually training it with new stuff? If it takes, say, a year to ingest approximately the whole internet, tacking on a single large codebase should probably take something like an additional few seconds if we consider what fraction that codebase is of the full training set. I know it's not that simple for a lot of reasons, but it doesn't seem obvious that fundamentally new methods are needed to do this.
To give a dramatic and superficial example, let's say you just found out you have been living inside the Truman Show. This impacts, if not everything, then a whole lot of what you know about your world. Instantly and completely. This cannot be "tacked on". This has to be integrated somehow.
Current AI doesn't work like that. Instead of building a tower of knowledge it sort of creates abstract cognitive landscapes filled with predetermined averaged out paths that, when followed, give good answers. Useful for a lot of tasks, but it doesn't strike me as the type of structure that can be manipulated with any kind of precision. Maybe that will change and/or I am wrong. Certainly the last one is very probable.
After all, if I just want to detect from text what color and brightness the user wants to adjust their lights to, it seems inefficient to use a model that's been trained on all of human knowledge, even if I'm sure it'll work just fine.
References found here: https://www.hopsworks.ai/dictionary/in-context-learning-icl
Can anyone offer an example of a free public-facing LLM which has been fine-tuned by adding much specific info about some narrow area? Say, one that knows all the PR about some car brand or fandom? Somebody must have tried that by now.
https://www.reddit.com/r/LocalLLaMA/search/?q=uncensored&res...
People have posted comparisons showing it works pretty well. Some foundational models have enough censorship built in that they still try to resist it. You have to prompt those cleverly to do what you want them to do.
This is a popular tool for finetuning Alpaca:
https://towardsdatascience.com/fine-tune-your-own-llama-2-mo...
{'previous': 'Marge Simpson: Ooo, careful, Homer.',
'character': 'Homer Simpson',
'line': "There's no time to be careful."}
and then used it to generate more Simpsons dialogue. With a data set so well matched to the goal, the algorithm barely matters. Remember that recent paper, "Copy is all you need."?[1] This is the ideal case for that approach.I'm thinking more in terms of loading in the detailed product descriptions and maybe manuals from a catalog, then letting users ask questions about how to do things with the products.
uhhh. I understand what was intended there but while fine tuning may reduce the rate of hallucinations and make hallucinations more plausible, it's not magic accurate and trust-worthy dust.
Unfortunately many people think this stuff is magic and care should be taken to not encourage people to confuse improvements with resolving the issue.
One way of characterizing the LLM accuracy problem is that it often looks very accurate and convincing even when it is emitting nonsense. If you cast the problem in those terms-- as a problem of looking more trustworthy than it actually is-- fine tuning actually exacerbates the problem.
Here's an example of a paper that shows good results from fine-tuning for instance: https://arxiv.org/abs/2305.02301
See https://twitter.com/animaanandkumar/status/16816906540601876...
I think over time it’ll become standard practice to train special purpose models on top of general purpose models. If I have a very specific task domain a model tuned to that domain will less often wander off and will be more likely to respond in the desired way.
Ok. How do I upload my data? Can't find any link.
Sign up for our newsletter. Nope.
And we went to Stanford? I don't care.
I'm not trying to be nasty, though it probably comes of like that. Assume I want to get started with your product and invite me in...
You should try a post on parameter efficient tuning next!
I know some systems also allow an extra fixed embedding parameter to the prompt, that can also be fine-tuned. But that is yet another thing that can be fine-tuned very cheaply.
The narrative goes, "look how awesome ChatGPT is, imagine how good it would be trained on just your company's documents".
Which 1000% misses the point. ChatGPT is because (a) it is trained on almost nothing short of the entire corpus of human language ever created. At > 1 trillion parameters, it can have ~1000 parameters for every human on the planet. Let that sink in. And then (b) because it has been subjected to an unknown but likely massive amount of human reinforcement feedback.
The idea that you can meaningfully impact the output of the model towards factual accuracy or logical correctness just by doing a small amount of fully automated training using a tiny corpus of company documents is seductive, but super far from robustly demonstrated as far as I'm aware. Yet this is the pitch being sold very often.
A more recent example is stable diffusion fine tuned on specific subjects.
Whether fine tuning can reduce hallucination is first of all a question which only pertains to decoders and which is highly dependent on how the model has been fine tuned.