RAG is a bit like having a pretty smart person take an open book test on a subject they are not an expert in. If your book has a good chapter layout and index, you probably do an ok job trying to find relevant information, quickly read it, and try to come up with an answer. But your not going to be able to test for a deep understanding of the material. This person is going to struggle if each chapter/concept builds on the previous concept, as you can't just look up something in Chapter 10 and be able to understand it without understanding Chapter 1-9.
Fine-tuning is a bit more like having someone go off and do a phd and specialize in a specific area. They get a much deeper understanding for the problem space and can conceptualize at a different level.
Fine tuning is just training. You can completely change the model if you want make learn anything you want.
But there are MANY challenges in doing so.
I regularly do fine tuning on a model with fine results and little damage to the base functionality.
It is possible, but it's too complex for the majority of users. It requires a lot of work per dataset you want trained on.
Models currently also have no way to update themselves with new info besides us putting data into their context window. They don’t learn after the initial training. It seems if they could, say, read documentation and internalize it, the need for RAG or even large context windows would decrease. Humans somehow are able to build understanding of extensive topics with what feels to be a much shorter context-window.
For instance I have a policy that I try hard not to say anything like "most people think that..." without providing links because I work at an archive of public opinion data and if it gets out that one of our people was spouting false information about our domain, even if we weren't advertising the affiliation, that would look bad.
This is a foundational problem that requires your data. The way you search Etsy is different than the way you search Amazon. The queries these systems see are different and so are the desired results.
Trying to solve the problem with pretrained models is not currently realistic.
Those are being worked on and RAG is the ducktape solution until they become available
Some kind of incremental fine tuning is probably necessary to keep a model like ChatGPT up to date but I can't picture it happening each time something happens in the news.
I think you’d get a close approximation of speaking with someone who was watching the game with you.
To reach their potential LLMs need to know how to use external sources.
Update: After some more thinking - if you required it to know information about itself - then this would lead to some paradox - I am sure.
When CL is properly implemented in an LLM agent format, most of these systems vanish.
Source: built a few products using RAG+LLM products.