https://docs.aws.amazon.com/sagemaker/latest/dg/jumpstart-fo...
It's possible to introduce new information by fine-tuning a new model on top of the existing model, but it's debatable how effective that is for introducing new information - most fine-tuning success stories I've seen focus on teaching a model how to perform new kinds of task as opposed to teaching it new "facts".
If you want a model to have access to updated information, the best way to do that is still via Retrieval Augmented Generation. That's the fancy name for the trick where you give the model the ability to run searches for information relevant to the user's questions and then invisibly paste that content into the prompt - effectively what Bing, Bard and ChatGPT do when they run searches against their attached search engines.
There are also newer techniques like or ROME that could edit individual facts, and you might also be able to get there when you are updating by doing a DPO tune of the old vs the new answers as well.
While I agree that RAG/tool use (with consistency checking) might be overall best approach for facts, being able to update/tune for model drift is probably going to still be important.
I'd also disagree about the training entirely from scratch - unless you're changing architecture/building a brand new foundational model or have unlimited time/compute budget, that seems like the worst option (and pretty unrealistic) for most people.
These open weights models can be retrained. Start with a foundational model like Llama2 or something and expose it to more recent training data that includes whatever updated information you want it to have access to. This is relatively expensive, but allows for big changes to the model.
If you have some relatively small subset of new information you want to bring in, you could build a Lora. Then either run your model with the Lora, or fold the Lora into your base model. This is relatively cheap, but fairly narrow in terms of your updates.
In the long run, it might be that Retrieval Augmented Generation (RAG) is the way to go. Here, your embeddings go into a vector database, and the model reads from there. Then you just need to update the database for the model to have access to new information.
This LLM stuff is new enough that anything like best practices are still being worked out. The optimal way to bring in new information could be a variant of one of the methods I mentioned above, or some combination of all three, or something else altogether.
- finetuning, as in restarting the models checkpoint and relearning it on the previous + new data
- adding extra neurons (e.g. LoRA adapters) at certain places and restarting learning
Oh, in classic machine learning there's also the "bagging/boosting classifiers" option, but I have no knowledge if that can be applied to a ANN.
The leaked Google "We have no moat" memo was very excited about LoRA style techniques, but it's not clear to me that it's been proven as a technique yet.
There are people (can point at a discord server) claiming it works for them and that they even sell finetuned models to business clients.
EDIT: I found one of the articles I tried to follow: https://www.mlexpert.io/prompt-engineering/chatbot-with-loca...
EDIT2: Ignore above. This seemed much more promising: https://www.youtube.com/watch?v=pnwVz64jNvw . Author provides consulting services and seemed very nice and approachable
For specifically knowledge you want it to be able to recall (like knowledge base articles or blog posts) vector database embeddings are best.
For knowledge you want it to operationalize, like being able to program in a new language the last resort is finetuning but this is not easy, requires massive amounts of high quality data, and is not generally effective for things which do not have a large amount of data to fine tune on (tens of thousands of pages worth of content).