From their big paper [2] "A foundation model is any model that is trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks; current examples include BERT [Devlin et al.2019], GPT-3 [Brown et al. 2020], and CLIP [Radford et al. 2021]."
[1] https://crfm.stanford.edu/ [2] "On the Opportunities and Risks of Foundation Models", Bommasani/Hudson/et al
One of the most well-known examples of a foundational model is the GPT (Generative Pre-trained Transformer) series developed by OpenAI. GPT models, like GPT-3 or GPT-4, are trained on large datasets containing diverse text from the internet, which enables them to generate human-like text, answer questions, translate languages, and perform various other tasks.
Foundational models are significant in the AI field because they allow researchers and developers to create a wide range of applications and solutions without having to train a new model from scratch for each specific task. This approach saves time, resources, and computational power while still providing a high level of performance across different tasks.