> Usage data: ChatGPT is said to generate on the order of 10 billion tokens of data per day – even before they opened their more compelling GPT-4o model to free users.
> Common Crawl (filtered) 410 billion tokens 60% of GPT-3 training data but only 44% of it was used i.e. 0.44 epochs (from the paper published May 28, 2020)
From Aligning language models to follow instructions January 27, 2022 [0]
> ...these models can also generate outputs that are untruthful, toxic, or reflect harmful sentiments. This is in part because GPT-3 is trained to predict the next word on a large dataset of Internet text, rather than to safely perform the language task that the user wants.
> To make our models safer, more helpful, and more aligned, we use an existing technique called reinforcement learning from human feedback (RLHF). On prompts submitted by our customers to the API, our labelers provide demonstrations of the desired model behavior, and rank several outputs from our models. We then use this data to fine-tune GPT-3.
> The resulting InstructGPT models are much better at following instructions than GPT-3. They also make up facts less often, and show small decreases in toxic output generation. Our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT-3 model, despite having more than 100x fewer parameters. At the same time, we show that we don’t have to compromise on GPT-3’s capabilities, as measured by our model’s performance on academic NLP evaluations.
> One way of thinking about this process is that it “unlocks” capabilities that GPT-3 already had, but were difficult to elicit through prompt engineering alone
Note about the mention of a 1.3B InstructGPT: They trained InstructGPT in a few different sizes including 1.3B, 6B and 175B [1]
From Training language models to follow instructions with human feedback March 4, 2022 [1]
> We start with a pretrained language model, a distribution of prompts on which we want our model to produce aligned outputs, and a team of trained human labelers. We then apply the following three steps:
> Step 1: Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution. We then fine-tune a pretrained GPT-3 model on this data using supervised learning.
> Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons
between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output.
> Step 3: Optimize a policy against the reward model using PPO. We use the output of the
RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO algorithm
> Steps 2 and 3 can be iterated continuously; more comparison data is collected on the current best policy, which is used to train a new RM and then a new policy
> The cost of increasing model alignment is modest relative to pretraining. The cost
of collecting our data and the compute for training runs, including experimental runs
is a fraction of what was spent to train GPT-3: training our 175B SFT model requires
4.9 petaflops/s-days and training our 175B PPO-ptx model requires 60 petaflops/s-days,
compared to 3,640 petaflops/s-days for GPT-3 (Brown et al., 2020). At the same time,
our results show that RLHF is very effective at making language models more helpful to
users, more so than a 100x model size increase.
From Introducing ChatGPT November 30, 2022 [2]
> We trained this model using Reinforcement Learning from Human Feedback (RLHF), using the same methods as InstructGPT, but with slight differences in the data collection setup. We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides—the user and an AI assistant. We gave the trainers access to model-written suggestions to help them compose their responses. We mixed this new dialogue dataset with the InstructGPT dataset, which we transformed into a dialogue format.
> To create a reward model for reinforcement learning, we needed to collect comparison data, which consisted of two or more model responses ranked by quality. To collect this data, we took conversations that AI trainers had with the chatbot. We randomly selected a model-written message, sampled several alternative completions, and had AI trainers rank them. Using these reward models, we can fine-tune the model using Proximal Policy Optimization. We performed several iterations of this process.
> ChatGPT is fine-tuned from a model in the GPT-3.5 series, which finished training in early 2022.
In all likelihood, the usage data collected from ChatGPT is being used to continuously train and update a reward model which is being used to continue fine-tuning and improve performance. As for sources of new data, the article lays those out pretty clearly. I agree with the article that models like Phi3 show that higher quality data is more important than simple volume of data and that paying experts to produce, edit and/or grade data will get you much further ahead than only focusing on increasing token count especially when RLHF can 100x the effectiveness of that data. The main reason for finding more sources of training data would be to increase the breadth of knowledge and fill in larger gaps that can't be tackled by a small number of experts. More internet slop won't do that so the contribution of scraped web data can be expected to steadily decrease.
[0] https://openai.com/index/instruction-following/
[1] https://arxiv.org/pdf/2203.02155
[2] https://openai.com/index/chatgpt/