No, it wasn’t, except under a very limited conception of “possible”.
No, it wasn’t, except under a very limited conception of “possible”.
It seems likely that humans gain that understanding by interacting with the world. Such data isn't available to train LLMs. Just including just basic sensory inputs like image and sound would easily increase training data by many orders of magnitude.
Assuming the ratio of equally-easily-accessible data to all data remains the same, and assuming that human data doubles every two years (that’s actually the more conservative number I’ve seen), there will be an order of magnitude more equally-easily-accessible data to train a future version on in around 6 years, 8 months from when GPT-4 was trained.
For instance, just extend the sequence length longer and longer. How low can you push down your perplexity? Bring in multi-modal data while you're at it. Sort the data chronologically to make the task harder, etc. etc.
The billion dollar idea is something akin to combining pre-training with the adversarial 'playing against yourself' that alphazero was able to use, ie. 'playing against yourself' in debates/intellectual conversation.
But lets say you reduced it to some sort of 'trying to prove a statement' that can be verified along with a discriminator model, then compare two iterations based on whether they are accurately proving the statement in english language.