Chinchilla's Wild Implications
lesswrong.com
lesswrong.com
I suspect that the more data modalities we add the less data would required, but that's not the whole picture either. For example, text-to-image generators often makes weird mistakes that look "unphysical", or objects that look like they're flowing into eachother. The reason is because these models (including DALLE) uses a simple UNET, which basically only sees textures. What it lacks is a human inductive bias that 2d images are typically representations of a 3d world, a world which contains largely discrete objects and physics. It makes these mistakes because it doesn't know what objects are, and need to brute force this idea from a ton of observations. Even simple cognitive abilities like object persistence requires time perception, which these models lack.
I think the fact that these models can make up for this deficit with a ton of data is very telling. There is a lot of low hanging fruit in integrating more data modalities.
What do you mean? If we send a robot to explore its environment, and train it by having it constantly predict the next video frame, wouldn’t it eventually learn the physics and therefore gained the “time perception”?
* The "bigger models" line of research is probably exhausted for now: "insofar as we trust our equation, this entire line of research -- which includes GPT-3, LaMDA, Gopher, Jurassic, and MT-NLG -- could never have beaten Chinchilla, no matter how big the models got"
* We don't really know how much written data is available. Even Google - which has access to data repositories like scanned books that they cannot share - seems to have trouble getting consistently sized datasets for unknown reasons.
* There seems to be as much (more?) written English available in books than on the entire web. MassiveText "scrape" = 506B tokens, MassiveText "books" = 560B tokens
https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla...
However it’s not clear if training for more than one epoch on deduped and well balanced data would help. I personally think it should and the reason people don’t do it might be because it’s too expensive.
Certainly when you have something as large as the datasets being used to train large transformer models, SOME of the data is already repeated. Why would one more epoch make it worse?
They make an argument that there might exist an unfortunate dup/unique data ratio in a dataset, where a model decides to memorize a frequently repeated chunk of data which is big enough to justify accuracy degradation happening for the rest of the data, but not too big to make memorization difficult (section 5.1). The degradation they show is substantial - almost as if going from 800M to 400M model.
Given that language models seem to mostly know when they're making correct predictions [1] this method might be useful for stretching the available datasets into something larger without falling into the same pitfalls repeating data would give you. And if you squint it kind of looks like daydreaming? When I'm learning a new skill I often find myself playing back scenes and internally running some simulations of what I would have done differently.
All of a sudden, the shortage of training data will no longer be a problem, since the amount of video data available dwarfs all other forms of data
the page loads perfectly for me on chromium 77, but after the bloated javascript finally loads, the entire content gets replaced with
> Error: TypeError: Object.fromEntries is not a function
/:
Kiwi source branch is Chromium 77.0 + Kiwi backported fixes, will show 88.0.4324.152 to websites for compatibility reasons (64-bit)