Bitter lesson wildly overstated in this context.
Bitter lesson wildly overstated in this context.
(had to look it up)
That may not be the intent of the original article, but over the past few years that’s what the phrase turned into.
Couch cushions?
A lot of pretraining is also choosing the right type of data, you don't want to just have it ingest garbage (although I read that some amount of garbage actually helps the model be more robust). Pretraining crystallizes a lot of the inductive biases that post-training builds on, so by crafting the right data mixture you can make it easier for it to start off with a good foundation. There is also a lot of focus on mid-training these days, which I understand is basically either the name for the synthetic data stage, or the SFT phase before all the RL
As GP said. More RLHF is in fact the bitter lesson.
My sense of the Sutton Dwarkesh interview was that he was calling out that he didn't mean just longer datasets, but rather learning through exploration and that's exactly RL.