There was this very interesting paper out of Stanford this last September about pretraining under the unlimited compute but limited data paradigm[0]. Pretty much exactly the same thing but with ~200M training tokens instead.
Still, just for reference, here's the paper I remembered: https://arxiv.org/pdf/2507.15857