It's true you don't need regularization if you've never seen the same data twice, but that's a similar regularization to early stopping. You'd expect the larger number of parameters would make the training error drop faster with fewer tokens as well due to improved optimizability. But rather than "larger models need less data", I'd say the take away is more "larger models need fewer steps to optimize training error". None of the models get "good" until they've seen a number of tokens similar to the number of parameters.
Unless you've run many epochs on small data sets and seen the same results, in which case that's pretty weird/cool.