> if you want improved performance, you still need more data
Not true. See figure 2: https://arxiv.org/pdf/2203.15556.pdf#page=5
The loss decreases with greater model size at the same compute budget (i.e. stopping sooner regarding training data). Also some rehearsal/multi-epoch training improves the forgetting rate (thereby improving performance substantially), which hasn't been taken into account by Chinchilla et al. because they train <1 epoch.