Maybe. In the context of natural language at least, Transformers require less and less data to reach the same result as you increase the number of parameters. No regularization needed. See Figure 2 in the paper Scaling Laws for Neural Language Models (2001.08361).
It's quite odd. Who knows if that will hold for other domains, like protein folding. It may very well be the case though, since AFAIK DeepMind's folding model used attention to reach these landmark results.