It should be fairly obvious that ‘curve fitting’ is a misleading category—these models are clearly learning highly meaningful latent spaces that no prior approaches ever did. But I would agree that the actual high-level ability to make causal inferences seems to be lacking.
Where I disagree with Pearl is simply with the idea that these stronger models won't emerge through future research. It's too early to say this, after barely a decade of large-scale AI research that has been undergoing continual rapid progress. Greater generality and more powerful models are some of the most well-established goals of the field.