The latter is always easier. Not to mention that the architectures are fundamentally curve fitters. There are many curves that can fit data, but not all curves are casually related to data. The history of physics itself is a history of becoming less wrong and many of the early attempts at problems (which you probably never learned about fwiw) were pretty hacky approximations.
> Hinton argues
Hinton is only partially correct. It entirely depends on the conditions of your optimization. If you're trying to generalize and understand causality, then yes, this is without a doubt true. But models don't train like this and most research is not pursuing these (still unknown) directions. So if we aren't conditioning our model on those aspects, then consider how many parameters they have (and aspects like superposition). Without a doubt the "superficial hacks" are a lot easier and will very likely lead to better predictions on the training data (and likely test data).