To elaborate on your point, it doesn't just require lots of data, it requires lots of data with a corresponding high effective sample size. If you want to forecast sales for next Christmas, I don't care if you have 2000 terabytes of granular orders and sales data, because the effective sample size for past observed Christmases is going to be like 3-4 (Christmases further than 4 years back may no longer be representative).
In these cases nice structural time-series models, which are in spirit not so different from what existed 20 years ago, will beat deep learning.