This is true for single time series, where we are predicting P(x_t+1 | x_0..t)
DL has advantages when you
a) have additional context at each time step, or
b) you have multiple related time series.
For example, consider Amazon who predicts demands for all of their products. At each time step, they know about inventory, marketing efforts, and could even model higher dimensional attributes like persuasiveness of the item's description with NLP.
It's also true they have items that are highly correlated. Skis, Snowboards, and Ski jackets all likely have similar sales patterns. Leveraging this correlation can increase accuracy, and is especially useful when you have items with limited history.
Including all of that context is hard with a statistical model, and whatever equation a human can come up with to combine them is probably worse than a learned, embedding-based DL model.
Statistical models are a great starting point & baseline for most problems, but as you add real world complexity beyond the general case time-series that's not as true.
I might not be aware of it, but I wish there were more benchmarks/research on higher complexity problems.