We love ARIMAs. That is why we put so much effort into creating fast and scalable Arimas and AutoArima in Python [1].
Regarding your valid concern. There are several reasons for the high computational costs. First, ARIMA and other "statistical" methods are local, so they must train one different model for each time series. (ML and DL models are global, so you have 'one' model for all the series.) Second, the ARIMA model usually performs poorly for a diverse set of time series, like the one considered in our experiments. The AutoARIMA is a better option, but its training time is considerably longer, given the number and length of the series. Also, AutoARIMA tends to be very slow for long series.
In short: for the 500k series we used for benchmarcking, ARIMA would have taken literally weeks and would have been very expensive.
That is why we included many well-performing local "statistical" models, such as the Theta and CES. We used the implementations on our open-source ecosystem for all the baselines, including StatsForecast, MLForecast, and Neuralforecast. We will release a reproducible set of experiments on smaller subsets soon!
[1] https://nixtla.github.io/statsforecast/docs/models/arima.htm...