Moirai: A time series foundation model for universal forecasting
blog.salesforceairesearch.com
blog.salesforceairesearch.com
- Time-LLM (https://arxiv.org/abs/2310.01728)
- Lag-Llama (https://arxiv.org/abs/2310.08278)
- UniTime (https://arxiv.org/abs/2310.09751)
- TEMPO (https://arxiv.org/abs/2310.04948)
- TimeGPT (https://arxiv.org/abs/2310.03589)
- TimesFM (https://arxiv.org/html/2310.10688v2)
- GPT4TS (https://arxiv.org/pdf/2308.08469.pdf)
Yet not a SINGLE transformer-based model I've managed to successfully run has beaten gradient boosted tree models on my use case (economic forecasting). To be honest I believe these foundational models are all vastly overfit. There's basically only 2 benchmarking sets that are ever used in time series (the Monash set and the M-competition set), so it'd be easy to overtune a model just to perform well on these.
I would love to see someone make a broader set of varied benchmarks and have an independent third party do these evaluations like with LLM leaderboards. Otherwise I assume all published benchmarks are 100% meaningless and gamed.
Not disagreeing with you, and I'm not a specialist, but it's funny that lot of papers seem to claim exactly the opposite.
"Prophet is a procedure for forecasting time series data based on an additive model where non-linear trends are fit with yearly, weekly, and daily seasonality, plus holiday effects. It works best with time series that have strong seasonal effects and several seasons of historical data. Prophet is robust to missing data and shifts in the trend, and typically handles outliers well."
As a sidenote/rant, it would be nice if all supervised TS benchmarks included "DLinear + RevIN" as the standard baseline, as in my experiments it tends to get about the same performance as all other new SOTA forecasting models. Most papers compare to the linear model without RevIN while they themselves use it, and only beat it because of that :) And in any case supervised training of transformers from scratch on datasets having less than 1M points is just stupid (so less raw data than a single image?). Less than 1B is still at least mildly stupid.
Here of course the angle is zero-shot so its somewhat excused from this, but it still would be interesting whether it can beat that supervised model combination.
Edit: oh you’re one of the authors — thank you, and congratulations!
https://en.wikipedia.org/wiki/Makridakis_Competitions
Makridakis and Hibon reached the sad conclusion that "statistically sophisticated and complex methods do not necessarily provide more accurate forecasts than simpler ones."
The Wikipedia article doesn't have that much detail on M5 or M6, but the M5 papers are in the International Journal of Forecasting[1] and M6 should be published later this year (there's already a preprint on arxiv [2]).
I recently spent some time looking into the history and results of the M competitions and had a chance to speak to Professor Makridakis about them, as well as the winners of each of the M6 competition tracks [3]. While the methods have become more sophisticated, some conclusions from M1 still seem to hold: in particular, that there is no overall "best" method, and that the winning method tends to be different for different types of data, time horizons, and evaluation metrics.
[1]: https://www.sciencedirect.com/science/article/pii/S016920702... [2]: https://arxiv.org/abs/2310.13357 [3]: https://mlcontests.com/state-of-competitive-machine-learning...
https://github.com/Nixtla/nixtla/tree/main/experiments/amazo...
This is the gold standard of forecasting tools.
Moirai stands for fates [https://en.wikipedia.org/wiki/Moirai] in Greek mythology
Whether or not the forecasts improves as a result of the additional covariates is still an open question which needs to be studied more -- we need to build better evaluations and benchmarks for this.