Chronos: Learning the Language of Time Series
arxiv.org
arxiv.org
In general, the M-Competitions (https://forecasters.org/resources/time-series-data/), the olympics of timeseries forecasting, have proven frustrating for ML methods... linear models do shockingly well and the ML models that have won, generally seem to be variants of older tree-based methods (ie. LightGBM is a favorite).
Will be interesting to see whether the Transformer architecture ends up making real progress here.
Transformer/ML models by themselves have a tendency to overfit past patterns. They pick up more signal in the patterns, but they also pick up spurious patterns. They're low bias but high variance.
It would be more interesting to compare an ensemble of transformer models with an ensemble of linear models to see which is more accurate.
(that said, it's pretty impressive that an ensemble of simple linear models can beat a large scale transformer model -- this tells me the domain being forecast has a high degree of variance, which transformer models by themselves don't do well on.)
Isn't that just dropout?
It might look like an ensemble because you’re selecting different subsets but ensembles combine different independent models rather than just subset models.
In my mind an ensemble is like a committee. For it to be effective, each member should be independent (able to pick up different signals) and have a greater than random chance of being correct.
I find a lot of these ML and DL libraries to be harder to troubleshoot beyond blind hyperparameter tuning whereas with stats I can tweak model, modify likelihood, etc. There’s also a lot of high value problems that have few data points these libraries tend to want at least daily data.
Also a followup question. With timeGPT and chronos advertised as "foundational time series models", do you think they have any value?
I’m not sure what to even make of a term like “foundational time series”. Does that just mean it’s widely used and known? You have to earn a role like that you can’t just declare yourself one.
(It depends on what you mean by "outperform" since metrics for classification and regression aren't always comparable, but I think I'm following the meaning of your comment overall)
> We tokenize text because text isn't numbers.
Text is actually numbers. People tried inputting UTF8 directly into transformers, but it doesn't work that well. Karpathy explains why:
My intuition is the following: transformers work really well for text, so we could try turning a time series into a "story" (limited vocabulary) and see what happens.
Text can be represented by numbers but they aren't the same datatype. They don't support the same operations (addition, subtraction, multiplication, etc).
I inherited a large time series JSON dataset in 2024. I've been successful in using the Observable Framework[1] by writing a Rust (rust-script) data loader[2] to parse and plot simple line charts[3] to visually see the data. There are hundreds of graphs over years of data so I would like to identify what graphs I should be paying attention to. My initial thought is to calculate metrics on each graph such as:
- Variability: how "spread out" are the data points from one another?
- Trend: direction of data path, up or down?
- Slope: are the data points increasing or decreasing?
- Level: where are the data points on the vertical axis?
What libraries, AI, databases, etc... would you recommend that would allow me to calculate these values? I am no data scientist and don't need forecasting but overall, I just want a dashboard that shows the most "important" graphs.[1] https://observablehq.com/framework/
[2] https://observablehq.com/framework/loaders
[3] https://observablehq.com/@observablehq/plot-simple-line-char...
edit: the x-axis is Time while the y-axis can be values such as duration, frequency, intervals
I already added linear regression marks that draws linear regression lines with confidence bands[1] to my Observable plots but they do not give me a “value” so I need to manually look at the graphs and read the red line.
[0] https://rc2e.com/timeseriesanalysis [1] https://otexts.com/fpp2/
Third edition: https://otexts.com/fpp3/
Load you time serie in a dataframe, and:
> - Variability: how "spread out" are the data points from one another?
So basically df.std(), with rolling variants for short term / long term.
> - Trend: direction of data path, up or down? - Slope: are the data points increasing or decreasing?
Just do a simple rolling linear regression of your data point against time.
I can see chronos working a bit better, as it tries to convert trends, and pieces of time series into tokens, like gpt does for phrases.
Ie. Stock goes down terribly, then dead cat bounces. This is common.
Stock goes up, hits resistance due to existing sell orders, comes down
Stock is on stable upward trend, continues upward trend
If I can verbalize these usual actions, it's likely chronos can also pickup on them.
Once again quality of data trumps all for LLM's, so performance might vary. If you read the paper, they point out a few situations where the LLM is unable to learn a trend, ie. When the prompting time series isn't long enough.
[1] https://aws.amazon.com/blogs/machine-learning/amazon-sagemak...
> In this work, we have focused on univariate time series forecasting since it constitutes the most common of real-world time series use-cases. Nevertheless, practical forecasting tasks often involve additional information that must be taken into account. One example involves covariates, that can be either time-independent (e.g., color of the product) or time-varying (e.g., on which days the product is on sale). Another closely related problem is multivariate forecasting, where historic values of one time series (e.g., interest rates) can influence the forecast for another time series (e.g., housing prices). The number of covariates or multivariate dimensions can vary greatly across tasks, which makes it challenging to train a single model that can handle all possible combinations. A possible solution may involve training task-specific adaptors that inject the covariates into the pretrained forecasting model (Rahman et al., 2020). As another option, we can build stacking ensembles (Ting & Witten, 1997) of Chronos and other light-weight models that excel at handling covariates such as LightGBM (Ke et al., 2017).
Probably just my own bias because it seems everything I deal with is at least MArP and anomalies are important to my use case.
I can see where this is useful for others, even Amazon suggests ARIMA or ETS if you don't have hundreds of related streams.
Is this more targeted at people who want more smoothing?
Or am I just missing something?
When these become publicly known and used, your system doesn't work any more because the prices now include whatever signal you had for yourself before.
e.g. If I have a good signal at predicting horizon 1 day, then it is in my interest to have many people trading it at horizon > 1 day, as they will push the price in my direction.