The unreasonable difficulty of time series forecasting
suzyahyah.github.io
suzyahyah.github.io
Whenever I teach people time series forecasting, I always point out that one of the biggest challenges is that you will always have values at prediction time that are out side the range of values observed during training (specifically the value of t).
In plenty of other machine learning and statistical modeling tasks this is not the case. You can train on every token you'll ever see and every pixel value you'll ever see, you can do regression analysis on every categorical value you include and an least an observation from within a range of every continuous and discrete value you'll observe. But with forecasting you will always have values you predict that are outside the range of anything you trained on.
You would run into similar problems if you tried to create a statistical model of the density of water given a temperature but your training data only included values between 0-100 C and you went out and started predicting values covering all the temperatures found on Earth.
For whatever reason, when time is a variable we somehow think it is immune from the obvious limitation of predicting on values outside of the range of values you trained on.
> You can train on every token you'll ever see and every pixel value you'll ever see, you can do regression analysis on every categorical value you include and an least an observation from within a range of every continuous and discrete value you'll observe.
Can you give an example of this? Lets say you are developing DLSS, you don't have the of a game that have not yet been developed.
> But with forecasting you will always have values you predict that are outside the range of anything you trained on.
A time series of my body temperature will only ever range from 20C to 50C. Outside of that range, I have bigger problems than my prediction being wrong.
1. Define {N = context duration, M = forecast duration} upfront
2. Select some time value T
3. Extract historical data whose timestamps lie in time interval (T, T+N+M)
4. Transform timestamp values to (-N, M) interval by subtracting T+N from each timestamp
5. Append timestamp-transformed data to training data
6. Goto 2
Or are you saying that people don't want to define N and M upfront?
I don't get this, time is usually not a covariate in ts models, so why is it a challenge?
Yes, it’s hard to predict markets. Because anybody who can successfully predict markets, does so, makes money, and changes the market so their predictions lose their edge.
Time series forecasts are a lot easier if you are forecasting, say, disk use in your servers or whatnot. (By “easy” I mean you can do a simple prediction and get useful insights.)
- there's no seasonal pattern to the matches, they happen sorta randomly.
- they drive increased query traffic in the hour or so before the game
- then during the game usage drops, sometimes to below "normal" depending on time of day and who's playing
So... now the accuracy of your forecasting tool depends on correctly predicting when world cup matches happen, and also who wins them!
edit: and this is just one recent example. others involve severe weather, national gameshows, earthquakes, and when you celebrate christmas.
But I’m sure the World Cup is still pretty relevant for operations.
As I mentioned in another comment, this can also be rephrased as "predicting data with values outside the range you trained on typically doesn't go well". If you tried to predict some health metric based on weight and height but you only had people under 4' 10" and less than 120lbs you wouldn't be shocked at all if it worked terribly when applied to American football players.
Time-series forecasting is hard because you are always going to be predicting based on data outside of your observed range ("forecasting" does go much better when you're trying to fill-in-the-blanks of things that happened in the past).
However, the aim of math in these situations is often to give explanations for intuitive impressions like "predicting the future is hard". The concept of NP-completeness gives one (very partial) explanation why certain computing problems are "hard", for example. So that theory doesn't "boil down to saying programming is hard". Unfortunately, I don't think the text really gives strong explanation in this case.
Consider for example the use case of forecasting the average speed on a road segment with a maximum speed of 70mph. Forecasting whether that will be 69.8 or 70.3 is not very relevant. What is relevant is forecasting when the speed drops below a traffic jam threshold. But the exact timing of that might be impossible to forecast due to the inherently chaotic behavior of traffic. Forecasting the probability of a traffic jam occurring may be more interesting to practitioners.
Interestingly market volatility often goes hand-in-hand with increased correlation between asset prices: https://en.wikipedia.org/wiki/Anna_Karenina_principle#Order_...
This corresponds to a restatement of Murphy's Law, namely "life is a bitch and then you die". When you most need a diversified portfolio, diversification is hardest to achieve.
Separately, I've wondered for some time if there might be some reliable way to predict non-stationary data. While I don't have the answer, it occurs to me that it will possibly be a non-statistical method due to the fundamental incompatibilities. However, it also occurs to me that, given enough information, every data-generating process actually could be predicted. For instance, in the stock example, if you could model every single input into the system of a single company's stock, including every variable affecting every human that might conduct a transaction of it (daunting and unrealistic as that might be, but this is a thought experiment), then I believe the problem of prediction stops being non-stationary and in fact becomes completely deterministic, if complex. In such a scenario, wouldn't you be able to accurately make your prediction? I believe that perhaps chaos theory could present us with some solutions here where pure statistics (or, rather, simple statistics) cannot.
Just my 2 cents..
There are chemical systems where Lyapunov time is small enough that you can only predict seconds or minutes into the future and astronomical systems that are nonlinear and chaotic but have a long enough Lyapunov time that you can make reasonable predictions for millions of years. For both of these scales, the Lyapunov time still bounds how far into the future you can expect your predictions to remain near to the actual behavior of the system.
I do not know the Lyapunov times of the financial markets. That said, mathematically-sophisticated professional analysis frequently get their predictions wrong in major ways, so I expect there is a pretty hard bound on predictive quality caused by a short Lyanpunov time of the markets themselves.
In the finance space, the stationary core is often some 'stylized fact' that you're hypothesizing will hold true. This could be e.g., the momentum factor, that if you strip away the noise, there's an underlying trend that will hold over an extended duration.
This is the approach I'm using at my job, which is incident detection with customer metrics. We're tagging our time series data with common features -- such as country, customer type, etc -- with the idea that we can do a graph-like search to find exogenous variables. We can also use this to identify time series that have a similar "data generating process" and are simply different "realizations" of each other.
We don't need great time series forecasts, just something that detects large deviations quickly. We can then add in an existing dataset of _known_ incidents, indexed by the same common features, as a training/validation set.
Beware of the Anna Karenina principle. Well behaved data might be explicable by the same common features, but often the anomalies all have unique characteristics.
this seems right to me. maybe another interesting approach would be a fusion llm+ts model that does multiple-input-single-output with input metadata and causality narrative. so it "thinks" about what data it has and how predictive it may be of the target variable and when something "interesting" occurs it uses the big priors to synthesize a good guess at what it would look like.
so you'd have something like the time series data plus textual narratives of the causality stories as the training data.
that being the case, it seems the next jump in performance would come from incorporating both metadata and metadata enriched causality and maybe that next jump in performance would be the most interesting jump from a practical system that is useful perspective. (it's more valuable for a system to predict outlier events than it is for it to do an excellent job at synthesizing ordinary behavior)
The takeaway might be that historic data of market data might often be enough to make a reasonable prediction. Only external data that nobody else has (used) can make your prediction better then the market.
Making a prediction with same accuracy as the market: easy Making a prediction with more accuracy then the market: very hard
that's basically the definition of an efficient market. demand forecasting and insuring against price shifts is the actually useful thing that the commodities futures markets do.
also a fundamental difference from other time series prediction problems. there are all sorts of weird dynamics that go on before one biosignal effects another, or one metric predicts a failure, where a reasonably efficient market price reacts quickly to well known exogenous factors.
Financial markets are not a natural phenomenon that exist unchanged regardless of whoever is observing them. Their dynamics continuously change in response to collective actions of all of the humanity.
Oil price changed quite a bit when the US attacked Iran. If you are trying to predict the price of oil, your model would have to be able to predict Trump ordering an attack on Iran. Does your model include a full simulation of the mind of the president of the United States (and every other person who have any kind of impact on the world events)? If not, then your time series forecasts are not going to be that great.
Also, one method works better than another is more meaningful than "predicting the future is hard".