One of the things that confused me is that regression models can be predictive, just like time series forecasting — they just do so in a different way. I tried to make this clear in the article (or maybe I’m not understanding what you’re saying).
In a regression model, you’re predicting target variables from feature variables. In a time series, you’re predicting the same variable from its past behavior. This is a subtle but crucial difference.
(And then you can do time series with covariates, which combines the two.)
My question is this. According to definitions, can the latter (f(X_t) = y_t) be a time series model if each row of data is a time step? It doesn't have any autoregressive terms in X, so I don't know if it categorically is a time-series model.
Not that this question even matters, it's purely a taxonomy/terminology question.
Although the term "regression" is a misnomer anyway, and often when people say "regression" they mean "linear model". And by "linear model", we mean specifically a model in which outputs/predictions are some fixed linear combination of the input.
It is however possible to interpret the Kalman filter as a kind of dynamic regression model. Check out here if you want a good math workout on that topic: https://stats.stackexchange.com/q/330696
(Another somewhat distinct meaning of the term "regression" is any model with a "continuous" outcome variable. This is usually in contrast to "classification", which is any model that has a "categorical" or discrete outcome variable.)
Suppose I have exogenous variables that vary over time, X(t). X is about 100 features. What are some methods I can apply onto X(t) to automatically engineer features that may be useful at predicting some noisy y(t)?
I want to simultaneously capture interactions/interdependence between the columns of X, as well as the autocorrelation structure of X.
If I treat X as merely tabular data, throwing it into a traditional regression model (e.g. XGBoost), it can capture the interdependence structure in X, but it will neglect the autocorrelation structure... Unless I manually engineer features that capture the autocorrelation structure in X (e.g. rolling/shifted/differenced features), but I want to explore methods that do that automatically.
I like this idea.
Practically, how would this look? Say X has 100 columns. Do we estimate 100 separate models f_{i}(X_{t}) = X_{i, t+1}, then generate 100 predictions for each time step, and then feed those 100 predictions into a regression to predict Y_{t}?
> cross correlation function of the X vs y
Is this supposed to be combined somehow with the f_{i} outputs?
Is this supposed to be combined somehow with the f_{i} outputs?
I'd rank the variables by their CCF, and use the top(n) to try to predict the series of interest.
Like, split Y in half, then use the X(1:(t/2)+n) to predict Y(t+n) to see if it works, and then if it works OK, actually model the top n X series and use them to really predict the Y.
It's a pretty manual approach, but you could automate it once you have a better idea what you're aiming for.
Usually our models are doing something like "Y = f(X) + E" where E is some unknown random noise and f() is the relationship that we are trying to infer from the data. We usually take X as "given" or "known", so in that case we are looking at Y conditional on some specific value of X.
If we are just trying to make good predictions, then we don't necessarily care about the structure among the components of X unless that structure tells us something about how Y is affected by X.
Imagine the following "true" relationships in the data, where E and H are unmeasurable random noise:
Y(t) = b0 + b1 * X(t) + b2 * X(t-1) + E(t)
X(t) = c * X(t-1) + H(t)
Knowing b0, b1, and b2 is sufficient to predict "Y minus random noise". Knowing c doesn't help us at all.If you're interested in obtaining good-quality estimates of b1 and b2, then you'll have a problem. That's because the direct effect of X(t-1) on Y is conflated with the indirect effect of X(t-1) on Y via X(t). But if you're just trying to make good predictions for Y, then you don't care as much about confidently distinguishing between b1 and b2.
That said, there are a lot of special considerations involved with timeseries data. There is a large number of specialized tools, techniques, and model families dedicated to time series modeling, which don't make sense to use for other kinds of problems. And all of those special time series tools exist to solve problems that do not arise in other modeling situations. So in practice, times series modeling is a distinct specialization from other kinds of modeling.
I should say that I enjoyed this post. And I think leaning into that confusion is my aim. In particular, my point of stochastic versus random is that they are more synonym than they are anything else. Just words that different groups came to use covering similar things.
Which is not to say that their aren't differences in the crowd that uses each term. I posit that most of the differences is in the aims of the crowd, and at the end of the day, you can get a lot of mileage by embracing the similarities. As opposed to the default of contrasting on the differences.
As a fun example, to me, if you view time not as just a number that always goes up, but as a number that cycles through the seasonal values, then it is easy to view as most any other feature. Similarly, the past is easy to envision as a feature of the present.
I do think the way you described a lot of time series analysis fits the fun read I had where Mandlebrot proposed a fractal view of time series predictions. Where you are looking for self similar behavior in the series data and reflecting/overlaying it on itself. But... as is probably guessable from the rest of my post, a lot of this is far outside of my comfort area. Love reading about it from a distance.
If you are using them to extrapolate (eg. Prediction) that should help you gauge how resilient you expect the model to be in prediction.
Obviously, for ARIMA the AR and MA parameters aren't very informative.
I use SARIMAX a decent amount, nonetheless.