A decoder-only foundation model for time-series forecasting
blog.research.google
blog.research.google
You can find recent papers from researchers about how their new transformers model is the best and SOTA, papers which claim transformers is garbage for time series and claim their own MLP variant is SOTA, other papers which claim deep learning in general underperforms compared to xgboost/lightgbm, etc.
Realistically I think time series is incredibly diverse, and results are going to be highly dependent on which dataset was cherry-picked for benchmarking. IMO this is why the idea of a time series foundation model is fundamentally flawed - transfer learning is the reason why foundation models work in language models, but most time series are overwhelmingly noise and don't provide enough context to figure out what information is actually transferrable between different time series.
That's exactly right it is bearly one step above playing "guess which number I'm thinking of" and acting amazed that if you play long enough you'll witness an occasional winning streak.
My god, model has learned to read your mind! ;)
This smacks of when very serious soviet scientists ran Telekinesis experiments and all manner of cold reading and charlatans. https://en.wikipedia.org/wiki/Telekinesis
Somebody should come up with a decoder-only foundation model for bending-spoons.
Interesting reading: https://www.kaggle.com/competitions/m5-forecasting-accuracy/...
You can use that description for system identification, i.e. build a model of how the "world" works. This can be translated into a two-part network architecture, one is essentially a world-model-informed (physics-informed, as its often called in literature) part for the known-unknowns, the other one is a bounded error term for the unknown unknowns (e.g. dense layers, or maybe dense layers + non-linearities that capture the fundamental modes of the problem space for reservoir computing). The world model is revised with another external cycle of meta-learning, via symbolic regression.
The unknown-unkowns bit that I choose is designed as a shallow network that can be trained online (traditional methods, but I'd like to see if the forward-forward algorithm from Hinton would work well for short-term online adjustments), or by well-known tools like particle filters / kalman filters.
The non-linearities (and the overall approach in general) resemble physics-informed dynamic mode decomposition (piDMD), which show remarkable resistance to noise (e.g. salt & pepper).
If you have simple timeseries, and not very complex hierarchical systems that change over time and show novel modes that you haven't encountered before, then piDMD is likely enough for what you need.
---
Essentially, what I describe is a multi modal model for timeseries + a planning step. (AlphaGeometry to the rescue?)
---
Like you say, time-series in general have an incredibly complicated domain with comparatively few data available.
For example, real-world complex physical systems (industrial plants, but also large-scale software systems) may have replicas of the same components with complex behavior, and no/few shared dependencies for reliability or other constraints (e.g. physically apart). These can be captured by transformers. Training will be much faster if you initialize weights like I describe above, and share weights among replicas. The physical structure also creates particular conditions on the covariant matrixes and on the domain of higher-level timeseries (ultrametric spaces, which changes how measurement and frequency behaves there and can lead to great simplifications, but also errors if tools like FFT are applied blindly without proper adjustments; much like in operations research / planning problems, symmetries are sought after to reduce complexity).
On the other hand, the next level that compose these building blocks often have graph structure and sometimes scale-free networks (e.g. if they represent usage or behaviors, rather than physical systems). I think we'll see graph neural networks shine on this front.
There are likely other kind of behavior that I haven't encountered yet in my work.
I think overall, we'll see planning/neurosymbolic used at the highest-layer, graph neural networks for scale-free networks and to optimize long-range connections (also when a dense model with dynamic covariant matrices would be too expensive to compute even in sparse form), and transformers or/with piDMD-like approaches for dense patches of complex behavior. I.e. graph models as a generalization for spatial locality to arbitrary spatial-like domains, transformers/piDMD or similar for sequence-/time- locality for arbitrary complex systems. (I wonder what kind of weights will they implement when trained together on problems that are fundamentally in the middle, where traditionally one would use wavelets... if you look a the GraphCast model by deepmind for weather forecasting, it looks quite similar)
Could you elaborate on this please? On why the transformer architecture lends itself well to this?
You can try pre-train a transformer to capture the behavior of a common part that is replicated, and then make replicas (sharing weights, or not depending on the problem) to train the whole ensemble. It works both for existing systems but also from high-fidelity enough simulations, or proxy systems that show the same range of behaviors (e.g. staging environments).
Even if the pre-train network part doesn't converge fully or capture everything, it can pre-condition the network and help training the whole ensemble faster.
---
For simple components, you can even just write your own simulations as custom NN layer (e.g. recurrent layers that take the current state and input, return the output and next state). It helps to avoid the performance bottleneck of going outside accelerators for simulations, or having to train too many small networks.
I'd generally just write my own recurrent layer, if the behavior is simple enough.
But you can also use existing code and tweak it cleverly: e.g. LSTM Cells can be pre-initialized to implement continuous-time markov chains, as a birth/death renewal process.
You can capture the behavior of a simple component in isolation, then use it in the whole. Either freezing it and adding an error correction layer (e.g. if the frozen part is quite big and replicated, you can share weights more efficiently), or not freezing it and letting it train further.
You can impose bounds on the complexity of the error correction, very much like LoRA you can design it as a low-rank matrix decomposition; together with the right loss (e.g. L1 or huber), it's another technique to ensure that the error correction doesn't drift too much away from the behavior that you can expect from the physics of the system (and when it no longer converges, it is a good indicator that you have model drift and new behavior is coming up... that's a way to implement robust anomaly detection).
---
PS: I do know about the bitter lesson... the problem with that is that it assumes you can throw more data and more training time to problems, and that they are stable or similar to what is in your data, this is not always the case.
However I cam imagine a kind of meta-learning foundation model that basically has a huge internal library of micro-features, and when you put a sequence into it, it matches those features against the sequence and builds up a low-noise summary of the data that it can use to make predictions.
That's of course heavily anthropomorphized, but it seems potentially in-scope for a transformer model.
The real problem with time series data is that you can't predict the future. Images and text are relatively homogeneous and exist within a kind of restricted space. "Time series" in general however could be just about anything, and there's not as much reason to believe that something like a "grammar of time series" even exists beyond what we already can do with STL etc.
This naturally comes to multi-model solution under one umbrella. Sort of MoE, with selector (router, classifier) and specialized experts. If there is something which can't be handled by existing experts then train another one.
the data is not a spherical horse in the vacuum. usually there is a known source which produces that data, and it's likely the same model works well on all data from that source. may be a small number of models. which means knowing the source you can select the model that worked well before. even if the data is from alien ships they are likely to be from the same civilization.
I'm not saying that it's a 100% solution, just a practical approach.
so while this seems persuasive, it's fundamentally about normal data which yields little value in extrapolation
A one-shot model that takes a prompt like "this sequence is a voltage of a PV-cell every hour" will have a chance though.
Such information will somewhat be embedded in the network (a new voltage graph will look similar to the voltage graphs it's seen before) but it should be better to make it explicit.
to be clear, I share these concerns
Maybe this one is better
I still think humans have better pattern recognition in the stock market than neural nets...at least for now
(Also, you don't just need to beat pure chance, because pure chance is not guaranteed — or I think likely — to result in net zero losses on average, since you are then in an information-asymmetric environment where your trading partner looks at data and you do not. But regardless, even if you "make" money but underperform the market, you are effectively losing money in opportunity cost!)
https://arxiv.org/abs/2111.00396 https://arxiv.org/pdf/2111.00396.pdf
Language models trained on a broad corpus make sense, as the language is similar across different domains. Time-series, however, are extremely different... Stock prices, heart-rates, brain-waves, digital signals... Patterns learned on a broad dataset should only introduce uninterpretable noise.
Still, I suppose we gotta keep an eye on this kind of work in case there's a tipping point.
Though mean absolute error is a bit of a weird measure. They're basically testing how closely a model predicts the median.
Clearly they had some reason to pick it over others, but I couldn't tell you why. Especially since ARIMA effectively minimizes RMS error.
I feel like this is what I'd want out of a DL time series NN. For just prediction it seems like overkill.
Can someone elaborate on what a grammar means in the context of time series forecasting?
To me, zero-shot would mean a model whose parameters was never tuned through training examples at all...
I can see how the machine learning definition is similar, but the whole idea of retraining and fine-tuning is foreign to a lot of subfields in GOFAI.
I guess attentional context is almost this, but LLMs don’t update their base model after a one shot inference.