As much as Transformers feel like the state of the art universal function approximators, people need to realize why they work so well for language and vision.
Transformers parallelize incredibly well, and they learn sophisticated intermediate representations. We start seeing neat separation of different semantic concepts in space. We start seeing models do delimiter detection naturally. We start seeing models reason about lines, curves, colors, dog ears etc. The final layers of a Transformer are then putting these sophisticated concepts together to learn high level concepts like dog/cat/blog etc.
Transformers (and deep learning methods in general) do not work for time series data because they have yet to extract any novel intermediate representations from said data.
At face value, how do you even work with a 'token window' ? At the simplest level, time series modelling is about identifying repeating patterns over very different lifecycles conditioned on certain observations about the world. You need a model to natively be able to reason over years, days and seconds all at the same time to even be able to reason about the problem in the first place. Hilariously, last week's streaming LLM paper from MIT might actually help here.
Secondly, the improvements appear marginal at best. If you're proposing a massive architecture change, removing observability & and explainability .........then you better have some incredible results.
Truth is, if someone identifies a groundbreaking technique for timeseries forecasting, then they'd be an idiot to tell anyone about it before making their first $Billion$ on the market. Hell, I'd say they'd be an idiot for stopping at a billion. Time series forecasting is the most monetarily rewarding problem you could solve. If you publish a paper, then by implication, I expect it to be disappointing.