I think the next jump will come from neurosymbolic approaches, merging timeseries with a description of what they are about as input.
You can use that description for system identification, i.e. build a model of how the "world" works. This can be translated into a two-part network architecture, one is essentially a world-model-informed (physics-informed, as its often called in literature) part for the known-unknowns, the other one is a bounded error term for the unknown unknowns (e.g. dense layers, or maybe dense layers + non-linearities that capture the fundamental modes of the problem space for reservoir computing). The world model is revised with another external cycle of meta-learning, via symbolic regression.
The unknown-unkowns bit that I choose is designed as a shallow network that can be trained online (traditional methods, but I'd like to see if the forward-forward algorithm from Hinton would work well for short-term online adjustments), or by well-known tools like particle filters / kalman filters.
The non-linearities (and the overall approach in general) resemble physics-informed dynamic mode decomposition (piDMD), which show remarkable resistance to noise (e.g. salt & pepper).
If you have simple timeseries, and not very complex hierarchical systems that change over time and show novel modes that you haven't encountered before, then piDMD is likely enough for what you need.
---
Essentially, what I describe is a multi modal model for timeseries + a planning step. (AlphaGeometry to the rescue?)
---
Like you say, time-series in general have an incredibly complicated domain with comparatively few data available.
For example, real-world complex physical systems (industrial plants, but also large-scale software systems) may have replicas of the same components with complex behavior, and no/few shared dependencies for reliability or other constraints (e.g. physically apart). These can be captured by transformers. Training will be much faster if you initialize weights like I describe above, and share weights among replicas. The physical structure also creates particular conditions on the covariant matrixes and on the domain of higher-level timeseries (ultrametric spaces, which changes how measurement and frequency behaves there and can lead to great simplifications, but also errors if tools like FFT are applied blindly without proper adjustments; much like in operations research / planning problems, symmetries are sought after to reduce complexity).
On the other hand, the next level that compose these building blocks often have graph structure and sometimes scale-free networks (e.g. if they represent usage or behaviors, rather than physical systems). I think we'll see graph neural networks shine on this front.
There are likely other kind of behavior that I haven't encountered yet in my work.
I think overall, we'll see planning/neurosymbolic used at the highest-layer, graph neural networks for scale-free networks and to optimize long-range connections (also when a dense model with dynamic covariant matrices would be too expensive to compute even in sparse form), and transformers or/with piDMD-like approaches for dense patches of complex behavior. I.e. graph models as a generalization for spatial locality to arbitrary spatial-like domains, transformers/piDMD or similar for sequence-/time- locality for arbitrary complex systems. (I wonder what kind of weights will they implement when trained together on problems that are fundamentally in the middle, where traditionally one would use wavelets... if you look a the GraphCast model by deepmind for weather forecasting, it looks quite similar)