For those who are unaware, previous "state-space models" (SSMs) took the classic state space model from control theory[a], parameterized its matrices, and applied it in discrete time to tokens in a sequence. Many people, including me, got excited by SSMs, given their linear cost... but they failed to live up to the hype.
The key difference now is that, instead of parameterizing the transformation matrices as before, the authors propose computing them from the data. In other words, instead of applying the same transformation at each step in time (which the authors call "linear time invariance"), they now dynamically compute and apply a different transformation at each step in time. Think of it as a new kind of nonlinear state-space model in which the linear maps are dynamically computed from the data at each step. The whole thing is very similar to RWKV[b], which, as the authors point out, can be formulated as a composition of these new "selective" SSMs. And RWKV itself is based on Linear Attention.[c] There's a lot to process in this paper.
More importantly, by establishing a connection between these linear-time DNNs and a well-understood model from control theory, the authors open the door to some exciting possibilities: Can these models be extended to the continuous-time setting? Which classic results from control theory apply to them? Can we draw from classic theory to improve our ability to reason about the behavior of these models? Other authors may have been the first to show these new linear RNNs can work well, but Gu, Dao, and others out of Chris Ré's lab deserve credit for making the connection to classic models.
I'm looking forward to diving in!
---
[a] https://en.wikipedia.org/wiki/State-space_representation#Lin...