Very cool ive read this line of paper originating from hippo, s4, hyena, mamba etc but can someone please explain how this isnt just an RNN/LSTM variant??
The way it keeps all the representation power of LSTMs is by having the transition vary with the input (but still be linear).