SSM models strengths are with continuous data like audio and video, they struggle with discrete data like text/ DNA. This newest architecture uses selective attention to try to address the weaknesses around discrete data with some loss in performance in continuous tasks, empirically shown here with audio. The empirical exploration was limited to smaller size models, the performance as larger scales is yet to be explored in practice.
I found my deepest understanding of the selection mechanism came from struggling with the discretization in section 2, followed by the deeper explanations of the variables involved in 3.5.2. This video gives excellent background to SSMs, along with a detailed walk through of the paper itself.[a]
I am still coming to understand S4, SSMs in general but the video suggested this annotated explainer that has been helping a lot[b].
I would also point out section 3.1 and it’s discussion of the tradeoffs between compression and effectiveness as particularly interesting.
I do wonder how many different GPUs / hardware architectures will be able to execute the optimizations that are described as critical. I think the nature of the optimizations is the part of the paper I understand least well.
The paper taken at face value looks very exciting. The promise of a very large context window with 5x throughput for inference would be huge if it proves to scale well. I do wonder if it will make sense to train SSMs without this selection mechanism for specific continuous use cases where it seems to perform better or if other architectures will prove to better serve those cases.
----