I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).
For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:
[0] https://aclanthology.org/2023.findings-emnlp.936/
https://arxiv.org/abs/2503.00392
https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...