Mamba is the new and hot "linear transformer", could one day replace GPT based LLMs and scale sequence length up to 1M. It uses a clever math trick to parallelize token inference in the input while keeping a constant state size for generating the output.