> Log-linear attention replaces the fixed-size hidden state with a logarithmically growing set of hidden states
Does this mean the models can be smaller too (on top of the primary benefit of being faster)?
Does this mean the models can be smaller too (on top of the primary benefit of being faster)?