The mamba paper shows significant improvements in all model sizes, up to 1b, the largest one tested.
Are there any reason why it wouldn't scale to 7b or more? Have they tried it?
Are there any reason why it wouldn't scale to 7b or more? Have they tried it?