Unfortunately its not very good at longer context lengths, which sort of defeats the point of efficient scaling with context. See https://twitter.com/arankomatsuzaki/status/16390003799784038...
Its also not really an RNN. The best way to describe the key time mixing operation is a normalized exponentially weighted moving average (EMA)- no non-linearity. Once viewed this way, its not surprising that it struggles at longer contexts- everything decays, and it has limited space to put things. Of course, it does have some clever tricks, and can choose to remember things for a while by upweighting them, but not forever.