This deepmind paper showed big limitations on nested structures with transformers, LSTMs actually did better:
Neural Networks and the Chomsky Hierarchy
Neural Networks and the Chomsky Hierarchy
It was a good bit older of a paper though, if I remember it's somewhat expected from the pure feed forward nature of them and limited circuit depth, where LSTMs have some recurrence.