Not to crit the paper: formally stating things is good. But it feels like the outcome was highly predictable. Inferences being drawn from the language model where the language is a real world language and not a restricted subset were always going to hit up against the complexities of real languages.
And I'd be mindful that conforming to Chomsky Hierarchies says nothing about the nature of universal language models, or {{GPT or turing machines} and emergent intelligence}. Which again, isn't to lay things on the authors. People have a habit (here, in general) of taking things out of context.
I kind of feel thats about it: If the analysis of the classifier architecture shows constructs which map to a machine capable of being the higher Chomsky class, then its likely that reflects the complexity of the languages being analysed.