"showing that multilayer transformers indeed cannot solve certain complicated compositional tasks. Basically, some compositional problems will always be beyond the ability of transformer-based LLMs."
Pretty sure this is just false and the paper doesn't show this. I could be misunderstanding, but it looks like the result is only about a single token/forward pass, not a reasoning model with many thousands of tokens like o1/o3