I'm not sure that the statement "some compositional problems will always be beyond the ability of transformer-based LLMs" is even controversial to be honest.
There's a reason all of the AI labs have been leaning hard into tool use and (more recently) inference-scaling compute (o1/o3/Gemini Thinking/R1 etc) recently - those are just some of the techniques you can apply to move beyond the unsurprising limitations of purely guessing-the-next-token.