Ask a LLM to spell "Strawberry" one character per line. Claude's output, for example:
> Here's "strawberry" spelled out one character per line: s t r a w b e r r y
Most LLMs can handle that perfectly. Meaning, they can abstract over tokens into individual characters. Yet, most lack the ability to perform that multi-level inference to count individual 'r's.
From this perspective, I think it's the opposite. Something like the strawberry-tests is a good indicator how far the LLM is able to connect individually easy, but not readily interconnected steps.