"Count the 'r's in the string 'raspberry'" is the same as "Count the joints in the Tasmanian sand spider": it means, "can you instance in idea and assess it and perform procedures over it - for example, can you count?".
We want to be sure it can count. We have seen it may guess, it may construe, it may recall - we instead demand it checks. We cannot trust it without that.
Nothing intrinsically more or less direct about the LLM's method than ours.
In my mind general intelligence is pretty much by definition a virtual machine, so the mechanisms behind thought are only relevant for the sake of efficiency (ie you can argue that LLMs make a poor basis for intelligence because tokens and natural language are a poor way to encode the world, but if you can run it on a big enough computer to counteract the inherent wasteful virtualisation then who really cares how it works under the hood?)
I will stop here before our analogies go too far.
I've never tried it and it might take some thought and effort to conduct an experiment to find out properly, but I would be interested in the answer.