Imagine you have a lot more computing resources in a multimodal LLM. It sees your request of count the syllables and realizes it can't do them from text alone (hell I can't and have to vocalize it). It then sends your request to a audio module and 'says' the sentence, then another listening module that understand syllables 'hears' the sentence.
This is how it works in most humans, now if you do this every day you'll likely make some kind of mental shortcut to reduce the effort needed, but at the end of the day there is no unsolvable problem on the AI side.
Me, Kubernetes Haikus, time taken 84 seconds:
----------
Kubernetes rules
With its smooth orchestration
You can reach web scale
----------
Kubernetes sucks
Lost in endless YAML hell
Why is it broken?
No.
I doubt you would fully trust a LLM to replace high risk jobs such as lawyers, doctors or pilots such that when something goes wrong as it is used unattended, there is no-one held to account for it to transparently explain its own mistakes and errors.
It is just nonsense to suggest that such systems are capable of ‘reasoning’ when it pretends to do so and repeats itself without understanding their own errors.
Thus, LLMs and other black-box AIs cannot be trusted for those high risk situations over a consensus of human professionals.