Researchers built the Winnograd Schema Challenge more than a decade ago to assess common sense reasoning, and LLMs beat that challenge task around GPT 4.
If you ask them in isolation they may write a script to solve it "properly", but I guess this is because they added enough of these to the training set. But this workaround doesn't scale.
As soon as I give the LLM a proper problem and a small part of it requires numeric reasoning, it almost always hallucinates something and doesn't solve it with a script.
If the logic/math is part of a larger problem the miss rate is near 100%.
LLMs have massive amounts of knowledge, encoded in verbal intelligence, but their logic intelligence is well below even average human intelligence.
If you look at how they work (tokenization and embeddings) it's clear that transformers will not solve the issue. The escape hatches only work very unreliably.
I have been broadly quite happy with gpt 5.4 xhigh's reasoning on things like performance engineering tasks.
But try asking your favorite LLM what happens if you're holding a pen with two hands (one at each end) and let go of one end.
Seems fine to me?
https://chatgpt.com/share/69bcd01a-a750-800d-95f5-3b840b9ee2...
https://gemini.google.com/share/edc223bb6291 (the try again gave a woman, oops)
Even Midjourney couldn't do it.