It's been shown time and time again that they'll produce a correct chain of reasoning when given a problem (e.g. wolf, goat, cabbage crossing a river; 3 guards and a door; etc.) that is roughly similar to what's in the training data but will fail when given a sufficiently novel modification _while still producing output that is confidently incorrect_.
My own recent experience was asking ChatGPT 3.5 to encode an x86 instruction into binary. It produced the correct result and a page of reasoning which was mostly correct, except 2 errors which if made by a human would be described as canceling each other out.
But GPT didn't make 2 errors, that's anthropomorphizing it. A human would start from the input and use other information plus logical steps to produce the output. An LLM produces a stream of text that is statistically similar to what a human would produce. In this particular case, it's statistics just weren't able to cover the middle of the text well enough but happened to cover the end. There was no "complex reasoning" linking the statements of the text to each other through logical inferences, there was simply text that is statistically likely to be arranged in that way.