1. They are single-pass and static - you "fake" short-term memory by re-feeding the questions with it answer 2. They have no real goal to achieve - one that it would split into sub-goals, plan to achieve them, estimate the returns of each, etc.
As for 2. I think this is the main point of e.g. LeCun in that LLMs in themselvs are simply single-modality world models and they lack other components to make them true agents capable of reasoning.
Based on those kinds of results an LLM should, in theory, be able to plan, analyze and suggest improvements, without the need for human intervention.
You will see rudimentary success for this as well - however, when you push the tool further, it will stop being... "logical".
I'd refine the point to saying that you will get some low hanging fruit in terms of syntactic prediction and semantic analysis.
But when you lean ON semantic ability, the model is no longer leaning on its syntactic data set, and it fails to generalize.