This is an issue of tooling, not intelligence. Language models absolutely have the power to process email and send (push?) code, should you give them the tooling to do so (also true of human intelligence).
> So basically any form of longer term tasks cannot be done by them currently. Short term tasks with constant supervision is about the only things they can do, and that is very limited, most tasks are long term tasks.
Are humans that have limited memory due to a condition not capable of general intelligence, xor does intelligence exist on a spectrum? Also, long term tasks can be decomposed into short term tasks. Perhaps automatically, by a language model.
Have you actually tried agentic LLM based frameworks that use tool calling for long term memory storage and retrieval, or have you decided that because these tools do not behave perfectly in a fluid environment where humans do not behave perfectly either, that it's "impossible"?