This is a bold claim. Today LLMs have not been demonstrated to be capable of synthesizing novel code. There was a post just a few days ago on the performance gap between problems that had polluted the training data and novel problems that had not.
So if we project forward from the current state of the art: it would be more accurate to say autonomously (re-)design, (re-)code and distribute whole apps. There are two important variables here:
* The size of the context needed to enable that task.
* The ability to synthesize solutions to unseen problems.
While it is possible that "most complex" is carrying a lot of load in that quote, it is worth being clear about it means.