But LLMs are nowhere near being able to do what you suggest, for anything that one person wouldn't have been able to do beforehand.
The issue is training a multi modal model that can make use of said gui.
I don't believe that there is a better general interface than text however so I won't bother.