Your critique about lack of grounding in these systems is an easy problem to solve. It’s as easy as teaching an LLM to associate words with real world objects or phenomena. Image-classification models, text-2-image models, audio transcription models, and many other modal specific systems already do this to some extent. And more recently there has been a push towards multi-modal language models(Deepmind’s flamingo), so this line of argument will be debunked very soon.
I actually believe GPT-4 will be multi-modal and it’s capabilities will dispel majority of these criticisms