I find it odd that the article doesn't address the apparent success of training with transformer based models in virtual environments to build models that are then mapped onto the real world. This is being used in everything from building datasets for self driving cars, to navigation and task completion for humanoid robots. Nvidia have their omniverse project [1], but there are countless other examples [2][3][4]. Isn't this obviously the way to build the corpus of experience needed to train these kinds of cross modal models?
[1] https://www.nvidia.com/en-us/industries/robotics/#:~:text=NV....
[2] https://www.sciencedirect.com/science/article/abs/pii/S00978...
[3] https://techcrunch.com/2024/01/04/google-outlines-new-method...
[4] https://techxplore.com/news/2024-09-google-deepmind-unveils-...