In terms of alternative strategies, Google DeepMind also has an amazing robotics team with lots of fantastic work for real-world robotics - including multi-robot generalists, showing positive effects when co-training one agent or model on multiple environments/bodies. Their prior work was very inspirational to us in SIMA! https://deepmind.google/discover/blog/scaling-up-learning-ac...
What would happen if, for example, the system were trained on simple 3D puzzle games, with natural language instructions such as “solve this puzzle” with a hint like “X or Y strategy might work”? I see that menu navigation is part of the training set. Can this thing learn to “read” or does it just learn the results of menu actions? If it can learn to play No Man’s Sky.. can it learn to play Zelda?
Video games might be even harder to predict sometimes, since the physics simulations have very strange edge cases. There are no physics glitches in the real world.
Massively, in a game running into a wall is perfectly normal and valid strategy to get close to it, in reality that will wreck you.
Or running on a fist sized rock has no consequences, in reality that destroys your foot. Reality is full of such extreme threats everywhere even in normal homes.
Video game physics can also be made more difficult - it's not a stretch to think about a reality simulator that dials these real-world effects up to 11 for AI to train in, at which point you need the meatspace bot and many things start to happen.
Tesla has an alternative. If you can get your devices widespread and can be recording observations and actions, you can collect huge datasets in the real world.
The first embodied, vaguely general, multi-task, useful and economical robots might just really open up the virtuous cycle of experience, learning, feedback and improvement. If I had to guess where it would come from right now, I'd pick Amazon warehouses.