Isn't this a real big deal? They manage to train a foundation model that connects the physical understanding gained from unsupervised training on images and videos with the real physical understanding necessary to control a robot. They have impressive videos to show for it. The approach seems indeed scalable and generalist. I've thought for a while that the keystone missing for household androids is connecting the understanding of large multimodal language models with physical hardware. To me this looks like exactly that. It actually makes me optimistic that we will see household robots within a decade. Now I wonder why Tesla, Figure etc. are messing around so much with Teleoperation if this indeed works. Maybe I don't understand what's going on.