https://blog.research.google/2023/03/palm-e-embodied-multimo...
That's always the question, isn't it? The article does a pretty convincing job of showing that at least in the given examples, it has pretty good "understanding" of what's taking place in the scenes and what makes them remarkable to people, and 7 years for comparison is going back a long ways, just the last 2 or 3 years has been where much of the most interesting progress has revealed itself.
Image segmentation, object detection and tracking, are all on display already here.
In the given images above, it may be clear from context (text? tags? exif info etc?) of images in its training data, that it's unusual for people to be dragged on a rope behind a horse, very unusual / dangerous for 747 sized airplanes to fly on their side, or houses to be lying on their side on a beach. And hence, describe such a view with "unusual", "dramatic", etc. Would it even need to understand a conceptual meaning of those words? Apply label, done.
Don't people work the same, in a way? Over the years we'll rarely see a house burning in person. We see news reports of such events mentioning people dead or severely burned. So after a while that 'training set' is enough to say: "person stumbling out of a burning house = something bad happened".
Yes, humans may then reflect on how they would feel if placed in unlucky person's shoes, and rush to alleviate that person's pain. Or cringe by the thought of it.
But in the end: maybe, just maybe, what human brains do isn't so special after all? Just training data, pattern match with external input & use results to self-reflect.
(that last step not -yet- covered by GPT & co)
Broadly it can be said that LLMs work by reducing uncertainty. Interestingly, human consciousness is also theorized by some to work the same way, react to the input in ways to reduce uncertainty.
I have a quote I like in this context, “Could you learn gymnastics just from watching videos?” However intelligent an internet trained model may be, I strongly believe you have to have some interaction with the real world to learn more complicated actions. So far, Pick and place, is all that’s been shown with the bigger models, so my hypothesis seems to be holding for now.
https://general-pattern-machines.github.io/
https://wayve.ai/thinking/lingo-natural-language-autonomous-...
It's just the straightforward application.
>I have a quote I like in this context, “Could you learn gymnastics just from watching videos?”
It's not like language models learn by some sort of magic osmosis. They're not just "reading" text or "watching" images. They learn by predicting, failing and adjusting neurons based on the data. Text is their world and they are interacting with it.