Idk how the Gemini demo (a thing that actually does work and do the things displayed, just not in the exact way shown) is "equivalent" to a literally non functioning car...
> The inputs to the AI are still frames, not video.
A video is just a sequence of frames. The input is always going to be frames when it actually goes into the model, and you don't need every single frame to understand what's happening in the video.
> The prompts were provided as text, not audio as shown.
That's trivial to do now. Using Whisper, you can just turn voice into text and do the exact same thing. They don't really need to demonstrate that.
So sure, they definitely embellished, made it seem real-time and as if it didn't need more target-specific prompting per task. But saying it is completely fake is foolish.