The Google video demo is hugely faked, to the point of being smoke & mirrors.
The inputs to the AI are still frames, not video.
The input images are different to the video frames, often substantially different.
The prompts were provided as text, not audio as shown.
Both the prompts and responses are significantly longer.
The latency is much higher than shown, which appears to be near real-time.
Etc...
The demo they showed would be a mindblowing improvement in LLM technology and applicability to use-cases such as controlling a home robot.
Instead, we were shown what is essentially a short science fiction movie "inspired" by possible future capabilities, not current capabilities.
Google literally said so: "The video illustrates what the multimodal user experiences built with Gemini could look like. We made it to inspire developers."
PS: Their other benchmark results are also highly suspect, because there is a high chance that Gemini trained on the exam question-answer pairs inadvertently. They even admit this in several places, such as in their technical paper. So... they knew the results are bogus and published it anyway!