> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.
I'm confused, videos contain images and audio ...?
I'm confused, videos contain images and audio ...?