It wouldn't be "video model", it would be an "anything that can be expressed in binary" model.
> Perceiver: General Perception with Iterative Attention
Biological systems perceive the world by simultaneously processing high dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. Perceiver is a deep learning model that can process multiple modalities, such as images, point clouds, audio, and video, simultaneously. It is based on the transformer architecture and uses an asymmetric attention mechanism to distill a large number of inputs into a smaller latent bottleneck. This allows it to scale to handle very large inputs and outperform specialized models on classification tasks across various modalities.