Generating videos and 3D models is _much_ more difficult than images. You can’t just train off videos from the internet in the same way, because they don’t have sufficient text labels to understand them like CLIP does.
FWIW, I don’t think the AI systems will generate a whole video by itself - it’ll be some form of image to image generation where an artist will render a rough sketch of the scene and the AI will fill in the details, frame by frame.