Reading this in the context of Sora, this is contingent on:
1. movie or tv show length generation needs to be reasonably affordable,
2. the generated output must look real beyond first look (text is readable, physics is correct, no dream-like morphing),
3. continuity between scenes
4. prompts can accurately produce the artistic vision from directors and producers faster than a human just doing it
5. the public will accept it
I am reminded of a scene in Westworld, where basically a someone can sit a desk and describe in natural language a scene to a computer for instant rendering. The user can now concern himself exclusively on the details of the story. Perhaps roles will transition to becoming full time story tellers, and the minutia of vfx, animating, rigging, 3d models, actors, cameras, set designers, trucks, food, lighting, electrical, location, licensing, permits, etc. all go away.