Even then, these videos are only like 50 frames long - and a real movie you would want to be hundreds of thousands of frames long.
Even then, these videos are only like 50 frames long - and a real movie you would want to be hundreds of thousands of frames long.
We can’t do it. AIs can sort of do it.
Latent diffusion models already demonstrated that operating on a compressed representation gives far better results, faster, but I don’t think we’re anywhere near the limit for what’s possible there. It’s no coincidence that this is how humans work.
They put an eye tracker on someone and captured their motion when walking in some rough terrain. You can sort of see that the person is focusing on the most likely place their foot will go next.
[1] https://www.youtube.com/watch?v=ph6uUHq3a-g
I think that we will discover that there is a more efficient way to encode temporal relationships, which appears to be "just throw transformers at it." My guess is that it will be in a more conceptual latent space that this attention will be applied.
Yes, but consider that most films are made up of many different shots, each of which are often just seconds long.
Obviously you could do 'human assisted' movie making where humans decide the storyboard and make directions for each shot, and then that isn't necessary.