Phenaki: A model for generating minutes-long, changing-prompt videos from text
phenaki.video
phenaki.video
The other two, also by anonymous authors using the same formatting, are:
AudioGen: Textually Guided Audio Generation https://openreview.net/forum?id=CYK7RfcOzQ4
and
Re-Imagen: Retrieval-Augmented Text-to-Image Generator https://openreview.net/forum?id=XSEBx0iSjFQ
There is a samples site for AudioGen, but it is currently flooded and inaccessible:
https://anonymous.4open.science/w/iclr2023_samples-CB68/repo...
But the architecture figures look like they have different styles. E.g. the Re-Imagen paper uses rows/stacks of small colored circles to represent output tensors, and colored rectangles of different ratios to indicate shape differences, where the phenaki paper uses stacks of squares for output tensors, and differently shaped elements to distinguish different kinds of components.
> Little girl, in a field, holding a flower. We zoom back, to find, she's in the desert, and the field's an oasis. Zoom back further, the desert is a sandbox in the world's largest resort hotel. Zoom back further, the hotel is actually a playground, of the world's largest prison. But we zoom back further--
It’s incredible that the joke was that his idea was simply impossible, and yet technology has advanced to the point where you can basically do it instantly.
Check it out alongside the project page to see the text that formed it alongside, or just watch it here:
https://phenaki.video/stories/2_minute_movie.webp
However, there are some flourishes and timing that are not indicated from the prompt text, and I think there is some manual tweaking at play (which is okay, it's still impressive).