Emu Video and Emu Edit, our latest generative AI research milestones
ai.meta.com
ai.meta.com
> To the best of our knowledge, this is the first work highlighting fine-tuning for generically promoting aesthetic alignment for a wide range of visual domains.
... unlike Stable Diffusion which did aesthetic fine tuning when it was released? Or like the thousands of aesthetic finetunes released since?
> We show that the original 4-channel autoencoder design [27] is unable to reconstruct fine details. Increasing channel size leads to much better reconstructions.
Is it not expected that decreasing the compression ratio would lead to better reconstructions? The whole point of the latent diffusion architecture is to make a trade-off here. They're more than welcome to do pixel diffusion if they want better quality, or upscaling architecture.
And then the rest of the paper is this long documentation that can be summed up as "we used industry standard filtering and then human filtering to build an aesthetic dataset which we finetuned a model with". Which, again, has been done a thousand times already.
I really, really don't mean to knock the researcher's work here. I'm just very confused as to why the work is being represented as new or groundbreaking. Contrast to OAI which documents using a diffusion based latent decoder. That's interesting, different, and worth publishing. Scaling up your latent space to get better results is just ... obvious? (As obvious as anything in ML is, anyway). Facebook's research isn't usually this off the mark. E.g. the Emu Edit paper is very interesting and contributes many new methods to the field.
[1] https://scontent-lax3-1.xx.fbcdn.net/v/t39.2365-6/10000000_1...
1. data is all you need to generate these amazing videos with the right gait (gait is something I focused on).
2. nobody doing new network structure, it is animatediff beefed up a little bit with temporal masking applied (neat trick, not a big leap from inpainting task we already see).
3. additional conditioning vector helps, and can be trained, look at these editing tasks!
These are pretty valuable for a looker like me to decipher what they did for Gen-2 or Pika Labs etc.
https://m.youtube.com/watch?v=NXX0dKw4SjI&pp=ygUII3Npbm50ZWs...
And when somebody comes along and fixes their program or reprograms what they did, they simply insert or change some of the prompts along the way and get a different effect.
When the characters add new data to the computer (like the episode where Geordi added the psycho profile of the enterprise engine designer), they're just tuning the foundational model with some new input data.
Yeah....that feels right for now to me.
> There are 5047 classifications of tables on file. Specify design parameters.
Interestingly enough, it seems existing AI models are already better than the Star Trek computer at dealing with ambiguity. Stable Diffusion would just generate a "normal" table and let you go from there.
I can think of solutions for the physical component or simulating the perception of a physical component
2030?
Also why do these AI people always end with "this does not replace anyone". Surely they do not believe this?
but there are lots of specialists I used to contract with in the ideation phase that I no longer do.
professional logo designers
testing out names of potential services
designers for landing pages for websites
additional coders for landing pages of websites
templates for powerpoint presentations
graphics for them
many many billable hours for lawyers for things I would have otherwise asked them about and thats totally a risk I’m willing to shoulder. now I simply have them implement unless they are not able to corroborate the legal view. In the past, I would have to explore several paths and then consider switching lawyers after I had all the information I wanted, having the subsequent lawyer implement without any knowledge of why.
some of these ideas generate revenue and I can get to that point far faster and cheaper
I can already code in the latest frameworks and have high proficiency in most media suites, but the media creation was not where I specialize or want to spend time on
so there is a general denial thats kind of useful, if a big company wanted tax breaks from a municipality they can say “look, jobs, we’re big on that”
but everyone knows whats happening
I didn't see a repository, but I think in this case, the paper is actually a perfect balance of detail? I think Meta benefits from startups building using their tooling (startups usually buy ads), and so the lack of a full implementation leaves a bit of room for startups to turn the work in to something a bit more production ready.
The cool techniques from the paper are:
Generating a bunch of example images in one go, and using CLiP to score your generated images
And mixing pre-built pipelines and grammars to execute common tasks.
These two ideas alone (with examples) give people in the space plenty to run with.
Great paper!
Maybe career wise? Should art have ever been really considered a career ? It was a nice side effect people might pay for it, outside of that ?
These are pretty great results though, don't you think?
If it at some point they go past looking terrible, will you think these in between "terrible" models were a waste of time?
The generated videos are aesthetically horrendous. I don't know what kind of mental gymnastics are going on that they can confidently describe something where the body shapes are nonsensically in flux with every change of frame (look at the eagle's talons, or the dog's leg movements as it runs) as "high-quality video".
Is generative AI hype blinding them to how hideous these videos are, or do they know and they just pretend like it's something it isn't?
But the best video generators were much worse than Emu Video; there was Make-A-Video[0] from Meta, and Phenaki[1] and Imagen Video[2] from Google.
[0]: https://ai.meta.com/blog/generative-ai-text-to-video/