Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
research.nvidia.com
research.nvidia.com
I imagine it will be the de-facto way AI generates video in the next few years, once huge models are smart enough to use tools and work in high breadth and depth jobs like humans do.
A major topic of research is to make differentiable renderers, which can then have their parameters tuned via gradient descent, much like how AIs are trained.
The problem is that most renderers are very discrete, and hence aren't differentiable. For the same reason that reverse rendering is difficult on such systems, it's hard to train an AI to use them also.
Having said that, I can imagine it happening. Look at Google's Alpha Star and related AIs, which can play games with discrete events at grandmaster level. You could train an AI to operate a 3D design program similarly. The input training could be recordings of actions taken by human designers, or simply train the AI by making it match target images. Start out simple, and then crank up the complexity over time.
Similarly, neural networks have already been shown to be able to directly simulate physics simulations (smoke, particles) with suprising accuracy.
I agree that logically your statement makes sense, however I'm not sure it matches what has been successful in the space so far. Tools may make sense for these NN's to use, however they probably won't look anything like the tools we use today.
possible because training models are fed by an abundance of existing human-generated content, which is ever increasing in quality, dirt-cheap, and easily accessible; which doesn't make AI Generated Content disruptive in this respect (my view).
not sure what you mean by "easier", since computers are better at following rules and people better at reasoning.
>> I agree that logically your statement makes sense ...
Isn't that the best case for AI generated content, logical-sense? or are you suggesting more (e.g 'source of meaning')?
I just mean that the approach sounds logical, although the best approach isn’t always the one that appears the most logical. This appears particularly true in AI - for instance diffusion models feel more like an incredibly surprising discovery rather than an approach that’s immediately obvious by applying logic.
> Not sure what you mean by "easier", since computers are better at following rules and people better at reasoning.
By easier I just mean we have had better results so far. Maybe easier wasn't the right word. On your second point - computers will be better at most reasoning tasks shortly.
But video has much stricter physical constraints than image, so it's not clear we can ignore the problem at all.
The stable diffusion community is getting impressive results generating keyframes in a batch and using existing interpolation mechanisms on the decoded frames.
https://www.reddit.com/r/StableDiffusion/comments/12o8qm3/fi...
That example is using control-nets to induce the animation they want but it would be quite easy to train a model on a sequence of video frames in the same layout.
But each advance has been an opportunity to unnerve people as they react to how much more real something looks, but still isn’t real.
I think the uncanny valley is now getting trained out of us.
We are all getting used to a continuum of real to stylistically unreal, with no more unexplored valley.
Except for the weird mistakes. Those will likely remain weird to us since most of us don’t want to immerse ourselves in worlds of uncountable fingers and third arms for long enough for that to start feeling normal.
About as exciting as Imagen-video, which was released eons ago.
I can't say I fully understand the mechanisms by which they achieve that, but it's clear that the powers that be have decided that the public cannot be trusted with powerful AI models. Stable Diffusion and LLaMA were mistakes that won't be repeated anytime soon.
We need a killer application for video generation models first, then I’m sure someone will throw a $100k at training an open source version.
Kickstarter is into the tens of millions these days [1]. I would assume some number of millions might be possible, if the right names were behind it.
[1] https://www.kickstarter.com/projects/dragonsteel/surprise-fo...
> people are essentially just pre-paying for a product
I think this is what the 4 nines of the user space wants. A pre trained, open source, model to work with.
Really? GPT-3 was released almost 3 years ago. Where is the public reproduction?
And don't say LLaMA. I've used it and it isn't even close to GPT-3, nevermind GPT-4.
Is there anything that incentivizes Nvidia to publish these results? Is it just needing to get papers out in the public for the academic clout? Something tells me that all this accomplishes is setting the standards of everyone who sees the possibilities, that "this will be the future", and a third party without the moral framework of Nvidia will become motivated to develop and openly release their own version at some point.
The most dangerous AI model today (in a practical sense, as people are actually using for shady stuff) is ChatGPT, which is closed source, but open to the public so anyone can cheat on their exams, write convincing fake product reviews, or generate SEO spams, etc.
The fact that a model is closed source doesn't change anything as long as it's available for use. Bad actors don't care about running the code on their own machine…
In this case, you need someone that can implement the method as described (hard!), and then you need someone with a setup better than a rented 8xA100 (expensive and not available on many cloud providers) to actually reproduce the model.
> (Team working on our variant, will be done when it's done and they are happy with it).
I _think_ he's saying that _his_ team is working on a similar model - and that they will release _that_ model "when it's done" (and not to to expect that to happen any time soon).
Just super vague, bordering on taking credit for work that NVIDIA did. Seems like he typed it out on his phone and/or is Elon-levels of lazy about tweets.
One of them even said: "Unfortunately we cannot release the weights. That's why I joined @StabilityAI to work on OS video models"
That almost everyone (other than, maybe, people specifically and very loudly selling the fact that they aren’t, e.g., Adobe) training base AI models is relying on the idea that doing so is fair use and doesn’t require a license is hardly news.
Some pretty experienced/expensive interns you got there, NVIDIA.
This is not particularly expensive considering the numbers these companies usually use, thanks to the fact that they only trained the temporal layers in the LDM and the autoencoder.
I remember when image generation was similarly derided.
I wouldn't count anything out at this point.
(Also, this was an intern project. There are a couple of senior staff researchers on the paper, but it's mostly interns.)
> Using DreamBooth [66], we fine-tune our Stable Diffusion spatial backbone on small sets of images of certain objects, tying their identity to a rare text token (“sks”).
I wonder how long that token will stay "rare".
No model, no code makes it a challenge to explore.
I think prompting is a bit past “trending on Artstation in the style of Greg Rutkowski” at this point..?
Like in this case? https://arxiv.org/pdf/2106.08254.pdf, figure 1
He's got to be the most cucked man on the planet.
There was also the time months ago when the art of a deceased artist (Qinni) was featured prominently on the front cover of an img2img/style transfer paper until the artist's sister requested it was taken down.