Lumiere: A space-time diffusion model for realistic video generation
lumiere-video.github.io
lumiere-video.github.io
The only way to describe this is bragging, advertising, or marketing. There are no reproducible processes described. While the diagram of their architecture may inspire others it does not allow for the most crucial aspect of the scientific endeavor, falsification.
There is no way we can know if Google is lying because there's no way to check. It should be assumed that every example has been cherry-picked and post processed. It should be assumed that the data used to train the model (if one was trained at all) was illicitly acquired. We have to start from a mindset of extreme skepticism because Google now routinely makes claims that cannot be demonstrated. When the performance of Gemini in bard is compared to GPT-4 for example, it falls far short. When they release a video claiming to be an interaction with a model it turns out it wasn't anything of the kind.
Ideally no organization would operate like this but Google has become a particularly egregious repeat offender.
We can gather that they are likely to be lying or cherry-picking examples to make themselves look better, since they were already caught faking an AI demo. In the world of actual research, if you got caught doing this, all your subsequent and prior work would be under severe scrutiny.
How did people get access to Gemini Ultra? Or are you talking about Gemini Pro, the one that compares to GPT-3.5?
This difference is purposely muddied, but important to understand.
That's not at all settled law. AI companies are hoping to use the fair use exception to protect their businesses, but it looks like it will soon be clarified the other way.
Wired summed it up: "Congress Wants Tech Companies to Pay Up for AI Training Data"
https://www.wired.com/story/congress-senate-tech-companies-p...
And Ars wrote "Media orgs want AI firms to license content for training, and Congress is sympathetic."
https://arstechnica.com/information-technology/2024/01/at-se...
"[Senator] Hawley expressed concerns that if the tech companies' expansive interpretation of fair use prevails, it would be like "the mouse that ate the elephant"—an exception that would make copyright law toothless."
"Using copyrighted works for monetary gain" refers to using art itself as the product. Knowing what Apple's logo is and making a logo in that style is not a violation of copyright. However using Apple's logo (or something strikingly close) is a violation.
The reason this is muddied is because legally artists don't really have a leg to stand on for "my art cannot be trained on by a computer" whereas they do have strong legal precedent (and actual laws) for "my art cannot be reproduced by a computer".
Training is the "production" of a derivative work (a model) based on the training data.
AI companies claim that this is covered by fair use, but this is simply a claim that has not yet been tested in court.
And even if courts rule in favor of the AI companies, it sounds likely (based on what I've read) that Congress will soon rewrite the law to support the artists' position.
You don't have the right to make a copy of an e-book and keep that file on your server/computer for the purposes of training AI. Copying that file onto your computer is in many cases already an act of copyright infringement.
This doesn't sound like a productive stance for science. You don't trust their result? It's fine to ignore all the claimed artifacts and you can just take the core idea. You don't have to assume anything malice to invalidate their so-called advertisement.
While this kind of stance might make you feeling a bit better, but will also make your claim political and slow you down if it happens to be true, given the history that many of Google's papers eventually have become a foundation of other useful technologies even though almost of all them didn't contain reproducible artifacts.
That's generally easier said than done. The dataset isn't available, and there's not really enough details in the paper to make it replicable even if you have the resources to do it.
That said, if this tech is as advertised, extremely impressive to me
To me this looks like the first good video generation model.
EDIT: Just noticed its by Google, NVM, will never be released publicly.
Google is sponsoring a lot of cutting edge research and sharing it openly. How cool is that? How long will it last?
We seem to value freedom of speech (and expression) only to a tipping point that it begins to invade other aspects of life. So far the noise and rate has been low enough people at large support free speech but newer information techniques are making it possible to generate a lot more realistic noise (faux signal, if you will) at higher rates (it’s becoming cheaper and easier to do and scale).
So while you certainly have a point I mostly agree with, we’re letting private entities policies dictate the limitations of expression, at least for the time being (until someone comes along and makes these widely available for free or cheap without such ethical policies). It does go to show just how much sway industries have on markets through their policies with no public oversight, which to me is concerning.
The OpenAI content policies are pretty strictly opposed to the holding and wielding of severed heads.
The themes maybe, but the forced positivity is frustrating. Trying to get stock ChatGPT to run a DnD-type encounter is hilarious because it's so opposed to initiating combat.
> Without jailbreaks ChatGPT will always give narration a positive twist
Most modern stories in Western literature have a positive twist. It is only natural that gpt's output will reflect that!
Example: ask ChatGPT any kind of innocent medical question, like if aspirin will speed up healing from a cold, and tell it NOT to begin it's answer by stating "I am not a medical expert" or you will kick a puppy. This works for most models, but not ChatGPT. It WILL make you kick the puppy.
I understand why they have to do things like this, but I'd really prefer the option to waive all rights to being insulted or poorly advised and just get the (mostly) raw output myself, because it does downgrade the experience quite a bit.
Fortunately we have Mixtral now.
Even if we wind up at a point where no one trusts photos or videos is that really a disaster? Blindly trusting a photo or video that someone else, especially some anonymous account, gives you is a terrible way to shape your perception of the world. Ensuring that less people default to trusting random videos may even be good for society. It would force you to think about where the video came from, if it’s corroborated by other reports from various sources and if you’re able to verify the events through other channels available to you. You have to do the same work when evaluating any other claim after all.
https://arstechnica.com/information-technology/2023/12/googl...
This was not Leonardo da Vinci's "Mona Lisa"[1], but Johannes Vermeer's "Girl with a Pearl Earring"[2].
https://github.com/lumiere-video
Nor did they claim it would. But I had to check anyway, and there wasn’t any link I could see to the GitHub profile. So here’s a link for anyone else that wants to check and don’t want to type the url of their profile manually from looking at the hosted website url.
Dynamic adjustment of fixed aspect ratio film imagery to non-native sizes without stretch or obvious distortion. Guess all the added edges accurately enough that audiences won't notice.
4:3 <-> 16:9 <-> 143:100 (IMAX) <-> 11:8 (Academy) <-> 3:2 (35mm) <-> 16:10 (tablets/desktops)
Make a new movie look like a classic b/w silent, then give it the correct frame.
Adapt any movie to smoothly work on IMAX displays.
I know we're all used to new releases like this coming very soon and very fast, but I'm amazed. I can't wait to have a software with this abilities. edit: nvm, it's by Google. I'll wait for an open source to be released.
[1] https://en.wikipedia.org/wiki/George_Washington%27s_teeth
That said, the idea is very interesting -- train the model to generate a small full-time representation of the video, then upscale on both time and pixels.
Essentially, we have seen models adding depth maps. This one adds a 'time map' as another dimension.
Coherence is pretty good, to my eye. The jankiness seems to be more about the model deciding what something should 'do' over time, where a lot of models struggle on keeping coherence frame by frame. The big insight from the Googlers is that you could condition / train / generate on coherence as its own thing, then fill in the frames.
I think this is likely copyable by any number of the model providers out there; nothing jumps out as not implementable by Stability, for instance.
It's rather impressive and quite quickly will likely result in a huge hoard of "make a movie with a paragraph" programs.
It's Google - It will probably go in a box and be a Rick and Morty gadget we never see.
It has a cool author format list I like. The 1,2,3,4,*,+ thing is nice for lead authors, institute attribution, and core contributors. I read so many astronomy and physics papers that are 10+ authors long, and I have no idea who did anything. The arXiv link for example shows no similar formatting.
It will probably be immediately used for abusive porn. Walking Woman Example: (5th variation) "Wearing no clothing"
There are a few important techniques to be refined such as keeping consistent subjects between generations but I could see many inconsistencies being made up for by applying existing methods such as separating the layers based on depth allowing more static images to be used or creating simple 3D models with textures where more depth is needed. With enough effort and skill someone could probably do it with existing technologies.
Do these models actually learn a 3D representation or do they just learn "something" that is good enough to produce an very convincing impression of 3D ?
Subquestion: if they don't learn 3D, can we say that models learning a 3D representation first will lead to even better productions ?
The second, but at the limit it's the same thing of course.
> Subquestion: if they don't learn 3D, can we say that models learning a 3D representation first will lead to even better productions ?
Generally speaking manual feature engineering almost always turns out to be a waste of time if you can just make the model bigger; this is called "the bitter lesson".
Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model
"... In this work, we investigate a basic interpretability question: does an LDM create and use an internal representation of simple scene geometry? Using linear probes, we find evidence that the internal activations of the LDM encode linear representations of both 3D depth data and a salient-object / background distinction. These representations appear surprisingly early in the denoising process−well before a human can easily make sense of the noisy images. ..."
I do enjoy the irony though of you copy-and-pasting a generic pro-AI rebuttal to a comment you didn't understand.
It is so easy to game and tweak examples, especially since there is a random component to them. For example, you could do a prompt 1 million times and only show the best response. Or you could use prompts that it’s optimized for.
The reason ChatGPT and Dall-e captured the public’s imagination is that the public could actually put in their prompts and see the results.
I felt like this was maybe 5-10 years away.
That doesn't work on a phone. I hoped they added an event handler for touching the animations. Instead they forgot they have a mobile OS and that they sell phones.
Everything is relative!
Years ago, I wouldn't even dare to dream it would be possible. It's nowhere near, what people are used to watch normally, but the fact it's even trying to compete is insane.
Also, ever since the Gemini marketing video shenanigans, I don't really feel like trusting whatever Google's research says they have, if I can't test it myself.
The video was released by Google product marketing for a launch to customers, not research.
I'm still somewhat confused by this one. I understand the community has decided to be harsh on Google for that video to draw a line - fair, truth in advertising, etc. -, but at the same time, we all had an understanding of where that tech is at currently and the pace it progresses at. Did anyone watching it really assume it was realtime? Can we not differentiate between technical publications and marketing anymore? Do we have to vilify everyone in an R&D department for the sins of the product marketing wing?
It was completely dishonest. Considering how trash Googles actual AI products are they deserve to be dragged even more over that video.
>We train our T2V model on a dataset containing 30M videos along with their text caption. [...] We evaluate our model on a collection of 113 text prompts describing diverse objects and scenes. The prompt list consists of 18 prompts assembled by us and 95 prompts used by prior works (Singer et al., 2022; Ho et al., 2022a; Blattmann et al., 2023b) (see App. B). Additionally, we employ a zero-shot evaluation protocol on the UCF101 dataset >
Google can publish whatever research they want, literally doesn’t matter literally changes nothing because they can’t turn it into a product anyone can use and never will.
The home assistant speakers aren’t making enough money to justify the large teams behind them. Thus we’ve seen significant layoffs on those teams in the past year.
BigCos are looking for other ways to reduce costs. Killing features is one way to do it.
There have also been situations where a feature is removed because of legal action; lawsuits alleging the features violates a patent.
Live updates givith, live updates taketh away!
Of course building products for actual users is no longer a “thing”, but think of the stock price.
I give it 5 years until it has been normalized to see AI generated TV/YT ads, 10 to 15 when traditionally made ones will be in the minority.
In the beginning just a bunch of geeks in front of computers crafting the prompts, later everyone will be able to make it.
It will probably be access to computing resources which will be the limiting factor.
But also you're making the mistake of extrapolating against the realities of the techniques.
Things may improve over time but prompts and random seeds aren't great for detailed work, so there are limitations which seriously limit the usefulness. "Everyone will be able to make it" is likely true, but the specialist stuff will likely remain and those users will likely be made more productive. It's those in the middle that will lose out.
That an industry is destroyed is neither here nor there. Sucks to have your business/job taken away but that's how the system works. That which created your business also will destroy it.
I think it might be within the realm of the possible to see 30 second videos at the end of the year.
The next step could then be infinitely long videos when frames are getting generated at 24 fps, as long as the ability is given that they are able to stick to a story and a visual style that makes sense. The story could evolve automatically from an LLM or be generated in real time by an artist, like a prompt every minute. In any case, we're not that far away from this, even if the first results will be more like trippy videos.
[video prompt: Two elderly people taking a stroll on a boardwalk, partaking in various boardwalk activities.] [AI gen voice: Suffering from chronic blorgoriopsy? Try Neuvoplaxadip by Excelon pharamceuticals. Reported side effects include... Ask your doctor.]
But on a more serious note, I vividly remember when GANs were the next big thing when I was in university and the output quality and variability was laughable compared to what midjourney and the likes can produce today (my mind was still blown back then). So I would be in no way suprised if we got to a point in the next decade where we have "midjourney" for video generation. So I wholeheartedly agree.
I also think the computational problem is tackled from so many angles in the field of ML. You have nvidia releasing absolute beasts of GPUs, some promising start ups pushing for specialized hardware, a new paper on more optimized training methods every week, mamba bursting on the scene, higher quality data sets, merging of models, framework optimizations here and there. Just the other day I think I saw a post here about locally running larger LLMs. Stable Diffusion is already available for iPhones at acceptable qualities and speed (given the devices power).
What I wonder about the most though is whether we will get more robust orchestration of different models or multi modal models. It's one thing to have a model which given a text prompt generates a short video snippet. But what if I instruct my model(s) to come up with a new ad for a sports drink and they/it does research, consolidates relevant data about the target group, comes up with a proper script for an ad, creates the ad, figures out an evaluation strategy for the ad, applies it and eventually gives me back a "well thought out" video. And all I had to do was provide a little bit of an intro and then let the thing do its magic for an hour. I know we have lang chain and baby AGI but they are not as robust as they would need to be to displace a bunch of jobs just yet (but I assume they will soon enough).
Single image genAI went from unusable to indistinguishable from reality in 18-24 months.
Me, scanning for a download link or a prompt to run the model and not finding any, excitement level medium
Me, realizing it's by google, excitement level zero
I would rather an actual animator create something beautiful for me rather than an AI spit out something that needs to be worked on by an actual animator ANYWAY.
I saw someone else say "I'm sure it'll be crap like all of the other AI stuff I've seen" but that's a naive view. Things that have been 100% created by AI, sure they're kind of boring a lot of the time. But this kind of tech gives people with a creative mind, but no money or time or resources to create a storytelling movie/video, the resources to do it. Obv ignoring the fact that Goog will never release this, if something like this did come out, it'd be game changing for a lot of people.
Think about something like RPG Maker. Yeah we've had a ton of random garbage come out of that platform but there were also incredible.
AI isn't just some garbage maker. It is a paint brush that enables people who are alone in their room to make something bigger than them.
FWIW, my kid also designed their own board game pieces in TinkerCAD and we 3D printed them. It's nothing special but it's frankly astounding how far kids can go now towards creating something not just imaginative but almost professional quality with the tools at their disposal now. For throwaway school projects. It may not be my kids, but I'm excited for what the next generation will be able to accomplish without massive capital requirements to fulfill their vision and create something.
Like we build these things and show them off, without any thought to the ramifications that they could lead to. Maybe I'm catastrophizing, but all this tech lately seems very unregulated/dangerous.
The target market is people and organisations who like/want/need the speed and low cost of generated "art" and prefer not dealing with external real world artists that need to be fairly compensated and will take time to produce an art piece.
Also laws are very murky on this for the moment (naturally, since it's a very recent new thing), and some consider that AI "art" can't be copyrighted. The EU is currently working on a new AI framework which will probably cover that.