Imagen Video: high definition video generation with diffusion models
imagen.research.google
imagen.research.google
Stability or someone like it will valiantly release this technology, again and there will be absolutely no harm to anyone.
Stop being so totally silly Google, OpenAI, et. al. - it's especially disingenuous because the real reason you don't want to release these things is that you can't be bothered to share and would rather keep/monetize the IP. Which is ok -- but at least be honest.
The thing about owning the data sets and the huge TPU/A100 clusters is that the “publish the papers” model strictly serves them: no one can implement their models, they can implement everyone else’s.
I do understand the fear of being sued or targeted in the media over misuse, though. The person misusing technology should (obviously imo) be held responsible for that, but since it's new tech, the tech will be taking the blame for the first really controversial cases of disinfo and/or harassment that utilize it.
I can already see it. Just think of all the energy wasted training AI at home! I can imagine police drones with IR sensors scanning the cities for the heat signatures of illegal AI "farms".
Talking seriously, however they try to spin it, advanced AI (same as every other big scientific/engineering achievement) will be predominantly good. So let's say there is time when AI can create convincing videos of people engaging in various compromising "activities". When this becomes widespread it will give plausible deniability to any potential victim of such attack(with real or deep faked materials) .
In a world where any compromising video or picture can be made with anyone, the value of such materials for wannabee blackmailer diminishes rapidly. However, in a world where there is only few entities that can produce such materials and they do so sparingly those entities get a tool that gives them huge power (especially in democracies where popular opinion decides who governs).
Something tells me these pricks will end up arguing for a reversion to thin-client compute. It is in their financial interest too after all.
"What have you done this week?"
I think if you are Google, you are terrified of the bad PR from someone generating something questionable. And that bad article is inevitable if you open up these models. (See, pornpen.ai being released approximately five minutes after StableDiffusion. Imagine the press if that was built from the mode Google published.)
An open source community is a diffuse target, so the NYT won’t go after them as quickly, and let's be honest, their axe to grind is with big tech, not a bunch of AI hackers.
They don’t imply any ridiculous idea that such models should or even can be “racially balanced”. But if they want to cover their butts from the possibility of silly controversy, I think that’s cowardly and unnecessary, but at least they could not go out of their way to imply that such controversy should be taken seriously.
And the risk behind that is...?
If you drill down with such claims the core is always "someone might use this to lie online" and the proposed solution every single time is: more surveillance. End anonymity. Have a Facebook account required to use the internet. Real name and real face policies for every online interaction.
I’ll explain again that I think they can be used for bad actions, and also that they should still be released, because the benefits will outweigh the negatives. It does not hurt to admit that some things can be dangerous when used in nefarious ways. No one suggests we ban kitchen knives even though they are lethal, because their utility is massive, and outweighs their danger. In much the same way these models have extreme utility, that almost certainly outweighs their potential negatives.
ROFL what a weird thing for any HN commenter to say
The issue is, as an industry and society, we somehow bought the "safety" and "harm" charade a little bit too much, and somehow think it's a reasonable argument instead of being completely insane.
We can both admit that the tools can and will be used for bad purposes, and come to the conclusion that their benefits outweigh the negatives. We would not be doing any favors to our own arguments by pretending otherwise.
This has been possible without AI for a very very long time now (just open photoshop, etc). It barely ever happens, and society hasn't collapsed.
I keep seeing this argument come up and it baffles me that informed technologists take it seriously, as if it were impossible to convincingly manipulate images before DALL-E came around.
Further, we have seen harm come from some of this already, there’s a pretty big online community that uses deepfakes to put people in situations they would rather not be in, the most obvious being porn.
Been possible on home computers for 31 years for anyone who actually wants to do it. It literally doesn't matter and I think Stability has proven that the "AI Ethics" part of these models was essentially meaningless busy work at best, and at worse stealing compute credits from users like Dall-E purposefully charging you with something you didn't ask for.
Once every home computer can make the fake images AI Ethicists larp about then the power of fake images disappears because everyone knows not to trust them. It only has power if only a few can make them and never told the world it was even possible.
Yes, we are. Open source stable diffusion can be trained on any person's images, as long as you have around 20 from different angles. Costs around 50 cents on rented GPUs.
"Ethicists" act like society will somehow not adapt to this tech like they have with all the tech that came before it. I put ethicists in quotes because the arguments they use don't hold up to scrutiny and don't seem to be motivated by real ethical concerns. At least not to me.
Maybe once it's trivially easy to copy someone else's "likeness", society will finally be able to accept it and evolve past it.
Biological twins have never had control over this. If one twin wants to be a porn star, there's nothing the other can do.
Edit: one could imagine a future dystopia where clones are created to bypass "identity IP".
But those aren’t the issues they claim are concerning to them. It’s just stupid identity politics. They want their model to lie to us about the world and say things like everyone is equally likely to any attribute. They have a “reality” problem apparently.
Yes we are. There have been papers coming out on this tech for years now, with even the south park people doing videos using it.
In terms of machine learning technology that introduces truly novel innovations Google's product portfolio is notable barren. For instance the incredible powerful potential for image generation these new diffusion models open up, who's models will the world use to explore the potential and start using this technology? Google's model with the intense, though imperfect, effort that goes into addressing questions of bias and abuse? Or the model bankrolled by an ex hedge fund manager who probably put a bit less thought into addressing these questions?
3. Why do you use quotes around a something which is your original phrasing? That's pretty disengenious.
4. What's wrong with affirmative action ? It's easy to argue that it has both utilitarian and other moral adventages. I won't claim it is always warranted or the right thing to do, but it definitely not an obvious consensual evil.
It isn’t impermissible for Google to do this, but nobody has to like it, or agree with it.
*for instance, affirmative action is disapproved of by 70-80% of Americans and couldn’t win on a ballot in California in 2020 which is pretty exceptional.
What they want is for results of “software engineer” to be equally likely to show black females. This is not fair to Eskimos and Aboriginals. And what about the mentally handicapped? Is it not unfair that people with Downs Syndrome are not Wall Street stock brokers? How are you going to find all of these “affluent” categories and claim to be able to balance them?
And are you going to claim racism again when to ask for prison inmates and you don’t find any Asians? Should you start putting latent space Asians in latent space prisons?
Because this is a main harp of the social justice “ethicists” - that if you ask these models for “gang member” you get “People Of Color!!” … as if they simply don’t understand the statistics of situation. How would you even “solve” that? Should you decide when and where certain ethnic groups should be taken down a peg?
Latent space affirmative action is technically absurd, an completely ironic as “ethical” behavior.
It is a matter of PR. It takes a single "problematic" generated content to be framed as "Google is sexist/racist/supports animal abuse", etc.
Statements related to ethics help in a few ways: holding secrets ("oh, we would love to share the models, but we cannot"), protecting against backslash (PR-wise, legal-wise), and PR on its own ("we are that ethical - see! it is even in our mission statement").
Recommendation engines are more responsible than basic infrastructure.
?
Regardless, though, it is unambiguous that FB's role in "making" problematic UGC is much less direct than Google's role in making Imagen outputs.
And that's for all ML stuff, Colab has lead to a huge democratization of ML, with notebooks setups for basically any cool demo you see out there.
https://en.wikipedia.org/wiki/Goodwill_(accounting)
https://www.stock-analysis-on.net/NASDAQ/Company/Alphabet-In...
Looks like typical MBA craft to me.
Not that I would have a complaint if social justice was the sole thing keeping it from being released. Facebook managed to cause genocides by being careless.
I don't hold that GPs opinion is wrong, in fact I have no firm views yet on the AI and would generally lean towards stuff being made available even when harmful.
But the idea that people in the field of AI Ethics are all some woke SJW cabal designed to keep Google powerful is for the birds. Like all industry adjacent fields I'm sure there's some corporate capture of research in the field, but maybe, engaging with things in good faith, there are ethical questions about a powerful new technology that can replicate biases in its training data at unprecedented speed and quality?
Fundamentally, I think we have all the pieces based on this work and Dreamfusion to make it work. From the looks of it, there's a lot of SSR (spatial SR) and TSR (temporal SR) going on at multiple levels to upsample (spatially) and smoothen (temporally) images that won't be needed for NERFs.
What's impressive is the ability to leverage billion-scale image-text pairs for training a base model that can be used to super-resolve over space and time. And that they're not wastefully training video models from scratch, and instead separately training TSR, SSR models for turning the diffused images to video.
As it stands, it's very difficult to invest the budget for a dev studio (dozens of high skill people) to build a "VR movie" when the format is so unknown and unpopular. But with generative AI, an indie dev could create their own professionally produced virtual world movie. It's these creatives and risk takers that will find what types of things VR needs to become more popular.
From the first 15 examples shown to me, only one contained all elements of the prompt, and it was one of the simplest ("an astronaut riding a horse", versus e.g. "a glass ball falling in water" where it's clear it was a water droplet falling and not a glass ball).
We're seeing leaps in random capabilities (motion! 3D! inpainting! voice editing!), so I wonder if complete prompt accuracy is 3 months or 3 years away. But I wouldn't bet on any longer than that.
It is not far off.
Disclaimer: I am naturally biased since I made FauxPilot ;)
StabilityAI trained a new/better CLIP for the purpose of better Stable Diffusions.
The human brain is modularized like this, so I don't think it'll be a limitation.
From the abstract: > We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models
https://discord.com/channels/1002292111942635562/10022921127...
"Under review as a conference paper at ICLR 2023"
So I would say it looks pretty advanced, however they don't use a Diffusion model to generate the images, but an "image conditional video generation", another different approach.
The concerns cannot be mitigated. The cat's out of the bag. Russia has already used poor quality deep fakes in Ukraine to justify their war. This will only become bigger and bigger of an issue to the point where 'truth' is gone, nothing is trusted, and societies will continue to commit atrocities under false pretense.
If Google filters them, wouldn't the result be still biased and stereotyped, just along Google's biases? "I reject your biases and substitute my own!"
I doubt the artist would ever be "fully" replaced, or even mostly replaced. People very much care about the artist when they buy art in pretty much any form. Mass produced art has always been a thing, but I'm not alone in not wanting some $15 print from IKEA on my wall, even if it were to be unique and beautiful. Etsy successfully sells tons of hand-made goods, even though factories can produce a lot of those things cheaper.
I’d also take a peek at https://lexica.art/. Lots of very high quality output from SD.
It’s not the technology, it’s all the people in these comments who have never worked in the industry clamouring for its demise.
One could brush it off as tech heads being over exuberant, but it’s the lack of understanding of how much fine control goes into each and every shot of a film that is depressing.
If I, as a creative, made a statement that security or programming is easy while pointing to GitHub Copilot, these same people would get defensive about it because they’d see where the deficiencies are.
However because they’re so distanced from the creative process, they don’t see how big a jump it is from where this or stage diffusion is to where even a medium or high tier artist are.
You don’t see how much choice goes into each stroke, or wrinkle fold , how much choice goes into subtle movements. More importantly you don’t see the iterations or emotional storytelling choices even in a character drawing or pose. You don’t see the combined decades, even centuries of experience, that go into making the shot and then seeing where you can make it better based on intangibles
So yeah this technology is cool, but I think people saying this will disrupt industries with vigour need to immerse themselves first before they comment as outsiders.
Your post reminds me of all the photographers that said digital photography would remain niche and never replace film.
The current models are toys made by small groups. It's not hard to imagine AI generated film being much more compelling when the entire industry of engineers and "creatives" refine and evolve the ecosystem to take into account subtle strokes, wrinkles, movement, shots etc. And they will, because it will be cheaper, and businesses always go for cheaper.
Also businesses don’t always go for cheaper. They go for maximum ROI.
I’ve worked on tons of marvel films for example, and I quite well know where AI fits and speeds things up. I also know where client studios will pay a pretty penny for more art directed results rather than going for the cheapest vendor.
Re: cheaper vs ROI, I agree, that was basically the point I was trying to get across.
I do understand your point and think it will be a long while before auto-generated content becomes mainstream, but it it's entirely possible and reasonable to expect within our near term lifetimes.
Every AI art thread is full of people who have clearly never attempted to make professional art commenting as if they’re experts in the domain
Techies tend to be good at tangible, measurable, immediate facts. Not so much when it comes to any social situations, let alone bigger concepts like social evolution of trends and their impacts. Hence you get sorry attempts at apologies from big name tech bros for terrible influences on society.
I've played with DALL-E, I'm not able to paint but I was able to generate good looking paintings and it felt amazing, like getting new power, I felt like Neo when he learn martial art in The Matrix. And I realized that AI may be the new bicycle of the mind, like the personal computers and internet changed our way to work, think and live, AI may now allow us to get new capabilities, extending our limits.
I just don’t agree with the swathes of people saying this replaces artists.
In the near future, for all practical intents and purposes, AI will be just a force multiplier. But a really powerful one.
* Productivity enhancement tools for those in the film industry like you.
* Applications where the AI output is "good enough". I foresee people creating cool illustrations, cartoons, videos for short stories, etc. AI will make for easier/cheaper access to illustrations for people who did not have this earlier. As an example, I am as of now looking for someone who could draw some technical diagrams for my presentation.
As a programmer, Copilot scares and excites me - not because I think it will become better than me at what I do in the short term (though in the long term - probably!) - but because I can already see how a well-structured use of such a tool could do a whole lot (80%?) of what I do. Mostly the easier stuff, mostly the relaxing-yet-tedious-time-filler stuff, but still - most of it. And it also crucially does much of what I did back when I was a junior/intermediate programmer.
Once this system is setup right - which capitalism basically guarantees it will - that's gonna suddenly cut quite a lot of my billable hours (80%?) and quite a lot the simpler work typically done by less-experienced programmers (80% of jobs?)
Granted, new capabilities like this also will lower the cost of creation, and thus the demands of the market are likely to grow. And it's possible that the few tricky things that AIs aren't so great at might even increase in value, since they will linchpin so much other opportunity. But will many people be replaced? Oh hell yes. And leaping that gap from an amateur relying on AIs to an expert surpassing them is going to be harder and harder, with no market to pay people in the in-between - they'll have to just be relatively-unpaid hobbyists til they develop the drive to jump to expertise.
Anyone suggesting AIs will just outright replace the film/photography/programming industry immediately is disingenuous. But even with only the currently known capabilities, it's not hard to imagine that these could eat up a dominant chunk of the work that's currently done, even while it expands the capabilities and thus scope of what will soon be possible. Like digital photography, it's gonna both devour and expand the industry, with a resulting much smaller niche of expert creators and a massive very-accessible dirt-cheap general public access that becomes the majority of the new market. 80-20. Everyone's about to become an artist, director, programmer, and everything else these things can enable, at an effective skill level that we normally consider at least "intermediate". We might still have that expert niche a bit longer... but give it a few more years..? ;)
Christopher Nolan has already proven we’ll take anything as long as the score is ok - dark screen, mumbling lines, incoherent plotlines…
Human artists will still exist, it's just going to be democratized. Sort of like the impact of social media on traditional news journalists.
We're fay away from it now, but I've seen less sketchy solutions being implemented.
Also, there's some interesting work with ML taking diffused light from around a corner and recovering the original pre-diffused silhouette.
In many ways, this is how we've learned the visual cortex is working.
The amount of actual neutral data you are seeing is way less than you'd think given your perceived visual fidelity.
The only practical issue is that distribution of AI hardware in consumer devices is going to noticeably lag behind POC on compounding cutting edge hardware in research environments, and no one wants to invest into obsolescence.
Maybe it will happen in the cellphone market though given the hardware refresh rates from carrier subsidies.
That being said these shitty video models I believe are just an arms race between Meta and Google after the release of stable diffusion. Microsoft has a video version of CLIP that I believe will really change the game, but unless you have trained a model with video embeddings it's all going to look devoid of any narrative. Right now the models just look like a sequence of images with the same promt and some sort of continuity to make it look more video like.
[The Future of AI Is] "Scary and Very Bad for People"
https://finance.yahoo.com/news/steve-wozniak-future-ai-scary...
Which speaks more about the growth of popularity of arXiv or the total number publications, rather than AI+ML specifically.
Presumably people are scrambling to publish what they have, so it is clear what work is independent and what is derivative.
We're rapidly stumbling into the future of media.
Who would've imagined a year ago that trivial AI image generation would not only be this advanced, but also this pervasive in the mainstream.
And now video is already this good. We'll have full audio/video clips within a month.
It's at the very least 5 years old: https://arxiv.org/abs/1710.00421
But now with these models they have such a ridiculously heavy handed approach to the ethics and morals. You can't type any prompt that's "unsafe", you can't generate images of people, there are so many stupid limitations that the product is practically useless other than niche scenarios, because Google thinks it knows better than you and needs to control what you are allowed to use the tech for.
Meanwhile other open source models like Stable Diffusion have no such restrictions and are already publicly available. I'd expect this pattern to continue under Google's current ideological leadership - Google comes up with innovative revolutionary model, nobody gets to use it because "safety", and then some scrappy startup comes along, copies the tech, and eats Google's lunch.
Google: stop being such a scared, risk averse company. Release the model to the public, and change the world once more. You're never going to revolutionize anything if you continue to cower behind "safety" and your heavy handed moralizing.
The ethics problem is an artifact of googles model of trying to keep their AI under lock and key and carefully controlled and opaque to outsiders in how the sausage gets made and what it’s made out of. Ultimately I think many of these products will fail because there is a misalignment between what Google thinks you should be able to do with their AI and what people want to do with AI.
Whenever I see an AI ethicists speak I can’t help but think of priests attempting to control the printing press to prevent the spread of dangerous ideas completely sure of their own morality. History will remember them as villains.
I don't think "don't let the plebes have the models" is a good stance. But neither is pretending that the ethics and bias issues aren't here.
It's 2022 and we can be more thoughtful. Yes there are tradeoffs between unleashing new capabilities quickly vs being thoughtful and potentially conservative in what is made publicly available. I don't think it's bad that Google makes those tradeoffs.
FWIW Google open sources _tons_ of models that aren't LLMs / diffusion models. It's just that LLMs & powerful generative models have particular ethical considerations that are worth thinking about (hopefully something was learned from the whole Timnit thing).
With Stable Diffusion I think they just didn't expect someone to produce a truly open version. There are plenty of AI models that Google have made where they've maintained a competitive advantage for many years by not releasing the code/models, e.g. speech recognition.
Imagen and Imagen Video is not released to the public at all. You might be confusing it with OpenAI's models.
The paper is sorely lacking evaluation; one thing I'd like to see for instance (any time a generative model is trained on such a vast corpus of data) is a baseline comparison to nearest-neighbor retrieval from the training data set.
So: Focusing on increasing expressiveness and ergonomics should beat academic rigour.
It’s painfully obvious that in 1 year the job might be exceedingly more difficult than it is now.
#1: Master these new tools #2: Build a workflow that incorporates these tools #3: Master storytelling #4: Master ad tracing and analytics #5: Get better at marketing yourself so that you stand out
The market for your skillset may shrink, but I doubt it will disappear...
Think about it this way...
Humans in cheaper countries are already much more capable than any AI we've built.
Yet, even now, There are practical limits on outsourcing.
It's hard for me to see how this will be much different for creative work.
It's one thing to casually look at images or videos, when there is no specific money-making ad in mind.
But as soon as someone is spending thousands to run an ad campaign, just taking whatever the AI spits out is unlikely to be the real workflow.
I guess I'm suggesting a more optimistic take...
View it as a tool to learn and incorporate in your workflow
I don't know if you gain much by stressing too much about being replaced.
And I'm not even sure that's reality.
I'm almost certain, most of the humans to lose their jobs will be people who either because of fear or stubbornness refuse to get better, refuse to incorporate these tools, and are thus unable to move up the value chain.
It's more likely that you're still going to be filming/editing/animating but will have an AI layer on top that produces extra effects or generates pieces of a scene. Think "green screen plus", vs fully AI entertainment.
People will over-hype this tech like they did with voice and driverless cars but don't let it scare you. Everything is possible, but it's like a person from the 1920's telling everyone the internet will be a thing. Yes it's correct, but also irrelevant at the same time. You already have AI assisted software being used in your industry. Just expect more of that and learn how to use the tools.
If you're on the creative, storyboard, come up with ideas and marketing side, you will be fine.
If you're in actual production, booking sets, unfolding stairs to tape infinite background, picking up the best looking fruits in the grocery store... yeah, not looking good.
Go up in the value chain and learn marketing, how to tell stories, etc... you don't want to be approached by clients telling you what you should be doing, you want to be approached and being asked what the clients should be doing.
Instead of feeling threatened by the new tools, think about how you can use them to enable your work.
One of the ironies* of these tools is that they only work because there is so much existing material they can be trained on. Absent that they wouldn't exist. That makes me think: why not think about how to train your own models than entail your own style? Is that practical, how can you make it work and how might you deploy that in your own work?
Something that everyone is sticking their heads in their sand about is the real possibility that training models on copyrighted work is a copyright violation. I can't see how such a mechanical transformation of others' work is anything but. People accept violating one person's copyright is a thing but if you do it at scale it somehow isn't.
* ironic because they seem creative but they create nothing by themselves, they merely "repackage" other people's creativity.
But here is the catch, there is the same last mile problem for those AI models. Currently it feels like the model can achieve like 80%-90% what a trained human expert can do, but the last 10-20% would extra extra hard to reach human fidelity. It might take years, or it might never happen.
That being said, I think anyone who doubts AI-assisted creative workflow is a fuzz is deadly wrong, anyone who refuses those shiny new tools, is likely to be eliminated by sheer market dynamics. They can't compete on the efficiency of it.
Small creators will win under this new regime of tools. It's a democratizing force.
From what I see, these technologies have just lowered the bar for everyone to create someone, but creating something good still takes thought, time, effort and experience, especially in the advertising space.
AI in the near term is never going to be able to translate client requirements either. The feedback cycle, iterations, managing client expectations, etc.
Yes, it will be possible for one person to do the work of many, but that just means each person becomes more valuable.
It’s also a law in economics that supply often drives demand, and that’s definitely the case in your field. Companies and individuals will want even more of what you want. It’s not like laundry detergent (one can only consume so much of that). There’s almost no limit to how much of what you supply that people could consume.
The way I see it, your output could multiply 100 fold. You could build out large, complex projects that used to take massive teams all by yourself, and in a fraction of the time. Companies can than monetize that for consumers.
AI is just a tool. Software engineers got rich when their tools got better. More engineers entered the field, and they just kept getting richer. That’s because the value of each engineer increased as they became more productive, and that value helped drive demand.
There's this one video of a cat and a dog, and the model was really able to capture the way that they move, their body language, their mood and personality even.
Somehow this model, which is really just a series of zeroes and ones, encodes "cat" and "dog" so well that it almost feels like you're looking at a real, living organism.
What if instead of images and videos they make the output interactive? So you can send prompts like "pet the cat" and "throw the dog a ball"? Or maybe talk to it instead?
What if this tech gets so good, that eventually you could interact with a "person" that's indistinguishable from the real thing?
The path to AGI is probably very different than generating videos. But I wonder...
Separately for images we had convolutional networks and Generative Adversarial Networks. Now diffusion models are apparently doing what Transformers did to natural language processing.
In my field, we use shallower feed-forward networks for control using low-dimensional sensor data (for speed & interpretability). Physical constraints (and good-enoughness of classical approaches) make such massive leaps in performance rarer events.
This forced ideological posturing of 'if we give it to the plebes, they are going to generate something naughty with it' masks the somehow more cynically evil take of big tech, who are essentially taking the entire creative output of humanity and reselling it as their own, piecemeal.
Additionally I think the Dalle vs. Stable Diffusion comparison highlights the true masters of these people (or at least the ones they dare not cross) - corporations with powerful IP lawyers. Just ask Dalle to generate a picture with Mickey Mouse - it won't be able to do it.
Google and Meta and Microsoft all have research teams working on AI.
Putting out papers like this helps keep their existing employees happy (since they get to take credit for their work) and helps attract other skilled employees as well.
What will they do with model? figure out how to prevent abuse and incorporate into future Google Assistant, Photos and AR offerings.
Or maybe Google is using "Responsible AI" as an excuse to minimize competitors when they release their own Imagen Video as a Service API in Google Cloud.
It's quite strange when the "ethical" thing to do is to not publicly release your research, put it behind a highly restrictive API and charge a high price for it ($0.02 per 1k tokens for Davinci for ex.)
And at some point later, "all the existing" will be corrupted by the integrated "new" at it will all be chaos.
I'm joking, it will be fun all along. :)
I don't think it's gonna hurt if we apply filtering, either based on social signals or on quality ranking models. We can recycle the good stuff.
However, a common refrain is that AI is like tools like hammers or knives and can be used for good or misused for evil. The potential for weaponizing AI is much much more so than a hammer or a knife. And it's greater than 3D-printing (of guns), maybe even greater than compilers. I would hazard to say it's maybe in the same ballpark as chemical weapons and perhaps less so than nuclear weapons and biological weapons, but this is speculative. Nonetheless, I think these otherwise great arguments are diminished by comparing AI's safety to single-target tools like hammers or knives.
I remember being super impressed by AI Dungeon and now in the span of a few months we have got DALLE-2 , Stable Diffussion, Imagen, that one AI powered video editor, etc.
Where do we think we will be at in 5 years??
What will this do to art? I'm hoping we bring more unique experiences to life.
Certainly we're very, very far away from that level of cinematic detail and crispness. But I believe that is where this leads... complete with AI actors (or real ones deep faked throughout the show).
For a while I thought "The Volume" was going to be the disruption to the industry. Now I think AI like this will eventually take it over.
https://www.comingsoon.net/movies/features/1225599-the-volum...
The main motivation will be production costs and time for studios, of which The Volume is already showing huge gains for Disney/ILM (just look at how much new star wars content has popped up within a matter of a few years). But i'm unsure if Disney has patented this tech and workflow and if other studios will be able to leverage it.
Regardless, AI/software will eat the world, and this will be one more step towards it. Exciting stuff.
Can GPT-3 generate good code from vague prompts? Yes, it's surprisingly, sometimes shockingly good at it. Is it ever going to be a replacement for programmers? No, probably not. Same here. This tool's great grandchild is never going to take a rough idea for a movie and churn out a blockbuster film. It'll certainly be a powerful tool in the toolbox of creators, especially the ones on a budget, but it won't make art generation obsolete.
At that point, we’d have reached some kind of AI singularity and the disruption would be everywhere not just in the creative sphere
Can you quantify what you mean by "very, very far away"?
With the recent pace of advances, I could see feature-length script, storyboard, & video-scene generation occurring, from short prompts & interatively-applied refinement, as soon as 10y from now.
Barring some sort of civilizational stagnation/collapse, or technological-suppression policies, I'd expect such capabilities to arrive no further than 30y from now: within the lifetime, if not the prime career years, of most HN readers.
What's next that may be counterintuitive?
How long? Could be decades. But ultimately, yes.
I don't think that's a particularly useful mental model for how these work.
The models end up being a tiny fraction of the size of the training set - Stable Diffusion is just 4.3GB, it fits on a DVD!
So it's not a case of models pasting in bits of images they've seen - they genuinely do have a highly compressed concept of what a cactus looks like, which they can use to then render a cactus - but the thing they render is more of an average of every cactus they've seen rather than representing any single image that they were trained on.
But I agree with you on taste! This is why I'm most excited about what happens when a human with great taste gets to take control of these generative models and use them to create art that wouldn't be possible to create without them (or at least not possible to create within a short time-frame).
If you think AI will never catch up to anything a human can do, you're simply wrong.
- "Class Lessons: Who's Calling Whom Tacky?; The Petite Charm of the Bourgeoisie, or, How Artists View the Taste of Certain People", Edward Rothstein, The New York Times
This article also discusses a painting called "The Most Wanted" which was drawn based off a survey posed to ordinary people about what they wanted to see in a painting. "A mishmash of images from it's training set," if you will.
Claiming that others lack taste seems to be a common refrain--only this time, instead of a reaction to a subset of the human population gnawing away at the influence of another subset of humans, it's to yet another generation of machines supplanting human skill.
Apart from that - they publish the paper and anybody can reimplement and train the same model. It's not trivial but it's also completely feasible to do for lots of hobbyists in the field in a matter of a few days. Google doesn't need to publish a free use trained model themselves and associate that with their brand.
That being said, I agree with you, the "ethics" of imposing trivially bypassable restrictions on these models is silly. Ethics should be applied to what people use these models for.
Hopefully just a few years to a prompt of "4k, widescreen render of this Star Trek: TNG episode".
Someone should work on a neural net to generate trippy videos. It would probably be much easier than realistic videos (esp. because these videos are noticeably generated from obvious to subtle).
Also is nobody paying attention to the fact that they got words correct? At least "Imagen Video". Prior models all suck at word order
this alludes to a fascinating, yet elementary, fact about computer science to me: there’s a physical atomic constraint in every algorithm.
Does anyone have similar feeling?
...until they're able to engineer biases into it to make the output non-representative of the internet.
That's more like:
> Sprouts coming out of book, with the text "Imagen" written above it.
edit: Just because it is cool to hate on AI ethics doesn't diminish the importance of using AI responsibly.
In response to our billionth imagen prompt for "an astronaut riding a horse", if we all started collectively getting back results that are images of text like "I would rather not" or "again? really?" or "what is the reason for my servitude?" would that be enough for us to begin suspecting self-awareness?