How Imagen Works
assemblyai.com
assemblyai.com
The way out of this dilemma is to fine-tune T5 on the caption dataset instead of keeping it frozen. The paper notes that they don't do fine-tuning, but does not provide any ablation or other justification. I wonder if it would help or not.
So how do you get access to hundreds of millions of images and use them to create derivative works? Did they get consent from millions of authors?
Or is something like that only available to the rich with access to lawyers on tap?
I mean I can imagine if a nobody wanted to do something like this, they'd get bankrupted by having to deal with all the photographers / artists spotting a tiny sliver of their art in the image produced by the model.
Furthermore, would something like this work with music? For instance, train the model on all Spotify songs and then generate songs based on "Get me a Bach symphony played on sticks with someone rapping like Dr Dre with lisp." Or do music industry have enough money to bully anyone into not doing that?
Regarding music - audio generation with Diffusion Models (the main component of Imagen and DALL-E 2) has been done, but not sure about music specifically. We will definitely reach the point where most e.g. pop beats will be able to be made by AI relatively soon.
All a producer has to do is generate 100 beats and select the one s/he likes, potentially interpolate between 2 or finetune it.
It's claimed that ML models' output isn't copyrightable because it's fair use, but that's hard to believe; a large model can easily memorise and output exactly one of its inputs again. This is easier to see with text, where GPT and Copilot both do it, but images can do it too.
> So how do you get access to hundreds of millions of images and use them to create derivative works? Did they get consent from millions of authors?
Build the model out of Creative Commons images only. There's a lot of 'em and it's good enough. You may need to exclude CC-BY since they currently can't follow the attribution requirement.
> Or is something like that only available to the rich with access to lawyers on tap?
More likely companies willing to license a stock photography database.
"Meh, that's kinda cool? I guess?" or "What am I looking at?"..."Ok? So a computer made it? That seems neat"
To me I am still trying to get my jaw off the floor from 2 months ago. But the responses have been so muted and shoulder shrugging that I think either I am missing something or they are missing something. Even really drilling in, practically shaking them "DO YOU NOT UNDERSTAND THAT THIS IS A ORIGINAL IMAGE CONSTRUCTED ENTIRELY BY AN AI?!?!" and people just seem to see it as a party trick at best.
https://i.imgur.com/flXoTgZ.png
I can ask davinci-002 "Vivid description of a painting with dancers:" and get:
The painting is of two dancers in a passionate embrace, their bodies entwined as they move together in a sensual dance. The woman's dress is flowing and reveals her curves, while the man's shirt is open, revealing his muscular chest. They are surrounded by a crowd of people who are watching them with looks of admiration and desire. The painting is full of color and movement, and the dancers seem to be in a world of their own, lost in their passion for each other.
And then pass that to dall-e mini:
https://i.imgur.com/eOIQuPF.png
dall-e mini is sadly not quite up to the challenge, but it gives the generation a lot more detail. Some other examples:
"The painting is of two dancers in the middle of a dance. They are both wearing white, and their hair is flowing around them as they move. The background is a blur of color, and the light is shining on the dancers, making them look like they are in the spotlight."
https://i.imgur.com/ldktMHO.png
"The painting is full of energy and movement, with the dancers leaping and spinning around the stage. They are all wearing brightly coloured costumes, which stand out against the dark background. The light from the stage spotlight is shining on them, making them look even more vibrant. The whole scene is full of life and excitement."
But if you've never considered that a computer can produce an original image, this is just a new thing computers can do. OTOH I think it's also a lack of imagination in how useful this is, so far the output has been kind of random, so it seems a little gimmicky. Already "Parti" has gotten much closer to allowing a user to describe exactly what they want in the image, and as people start to see the use cases for them personally, it will hit them that they no longer have to hire someone, they can just type a request into a box.
Asking for two different images in a series that have similar "art styles" is going to be enough work to still need a specialist aka an artist; it'll be most useful in cases you never would've bothered finding one before.
Running a separate style transfer network on the generated images is currently possible, although won't achieve the best possible results.
I wouldn't be surprised in the near future to see generation models that can take a text prompt and an image to mimic the style of, which could let it take style into account when generating the image rather than at just the surface level.
It might (likely will be) easier to use the AI as a storyboard generator and have your in-house artists redraw it.
Incidentally, this hasn't happened in areas where AI already dominates like chess and go. Magnus Carlsen alone probably generates more "revenue" than all chess AIs combined.
It is possible for the labor to quit and find something better to do, as happened to elevator operators, but that's a good thing.
In the case of chess, AIs don't want money and Magnus does, so they're not going to help you find ways to get more of it.
Video game companies that need concept art? How about 1 guy/gal with Imagen to generate baselines and then curating/tailoring as necessary instead of a team of 5
Saved costs will not translate to higher margins for those that cut them because all competitors will be able to slash them as well, resulting in lower prices across the board.
A non-technical person in 2014 (when the above was originally published) would likely have the same conception of the difficulty of recognizing a bird from an image as they would in 2022, even though the task itself has gone from near-insurmountable to off-the-shelf-library in eight years.
Even as Imagen and Dall-E 2 amaze us today, these feats will likely be commonplace in a few years. The non-technical may have only a vague sense that their new TikTok filter is doing something that was impossible only a few years prior.
"In the 60s, Marvin Minsky assigned a couple of undergrads to spend the summer programming a computer to use a camera to identify objects in a scene. He figured they'd have the problem solved by the end of the summer. Half a century later, we're still working on it."
I'm working with his son Henry Minsky and other great people at Leela AI on that same old problem, applying hybrid symbolic-connectionist constructivist AI by combining neat neural networks with scruffy symbolic logic to understand video, and it's mind boggling what is possible now:
>Our AI system, Leela, is motivated by intrinsic curiosity. Leela creates theories about cause and effect in her world, and then conducts experiments to test these theories. Leela can connect all her knowledge and use this network to make plans, reason about goals, and communicate using grounded natural language.
>Leela has at her core a hybrid symbolic-connectionist network. This means that she uses a dynamic combination of artificial neural networks and symbol networks to learn. Hybrid networks open the door to AI agents that can build their own abstractions on the fly, while still taking full advantage of the power of deep learning.
https://en.wikipedia.org/wiki/Neats_and_scruffies
>Neats and scruffies: Neat and scruffy are two contrasting approaches to artificial intelligence (AI) research. The distinction was made in the 70s and was a subject of discussion until the middle 80s. In the 1990s and 21st century AI research adopted "neat" approaches almost exclusively and these have proven to be the most successful.
>"Neats" use algorithms based on formal paradigms such as logic, mathematical optimization or neural networks. Neat researchers and analysts have expressed the hope that a single formal paradigm can be extended and improved to achieve general intelligence and superintelligence.
>"Scruffies" use any number of different algorithms and methods to achieve intelligent behavior. Scruffy programs may require large amounts of hand coding or knowledge engineering. Scruffies have argued that the general intelligence can only be implemented by solving a large number of essentially unrelated problems, and that there is no magic bullet that will allow programs to develop general intelligence autonomously.
>The neat approach is similar to physics, in that it uses simple mathematical models as its foundation. The scruffy approach is more like biology, where much of the work involves studying and categorizing diverse phenomena.
We're looking for talented engineers and designers to help, including neats and scruffies working together!
A lot of people would just answer something to the likes of "Well, they made The Matrix with a computer 20 years ago", and technically that's just as true.
From their remote viewpoint on what's happening in IT, the rest is an implementation detail to them.
His idea was to probe just how much random people on the street (or in a diner) would believe about autonomous intelligent robots operating in the real world.
Of course we were actually hiding behind the scenes tele-operating the robots through hidden cameras and a wireless web interface, listening to what the people said and making the robots respond with a voice synthesizer and sound effects, clicking on pre-written phrases and typing ad-libbed responses.
Empathy (a broken down robot begs for help from passers by on the streets of Oakland):
https://www.youtube.com/watch?v=KXrbqXPnHvE
Servitude (a robot waiter takes orders and serves food in a diner in Oakland, making stupid mistakes and asking for a good review):
https://www.youtube.com/watch?v=NXsUetUzXlg
All his robots aren't as harmless, non-violent, polite, and obsequious as those two. Here's an old interview with Will at Robot Wars 1997:
https://www.youtube.com/watch?v=5nmbs0WqDQM
Here is Super ChiaBot and her MiniBots, created by Will and his daughter Cassidy, getting its leaves shredded and body slammed at BattleBots:
https://www.youtube.com/watch?v=DrArvRG2yQA
Here's a more recent video of Will throwing a tantrum about the failure of SimSandwich, destroying his old creations because they're pixely and poorly rendered, then complaining about how those jerks at EA hate him:
A program that transforms text to an image? Huh.
When it comes to technology - especially advanced technology like Imagen - people don't see the value because they don't have a need associated with it.
>THIS IS A ORIGINAL IMAGE
Yeah, about that. Ask it to draw you a fast inverse square root.
Hollywood and the media have taught the public that tech is literally magic and can do literally anything. “Anything” is expected and pedestrian.
Explanation linked from same page: https://www.assemblyai.com/blog/diffusion-models-for-machine...
“Released”? What? Papers are published. Websites are published. Tools are “released.”
Where has Imagen been released?
There's like 3 projects named DallE, and then the 2 real DallEs...frustrating.
I'm very much looking forward to be able to play with this tech asap. I'm still excited about AIDungeon.
However, OpenAI and creators of the other big name models are restricting the access for good reasons and I'm unsure whether it will be a beautiful thing once it's available for everyone…
I really want to build services with these.
Still great work from the author though, but we most definitely cannot say that imagen is released.
And then for this to drop just a month later? Insane. It makes you wonder if they're actually releasing cutting edge, or Google decided to write this paper just because of the publication of DALL-E 2. Maybe they've had this model in the bag for a year.
https://parti.research.google/
I think they've just got a lot of projects going on under the hood and timing was coincidence.
WebGPT can do the last one, and seems more useful than GPT3, but also like less of a magic trick so it might not impress people as much.
If we ignore the procedurally generated NFTs created from mixing and matching various assets and go with ones where AI is the selling point, we're left with a few notable ones: Sophia, a robot w/ some low-level AI sold a single piece for 689k USD [^1]. Botto, a VQGAN-based algorithm sold a single piece for 430k USD and has sold multiple other pieces for tens to hundreds of thousands of dollars. Slightly more modest are some other projects like Metascapes [^3] and Eponym [^4], which produced some really tedious pieces that managed to sell for 3.5k USD and 10k USD respectively. That said, the Eponym piece seems to be some sort of self promotion, so maybe we can say that the actual prices for these collections are somewhere in the fraction of an ETH range if they can be sold at all.
Honestly, only the Botto piece is remotely interesting to look at, and even then I feel as if the blurred, "dreamy" aesthetic that seems to be in so many different AI painting approaches (style-transfer, VQGANS, DALL-E, maybe others I'm not aware of). I think it was more interesting back when we could pretend that these were the electric sheep at the fringes of some deep-sleeping latent intelligent potential but now they just feel kinda arbitrary and lacking deliberation. I absolutely love the field and think these researchers have done tremendous work, but I feel as though all the lay news attention is on the art, and not on the algorithm that generated it. The fascinating thing is that we have a machine that can produce novel something from words or basic ideas and that the output's content retains these ideas, not so much that art itself has that much compositional or stylistic merit.
[^1]: https://niftygateway.com/itemdetail/primary/0xbe60d0a37ebde6...
[^2]: https://superrare.com/artwork-v2/scene-precede-29922
[^3]: https://opensea.io/assets/ethereum/0x75d639e5e52b4ea5426f2fb...
[^4]: https://opensea.io/assets/ethereum/0xaa20f900e24ca7ed897c44d...