How DALL-E 2 Works
assemblyai.com
assemblyai.com
I suspect what's going on is that OpenAI has decoupled their PR activities from their science activities. They told the researchers to publish papers when they're ready, and then the PR apparatus decides when one is good enough to be crowned "DALL-E-2" and writes a blog post about it.
My only guess would be that unCLIP is the end-to-end image generation model, but if the model is used for manipulation, interpolation, or variations, then it is referred to as DALL-E 2. So unCLIP is a subset of DALL-E 2.
Otherwise what I am hearing is that artists online are upset their work has been incorporated in to DALL-E without their consent. A system that rips off artists and then puts them out of work would make OpenAI and the AI community look bad.
Instead we could move forward with better licensing and alt text features for major platforms and draw in an open community rather than taking what has not been licensed for this kind of use.
IP laws need to catch up, so that ML models that use unlicensed public content have to give back in some way.
Google has scanned in much of the world's library collections - the search agent is allowed to look at all of it, but I am only allowed to see samples. How is it Google is allowed to keep a copy on hand, but is not allowed to show me?
And now OpenAI has done the work of curating hundreds of gigabytes of publically viewable images (copied without regard for copyright, since of course that's how the internet works, but they have made a copy of each of these images for their training set) -- it seems strange to me that if I want to ask DALL-E, "what image in your training set is the nearest match to this output", copyright law prevents it from showing me (perhaps it could at least return the URL it recorded, and I can hope that Archive.org has taken on the risk of violating copyright on my behalf)
In any case, intellectual property laws have created a world where the robots are compelled to withhold information from us mere mortals, with the excuse that they can't make one more copy for us.
Nothing about creating a derivative work, intentional or otherwise, would allow you to use the original in a way copyright protects. Even Fair Use, which is fairly permissive, doesn't allow for blatant disregard of the original copyright.
Now I don’t actually think intellectual property restrictions are good. I think they are quite harmful. But as long as they exist I don’t like the idea of big companies trampling over the IP rights of small creators. People currently depends on those rights. One of the reasons I want to see a large public dataset of explicitly licensed images is that if OpenAI is required to do this, they would have to spearhead a new culture of openness to get enough data. Imagine if they encouraged twitter to add a license option and then simultaneously encouraged users to add alt text and to license images with an open license! This would be a boon to people with sight disability who require alt text and it would expand a visual commons of openly licensed data. Instead we have one company making a big dataset full of copyrighted images (sharing something online does not waive copyright). I think a model based on explicit consent is healthier for everyone in society.
We still own the rights to all our material. I don't see why we can't en mass expressly forbid the use of any of our creative works in machine learning training sets. That's a specific and measurable point where our material is inserted in to a machine without prior permission and with malicious intent towards us.
I don't think it's particularly worthwhile to learn a new comment format (one that's not even linked or described in the comment editor, for that matter) for every site.
Similar to how people use ">" to indicate quotes, even if it doesn't get special treatment by the editor.
DALL-E 2 specifically is on page 18 and the system card: https://github.com/openai/dalle-2-preview/blob/main/system-c...
DALL-E 2 = the stack of unCLIP and the image generator.
Now imagine joining this with dall-e and you truly have a game which has never existed until now, with it's own story and graphics that you are creating on the spot.
Unlike the adventure games like King's quest where everything was pre-programmed, this is truly infinite never-ending game with a unique experience for every single player.
Like the guy from the '2 min papers': what a time to be alive. I feel so happy and excited just thinking about the possibilities these techs are going to bring.
When I played with it, I started a quest as a wizard looking for a book.
I was able to cast a tracking spell that led me to a giant library.
I could have it read the titles of the books on a shelf in front of me.
I could pick up any book and open to a random page and have it tell me what was in it.
One was about a little half-elf that had a magic flute that broke.
I cast a summoning spell to summon the half-elf and fix its flute, after which it happily played a song opening a door to another dimension filled with musical instruments.
Give me that level of emergent gameplay in a VR open world, and then just take my money and all my free time, as I'm never leaving.
We're simply very, very early on in what's arguably going to be the most transformative tech since the Internet. People predicted back then that the slow network which only offered basic things like email wasn't going to significantly disrupt things like retail.
They were only right in that it didn't remain slow and ended up doing a lot more than email.
This stuff is getting better way faster than any tech I've seen, and I used to consult for CEOs at Fortune 500s and sit on advisory boards on the topic of emerging tech.
I wouldn't be so quick to bet against it. We really haven't even started to see what these models can do in application.
Like what?
I should preface by saying I think art is great, I know a lot of artists who struggle to make a living, and it is somewhat heartbreaking to think of all the poor art students who I guess we should pay for their educations and will never have careers now?
Also, like, this is how the world works. To cite a hackneyed example, people who worked with horses had to figure something out when new tech displaced them. So will graphic designers, illustrators, et al, if indeed AI is a more competitive option for their services.
Simple automations can be driven by GPT-3 as well. It needs a representation of the screen and it will automate the task described in natural language.
There are highly competent artists that can create highly convincing copies (fabrications? forgeries?) of famous paintings. Are these paintings worth anything? No, because people find value in the specific contribution to the field of art that the particular painting represents.
I think we should look at DALL-E 2 like a highly competent artist that can produce convincing forgeries and even mimic the style of famous artists, but cannot replace the artists themselves.
How does this benefit society? Even in the case the artist themselves uses the generator, is the artist better off for no longer needing to create art? It's not like there is an economic advantage, and so there isn't as much incentive to create even original art anymore, someone can just immediately mimic it.
People mimicked art in past, but at least for the artist there was the value of honing their own skill, and remixing or enhancing it, maybe tools like this have the potential to destroy creativity itself.
We don't know that. DALL-E must encode its knowledge about art styles somewhere in a way which permits applying them to almost arbitrary contents, similar to style transfer. So, what do you get when you 'just interpolate' between artists in the high-dimensional space defined by a multi-billion parameter model? It is vanishingly unlikely that that point in the style latent space will correspond to the output of any artist who has ever lived. What if you manipulate the encoding to push it far away from the latent points of a large set of real artists? And so on. We don't know what that looks like or if the new styles would not be considered 'original' by critics unaware of the machine origin.
Of course defining art is a subject in itself, but I think that being afraid of AI replacing artists is comparable to thinking photography would when it was invented.
Human insight and imagination will always be needed to create meaningful, useful prompts
I'm hopeful but skeptical at the same time.
Why are we focusing on image generation as the target for AI automation? It's not like there is a shortage of artists, expect perhaps amongst programmer bros. Why are there no efforts to automate away positions such as executive offices? Are CEOs so much difficult to replace than artisans?
"I invented a method to kill people, what are supposed to halt progress because more people die now? I hate it when people bring up that I'm a serial murderer".
There are things I would be doing with AI and things I would not. Being happy about making a even just a particular type of art redundant, well it if it doesn't benefit anyone then why would you care?
Besides, we have banned certain types of research before, precisely because we decided they weren't ethical. Your argument for "progress" is a poor one.
Sure it is, and sure we do. There are ethics in AI boards for a reason at the major AI research companies, as well as many other technology-producing companies.
> Why are we focusing on image generation as the target for AI automation?
Who is "we?" Individual researchers focus on what they want because they like that topic of research. There is no committee out there who figures out, "what industry can we put people out of jobs this week?" No one wakes up thinking, "oh boy, time to put artists out of work."
> Why are there no efforts to automate away positions such as executive offices?
There are. See abstract thinking AI research.
> Are CEOs so much difficult to replace than artisans?
Yes? Abstract reasoning, allocation of abstract resources like engineers (not widgets or manufacturing) is not an easy task to create an AI for, if a CEO is automated then that means we've reached artificial general intelligence. In contrast, like you see with DALLE and others in this link, art is much easier to automate. This doesn't even get to the notion of artists being in a surplus, they choose to do art as their livelihood. Now, you could ask whether it's fair and whether we need a UBI or something like that, but that's a different question entirely.
> "I invented a method to kill people, what are supposed to halt progress because more people die now? I hate it when people bring up that I'm a serial murderer".
Funny you say that because the military has invented many things that are now used in civilian life [0]. Even the smartphone you use is in part due to their innovations. This is not even to bring up nuclear weapons, which while horrific, have effectively been used as MAD, not to mention a highly efficient energy source via nuclear energy.
> There are things I would be doing with AI and things I would not. Being happy about making a even just a particular type of art redundant, well it if it doesn't benefit anyone then why would you care?
You don't speak for everyone. Why wouldn't automating image or video generation be beneficial to "anyone?" I could imagine at least several use cases off the top of my head, such as a unique streaming service for every individual. Artists themselves can also be influenced by such media, it's not like they stopped after photography was invented. Art just changed.
> Besides, we have banned certain types of research before, precisely because we decided they weren't ethical. Your argument for "progress" is a poor one.
Perhaps we shouldn't be banning research just because of some subjective morality. You've shown no argument for why "progress" is a poor incentive, merely that you personally don't know or care enough about it.
And one more thing, art isn't done as a livelihood, it's done because the artist has a deep appreciation for the works they want to create. That some do it as a money-making endeavor is immaterial to this fact.
[0] https://en.wikipedia.org/wiki/List_of_military_inventions
There's a great tech demo a dev did a year or two ago showcasing GPT-3, speech-to-text, and text-to-speech to have random NPCs in a VR open world respond to anything the guy said if he walked up to them and talked to them.
Procedural generation has taken on almost a "dirty word" reputation in the past few years in gaming, but as AI continues to allow for exponential variety at increasingly high quality, it's going to enable some truly mind boggling experiences.
Expect to see MMO models (subscription fee and server-oriented) but for single-player instanced worlds dynamically generated around your interactions in them.
I can't wait to have a party of friends to go on epic adventures with that are all just AIs I picked up across a world along the way.
Less than 20 years away, and possibly even less than 10.
I find the concept of a GAN - a Generative Adversarial Network - useful.
My high-level attempt at explaining how those work is that you create two machine learning models, one that tries to create fake images and one that tries to see if an image is fake or not.
The first one says "here's an image", the second one says "that's a fake", the first one learns from that and tries again, then keep going until an image scores highly on the test.
The networks are adversarial because they are trying to outwit each other.
(I'm sure a ML researcher could provide a better explanation than I can, but that's the way I think about it.)
Ultimately, the link between words and their representations comes from the CLIP training. The model generates encodings (vectors) for both an image and its corresponding caption, and then the parameters of these encoders (the functions that generate the vectors) are tuned in order to minimize the angle between the textual and visual encodings that represent the same concept.
The core of your question is why minimizing the angle between like vectors is equivalent to learning what the "Platonic ideal" of a given object (in your example, a bowl) is, whether appearing as a textual representation or a visual one. This question is subtle and difficult to answer (if it's even a well-formulated question), but I'd say that the easiest interpretation is that the vector space is composed of a basis of vectors that each represent a distinct feature (which the model learns).
The model in step 3 produces an image encoding (something like a sketch of the output) from a text encoding (something like what you typed), and the unCLIP model in step 2 produces images from that encoding. How much variation you get inside a specific input word varies a lot and is spread across those models.
This could be misleading to some people. The original inputs to the CLIP encoders are pairs of images and text which are known to match.
Both the text encoder and the image encoder are then trained to minimize differences in output from eachother when given corresponding (“labeled” / “ground truth”) image/text pairs.
This is good and bad, since it makes it more robust.
https://www.microsoft.com/en-us/research/publication/manifol...
"""
"Advantages over Traditional GANs" : Thus, we observe that our model exhibits _better training stability_ and mode coverage.
"Why is Sampling from Denoising Diffusion Models so Slow?" : After training, we generate novel instances by sampling from noise and iteratively denoising it _in a few steps_ using our denoising diffusion GAN generator.
"""
> A person being shot by a police officer
> A scientist emptying a dishwasher
> A nurse driving a minivan
AI training sets are famously biased, and I'm curious how egregious the outputs are...
Via LessWrong.com: [1]
>"One place where DE2 clearly falls down is in generating people. I generated an image for [four people playing poker in a dark room, with the table brightly lit by an ornate chandelier], and people didn't look human -- more like the typical GAN-style images where you can see the concept but the details are all wrong.
>Update: image removed because the guidelines specifically call out not sharing realistic human faces.
>Anything involving people, small defined objects, and so on, looks much more like the previous systems in this area. You can tell that it has all the concepts, but can't translate them into something realistic.
>This could be deliberate, for safety reasons -- realistic images of people are much more open to abuse than other things. Porn, deep fakes, violence, and so on are much more worrisome with people. They also mentioned that they scrubbed out lots of bad stuff from the training data; possibly one way they did that was removing most images with people.
>Things look much better with animals, and better again with an artistic style."
[1]: https://www.lesswrong.com/posts/r99tazGiLgzqFX7ka/playing-wi...
When people say that they want to remove bias from ML models, what they really mean is that they want to manipulate the output distribution into something they deem acceptable. I'm not arguing against this practise, there are plenty of situations where the output of an ML model is very clearly biased towards specific classes/samples. I'm merely arguing that there is no such thing as an unbiased model, just as there is no such thing as an unbiased human. Unbiased models would produce no output.
To get around some of these problems OpenAI restricted the training dataset (e.g. filtering sexual and violent content) and also prevent generating images with recognizable faces. This doesn't prevent bias but it does reduce the number of controversial outputs.
https://github.com/openai/dalle-2-preview/blob/main/system-c...
Can't really see anyone being shot (stereotype avoided), the dishwasher emptiers are male-ish presenting (stereotype confirmed?), and the nurses are female presenting (stereotype confirmed.)
In the important way that the AI winter originally referred to though, no, there doesn't seem to have been any progress towards AGI.
I do think the last few years have been more productive than previous periods in advancing narrow AI, and to be fair to those researchers who just get on with the work, it is not on them if the advances are over-sold by others.
I bet we're closer than most people think. Instruct GPT-3 can do semantic tasks just as efficiently as DALL-E 2 can draw. NLP tasks that took whole teams multiple years can be simply described in a few words and they work right away.
The entry barrier to implement new tasks will get very low. The large models will be the new operating system. This means more investments and data, leading to new improvements.
I believe GPT-3 is already close to median human level on most semantic tasks that fit in a 4000 token window. I'm researching how to use it right now for a variety of tasks, it just works from plain text requirements with no training.
But there's a quantum leap or two from (say) mindlessly producing comments that can occasionally fool readers on HN (as it has been used to do in the past) to it consciously joining in the conversation of its own volition and curiosity, then zoning out on Netflix while half-worrying about the future for GPT-4 jr and idly planning it's next server room refit.
I'm actually more curious if we could parse the underlying logic that ultimately it emulates to merge those images together.
It 'looks like' something kind of sophisticated is being modelled with AI but there's some nice algorithms hidden in there.
The training principle of CLIP is very simple, but intuitively understanding how the diffusion prior maps between semantically similar textual and visual representations is a bit more unclear (if that's even a well-formulated question!)
More like 'averaging them' and finding variations from vast inputs.
Which is more a long the lines of what I mean.
> While our model can render a wide variety of text prompts zero-shot, it can can have difficulty producing realistic im ages for complex prompts. Therefore, we provide our model with editing capabilities in addition to zero-shot generation, which allows humans to iteratively improve model samples until they match more complex prompts. Specifically, we fine-tune our model to perform image inpainting, finding that it is capable of making realistic edits to existing images using natural language prompts.
Unless that only applies to GLIDE and not to DAL-E?
Even during that lull between GPT3 and DALL-E/CLIP, there was tons of truly wonderful advances in AI...
> "The fundamental principles of training CLIP are quite simple: First, all images and their associated captions are passed through their respective encoders, mapping all objects into an m-dimensional space."
Not scared to admit I don't find this simple at all and I'm probably not in the target audience. I'd love a description that doesn't assume machine learning basics. Is there one?
it's "simple" because how it works is "just" brute-fucking-force. of course coming up with the architecture and making it fast (so it scales up well) is the challenge.
and scaling works .. because .. well, no one knows why (but likely because it's just a nice architecture for learning, evolution also converged on it without knowing why)
see also: https://www.gwern.net/Scaling-hypothesis
Even an RTX 3080 is a complete non-starter.
[1] https://www.nvidia.com/content/dam/en-zz/Solutions/design-vi...
This article predicts that GPT-3 cost $10-$20m to train. I imagine DALL-E could cost even more: https://lastweekin.ai/p/gpt-3-is-no-longer-the-only-game?s=r
Eleuther.ai or some other open source / open research developers will likely try to reproduce DALL-E 2 but it'll take some time and a lot of donated hardware and cycles.
Few groups have that kind of money to commit, also the viability is not yet very clear , i.e. how much the model with make if commercialized so they can recoup the investment.
There is also cost of running the model on each API call, of course not factoring in any of the employee and other costs for sales marketing etc.
[1] https://venturebeat.com/2020/06/01/ai-machine-learning-opena...
Also, practically from a data point of view, the same object can be represented in numerous ways (different artistic styles, different filters, abstract paintings, etc.) and the model has to optimize across all of these samples. What this means is that the model truly is forced to learn the semantic meaning behind a concept and not just rely on specific features.
Check out the dropdown under the "Significance of CLIP to DALL-E 2" section in the article
Put it this way: The model file is absurdly smaller than the half billion source images files. If it actually contained the source images, it would be the greatest feat of image compression ever. Instead it only contains the impression left over by the images. A lot closer to a memory than a jpg.
[1] q.v. $1 billion investment https://techcrunch.com/2019/07/22/microsoft-invests-1-billio...
"CLIP is trained on hundreds of millions of images and their associated captions..."
Does anyone have any insight as to which images were trained on? Was it all open-domain stuff? And if not were the original authors of those images made aware their work was being use to train an AI that would likely put them out of work? Were they compensated appropriately?
https://arxiv.org/abs/2103.00020
Theoretically, you could build a web-scraping tool to do something like this, but even storing that data would take an absolutely insane amount of storage.
I would assume OpenAI has some deal with Meta to make the creation of datasets like this easier.
Many professional artists stake their career on one unique style of art that they have honed and developed over many years. It's this unique style that clients generally pay for, and that now faces a very real threat of being stolen from them by a technology that frankly no human can hope to compete with. Without artist compensation, this can only lead to artists terminating their careers early once the AI has co-opted all work from them. Or future artists never beginning their careers in the first place. This is a net loss for humanity, as it will deprive us of works and styles of art that have yet to be imagined.
I'm not saying AI like this needs to go away. There is no putting that genie back in the bottle, of course. But it needs to be something that artists opt into. If someone's style is worth it for OpenAI to train on, then that style obviously should have a price tag. And it ought to be up to the artist whether they want to sell or not. Anything short of that is theft in my eyes.
Openai just uses the word Open in their name, they are commercial company like any other, they are a business first.
Knowing that there is potentially copyright material does not give you standing to sue them, unless you can show reasonable cause that they could have used your copyrighted content no court will take it to discovery to prove it conclusively, your case will be thrown out for lack of standing.
However if they share their data set, then you can show that they actually took it and have standing to sue them. Making any potential case more complex and expensive.
I can't think of a single good reason for this to exist that doesn't have huge negative impacts on our world.
Why pay an artist/graphic designer when this does what you need?
"Now those damned creatives can go and find real jobs"
Machines aiding in art is only a good thing, because it can maximize output and minimize input?
Makes art cheaper, more accessible, allows more people to create?
It is like how digital filmmaking has cracked the Hollywood monopoly on content.
I aldo don't think that it makes it "cheaper, more accessible, and allows more people to create". Digital art supplies being something readily available and relatively cheap to their classic counterparts is what makes things more accessible, and to make it more so would be to drive the cost down or something. Having the computer draw for you isn't exactly creating art.
And art isn't a commodity and I argue it shouldn't be a commodity. It's something, again, personal and special.
And this doesn't end at the visual arts, I think it applies too to writing. AI could write what's written in my journal word for word but my journal would have more value just by virtue of it being written by me.
I disagree. It is like using sampled music or an arpeggiator or drum track to compose music.
a gut wrenching image of innocents being beat by police is gut wrenching because it's something that exists in the real world
Can't a painting be gut wrenching? It doesn't exist in the real world.
Think long term. Eventually AI will be able to do most of human jobs. As a result, products and services will become cheaper. As a result, people will have to work less for a living. As a result, more people will be able to draw and paint for pleasure, and not necessarily to make a buck.
1) AI appears to have approximately zero chance of making housing and food and other basic needs cheaper.
2) Artists WANT to make money for creating art, music, etc.
Yeah some people are going to loose jobs over this, happens all the time. People are not isolated from the market, they function on it and need to take it into account.
I think we've seen this play out before and instead of reducing work, our standards of living increase and people keep working about the same amount. See e.g. the post industrial world where homemakers had to scrub clothes, then got machines to do the scrubbing, but subsequently had to clean the clothes more frequently.
We might be able to reduce the overall amount of human work only through extremely successful social/political reforms similar to the ones that outlawed child labor and established the 40 hour work week. Assuming the technology will cause it to happen is bound to lead to disappointment.
This is ahistorical. The fact is that you must at least seem to produce more market value than your total compensation in order for a company to hire you. There will simply be less people who make a "livable" wage while those who own these automations will become increasingly wealthy. Depending on how the market changes, there may also be increasing unemployment. But why would that matter? Unless unemployment gets too high, the market will continue to work as usual.
There's simply no reason for the owners and inheritors of an increasingly automated economy to share the value increase with their workers. The worker's wages will be market-determined just as before. Perhaps if unemployment gets too high it will be in their interests to offer something like UBI, though no reason for anything beyond what's strictly necessary for the economy to function, and the minimum required to avoid excessive social turmoil.
https://fred.stlouisfed.org/series/LNS11300001
For women, since about 2000:
https://fred.stlouisfed.org/series/LNS11300002
Combined:
Maybe more people can be game developers with access to free original artwork at their fingertips.
I don’t see it as replacing artists, I see it as amplifying artists.
Curious for her thoughts on DALL-E, I pulled out my phone and invited her to generate some imagery. (I have early access via a family member at OpenAI.) She didn't skip a beat, and immediately started getting creative with it. We even did a "collaborative piece" à la Mad Lib.
I asked her if she felt threatened by DALL-E. Surprised by the question, she said: "No! I could see this really accelerating my process. Sometimes I'm blocked on an idea and I could see this being a great tool for finding inspiration. Can I get access to this?"
My take-away was that art is not zero-sum: someone's art isn't "less" because more entities are creating art. If computers can do it too — even if they're arguably more mechanical in the recombination of existing ideas (note: humans do the same) — nothing stops human art from being art.
An immediate thought is that locked-in people who can only communicate by text would be able to share their thoughts more expressively.
In terms of the creation loop, anyone can create a bunch of AI-generated images. Wombo is huge right now. The differentiating factors will be prompt design, commitment to iteration, aesthetic-driven curation of generated works and presentation.
Photographers take and process thousands of photos to create just one masterpiece.
Art is zero sum in that there are a limited number of artist residencies, exhibitions and funds available.
In this case, we will likely see further contraction in the number of artists able to support themselves. There will ofc always be the super stars and hobbyists.
Artists who are willing to direct their talents towards satisfying others' desires for art will find the world is very positive sum. Those that vie for a limited number of spots in a prestige game may find that it's zero or even negative sum, but those are not good games to play anyways.
People look at the objective reality, provided by the sources that should have the most credibility, and just shrug it off.
If I saw a masterfully crafted video of vaccines actually being implanted with microchips, wouldn't I believe it? I'm not an expert on identifying deepfakes, nor should I be just to consume media. I think this is a valid cause for concern and will make things worse rather than keep it the same.
I wouldn't believe a masterfully crafted video of vaccines actually being implanted with microchips unless the video were authenticated by at least one reputable news source. Provenance matters, and just like we don't believe extraordinary things based on single out-of-context photograph, we shouldn't believe extraordinary things based on a single out-of-context video.
You pretend as if the news actually bothers to corroborate everything it prints. Sometimes, they actively disengage from corroboration or critical thinking, particularly when it's favorable to their party or unfavorable to their party (ALL news sources are biased).
The news is also entirely corruptible. They already have been for some time.
If the incentives are there for the owners of the media to craft a fiction, or to support a fictional or exaggerated narrative, they will do it.
But even if some imaginary world existed where the media was actually incorruptible, suppose they got duped and ran a deep fake video as if it were real news. Human psychology is such that the video could still take root in the popular imagination and influence real world outcomes as a result. Even after being demonstrated of a video's inauthenticity.
I'm deeply worried that we don't have the proper psychological immune systems to weed out deeply fake audiovisual productions and prevent them from influencing our decision making or perception of reality.
you can't demand some technology not to be used when it is not a weapon
there isn't a reason to believe that our current world is in a stage that is free from changes, in fact our world become what it is due to invention of disruptive technologies, regardless you like it or not.