Playing with DALL-E 2
lesswrong.com
lesswrong.com
edit: fixed typo
The thing is that current has been producing more and more things that stay at the talking dog level (as well as other definitely useful things, yes). So the impressiveness and the limits are both worth mentioning.
The only real problem with them is that they are conceptually simple. It's always just a subject in the center, a setting and some modifiers. It's doesn't produce complex Renaissance paintings with dozens of characters or a graphic novel telling a story. But that seems more an issue of the short text descriptions it gets, than a fundamental limit in the AI.
As for the AI-typical image artifacts, I don't see that as an issue. Think of it like upscaling an image. If you upscale it too much, you'll see the pixels. In this case you see the AI struggling to produce the necessary level of detail. Scale the image down a bit an all the artifacts go away. It's really no different then taking a real painting done by a human and getting close enough to see the brush strokes. The illusion will fall apart just the same.
Reminds me of the notorious Dropbox comment, or even the iPod one on Slashdot.
I agree these pictures are amazing and any nitpicking is distracting from the actual impressiveness.
Imagine telling someone in 2002 that they can write a description of a dog on their computer, send it over the internet and a server on the other side will return a novel, generated, near-perfect picture of that dog! Oh - sorry, one caveat, the legs might be a little long!
<side eye>
https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/bcf... https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/5c8...
DALL-E definitely does not encode the mechanics of cloth and yet there are regions of perfectly consistent geometry and lighting in the clothing wrinkles.
I have a strong suspicion once these models start getting into mainstream hands it's not going to be uncommon for people to find the input photos that were copied. What do we do when a popular image from DALL-E turns out to be an artist's existing creative work with a style transfer applied?
The most interesting thing about DALL-E is its understanding of the prompt, objects, concepts and scenes. Which has improved significantly but is still very far from perfect.
The rest is a rorschach test that abstracts the flaws away.
I'd be almost as impressed with it if it spewed out stick figures.
In fact I wish it had an unstylized mode, so one can more clearly see and think about the way it composites things.
Having said that the way the visual elements are integrated together without painfully obvious seams or blurring is impressive.
And nor do you, so how are you so sure at a glance that it's consistent? It's probably only perceptually convincing.
Betting that it must be cheating because it's good is classic AI denial.
But I'm also really interested in the possibility of finding obvious traces of specific input images on the output. The fact that "The Emperor of the Galaxy, byzantine mosaic" generates something strongly resembling Jesus may be a sign of this. Maybe someone with more AI knowledge is able to explain how possible/likely that is.
I would say that the merit of the technique is not in generating new images, but in correctly associating the images and the concepts they represent, at the right level (art styles at the low-level pixel representation, and objects as recognizable high-level features).
I agree that it will be common to recognize where some reused elements come from, in special for small concept domains or art styles; this image generation is largely making clever collages. However, at a sufficiently fine-grained level, it is not unlike how most human artists create art.
[1] https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/855...
Even if DALL-E was copying certain segments of an identically in order to compose new images, it is a huge advance. On top of that, much work has been done on GANs and other generative algorithms to investigate how similar output images are to training images.
And before anyone even says it… obviously they have to try and prevent child porn. That’s settled US Law with regards to artistic depictions of underage sex. However everything else they are preventing (and would ban you for if they caught you deliberately found a way to trick it into producing) is protected artistic expression under freedom of speech, first amendment stuff.
This artistic tool is built to prevent you trying to create content containing certain concepts which is a fundamental new factor in how artistic tools shape the art that is created with them.
And in case my viewpoint isn’t clear, I don’t think it’s a good thing, and that we should not allow it to become normal. I don’t object to its existence, because I understand corporate politics and how it leads to this. But this sort of crippled tool has all the concerning potentials of “newspeak” in Orwell’s 1984, except to art not to language.
But is the comparison to 'art' apt?
The platonic ideal is that this will work like the holodeck from the Enterprise, but the skill is all in the machine, there is no effort or skill for the consumer aside from the decision what she wants to see.
I can enter "monkey on a bicycle" in the Google image search and Google shows me pictures of monkeys on bikes. Dalle works exactly the same. Is Googles image search for that reason like a paint brush? Should it make illegal content available? I don't think so, it is a content service. A future Dalle-22 may be virtually indistinguishable from Youtube or Pornhub and A) it should be able decide what it wants to be and B) be bound by the same laws like youtube or pornhub.
This is the way it’s an artistic tool. You can say “monkey on a bicycle” and sure you get it randomly generating stuff, but if I ask for “capuchin monkey riding a red schwin bicycle with whitewall tires and a basket on the front handlebars” I’ve used the tool to craft a specific image I’ve constructed in my imagination, it becomes a tool to take what I have imagined and realise it, it’s an artistic tool just like the photoshop contextual fill tool is an art tool, just way more advanced.
As for future iterations… I don’t see any reason to assume that “Dall-E 22” or even “Dall-E 44” will somehow gain sentience and have tastes or be capable of deciding what content it want to produce for us. The “Tastes” of this model are determined by its training data, as is what it is capable of generating, you mentioned PornHub and that’s a great example, no matter how good the model gets at generating photorealistic things from descriptions, if they don’t include anything in the training data that’s labeled as “dildo” then the mode will have no way of knowing what to generate and will just produce randomised nonsense… so again you would be forced to use it as an artistic tool, to describe the scene constructively like they do with their dead horse example, “horse sleeping in a lake of red liquid” produced an image that looked like a dead horse in a lake of blood. If you have to do this then you are “painting” a scene by writing an elaborate description of everything in it and are using the model as an artistic tool in order to produce the visual image you want.
I'm not clear what you mean by this. The laws regarding CSAM vary a lot across the USA because of the different states and then federal law on top. After a federal Supreme Court ruling that virtual images didn't meet the definition of CSAM a lot of jurisdictions modified their statutes to make constructed images also illegal (e.g. cartoons like "lolicon" come under this). It really comes down to where in the USA you live, and remembering that for most Internet crimes the state can charge you and also the federal government can also charge you. I wouldn't screw with these laws either. If I remember correctly there is one US state that has a mandatory minimum of 100 years in prison for a single image.
I think that ship has finally sailed with this recent release that allows deage-deepfakes:
$ cat /proc/meminfo
MemTotal: 16269536 kB
Thanks, earlyoom looks interesting. But as I know where my memory usage goes (Firefox, mostly) I think I'll just write a script to kill Firefox if I run out. Much better than waiting for the scheduler to kill KDE out from under me!So while it’s a fair position to take that banning it is absurd… as a Corporation in US legal jurisdictions such as I believe California, Washington and Delaware, they will have to comply with state law regarding such matters in three separate jurisdictions, easiest way being to completely prevent the objectionable content entirely.
I think it makes perfect sense to ban child porn, and strongly support it.
Given the fact that millions upon millions of children play incredibly violent video games without feeling the urge to go on killing sprees, I'd say what you're saying is completely specious, but I can understand that critical thinking can be difficult.
You can clearly see [1] how the glyphs are created by approximately merging several letters in one place to create a new symbol, while maintaining the overall structure of words and paragraphs within a well-composed layout in the page. The periodic table [2] seems to be doing the same, where the layout is the structure of colored boxes and two sizes of letters within them.
Image composition seems to be doing something similar, learning to associate concepts with their visual representations at the right level, and merging already-seen examples in the right proportions to create a novel image. Cats and vampires are represented by their distinctive features, arms and legs are correctly positioned as parts of the body according to the action they perform, and instructions of style (either by artists like "Pieter Brueghel", art styles like "digital" or "mosaic", or even camera settings like ISO exposure [3]) are translated into lower level pixel representations of color, shapes and shadowing.
My hypothesis is that if you included examples where the letters are taught one by one, like in kindergarten primers, it may be able to learn those concepts as well and generate better painted text (although I'm not sure it could make the jump to "understanding" the relation between their role as input instructions and as output image text).
[1] https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/855...
[2] https://39669.cdn.cke-cs.com/rQvD3VnunXZu34m86e5f/images/157...
[3] https://www.bramadams.dev/projects/dalle-tricks#let-there-be...
These fall into The Valley for me. I don’t know why, but they do.
Don't get me wrong, DALL-E is remarkable. But it's remarkable in almost the exact same way ELIZA was remarkable in 1966, and Markov Chain generative models were in the 1990s. All of these, given the context computational powers at the time, were both miracles and parlor tricks at the same time.
The trouble is that these demos are impressive, but not much has changed in how the world works. I work in AI/ML and every places I've seen the practical applications of these techniques has ranged the equivalent to adding sprinkles on a cake, to completely useless engineering nightmares that would ultimately create more value with their removal.
Yea, a 3.5 billion parameter model is neat, but we know that in 1970 when we couldn't imagine training such a thing. The problem is that we've made essentially no progresses other than showing when you pour incredible human and physical resources into a pot the result looks cool.
But when you do the accounting, and ask yourself "what has really changed with all these fantastic innovations in machine learning" the answer is surprisingly little.
Dall-E 2 is the type of cool that will be figure 12-3 in a undergrad textbook in 20 years. Students will go "oh that's cool" and turn the page.
Now we can do that in a model that runs locally on your phone. This is a pretty big deal.
As far as something to play with to get interesting ideas, and to take a first few cuts at implementing them, DALL-E is great. Then let's bring the human's technical skill and ability at curation into the mix.
Harry Roolaart
https://twitter.com/nickcammarata/status/1511861061988892675
And as said before, this is only the beginning.
20 years ago (2002) Deep Blue had beating reigning world chess champion Kasparaov was old news.
Unsolved problems were things like unconstrained speech-to-text, image understanding, open question answering on text etc. Playing video games wasn't a problem that was even being considered.
I was working in an adjacent field at the time, and at that point it was unclear if any of these would ever be solved.
> In the end they all fell to specific methods that did not provide general progress.
In the end they all fell to deep neural networks, with basically all progress being made since the 2014 ImageNet revolution where it was proven possible to train deep networks on GPUs.
Now, all these things are possible with the same NN architecture (Transformers), and in a few cases these are done in the same NN (eg DALL-E 2 both understands images and text. It's possible to extract parts of the trained NN and get human-level performance on both image and text understanding tasks).
> While the current progress on image related problems is great, if it does not lead to general advances then an AI winter will follow.
"current progress on image related problems is great" - it's much more broad than that.
"if it does not lead to general advances" - it has.
That "count the objects" app that can tell you how many items you have in a photo seems like a very practical application that wasn't possible with traditional CV before ML.
Not far off from image recognition, save for the background segmentation ;)
machine is briefly described here, I've seen a more thorough breakdown somewhere out there... https://distributedmuseum.illinois.edu/exhibit/biological_co...
And calling it "curve fitting" is just disingenuous. There's quite a lot of evidence that "curve fitting" will scale for at least a few more orders of magnitude. Who knows, maybe the human brain is just "curve fitting".
No it hasn't. And the problems are clear - it's impossible to express anything in the rigid hierarchies that symbolic AI requires.
Representing "Britain" in geographic, language, political and economic hierarchies does not allow a model to do any reasoning about what "British sense of humour" means.
"Softer" structures that represent concepts as a "blob" in a multi-dimensional space is clearly a better approach (aka embeddings, and the even better representations that more complex models use are even better).
Representing "Britian" as blob in a multidimensional space that is adjacent to concepts like "satire" and "surreal" as well as people like "John Cleese" lets a model reason about what "British sense of humour" means without being specifically trained.
As GPT-J[1] says when prompted with "A good example of the British sense of humour":
A good example of the British sense of humour is found in George Orwell’s novel, The Lion and the Unicorn. It’s a satire on socialism that is much more sophisticated than anything in contemporary Leftist intellectual thought. It was written during the war, in 1944. The story opens with a visit to a pub in a fictional village in England. The pub is named the Unicorn, but Orwell (or George, as he calls himself) has decided to call it the Lion. He explains:
“The Lion is a pub, just like any other pub, where people drink in the evenings, and talk about their daily business and their hobbies and where they exchange ideas, views, points of view. But the difference is that nobody in the Lion ever argues about anything. They just sit there, saying nothing, drinking nothing, not even beer. People can come to the Lion and buy beer, and leave the Lion and not buy beer. Beer is freely on sale in the Lion, but nobody ever buys it.
(One should note that George Orwell's "The Lion and the Unicorn" is nothing to do with a pub where no one buys beer. BUT the joke is kind of exactly like a British sense of humour).
DALL-E 2 is useful now. It's easy to imagine it replacing 99designs, Getty Images, and most of the digital art services on fiverr. I can't imagine what is going to happen in the world of print-on-demand tshirts.
And of course porn... what a market.
I don't think anyone (yet) imagines all the humans at the NYT will be replaced by GPTX+DALL-E. We still have editors even though humans author the articles.
If you look closely at most of the images, they don't look quite right. There are usually artefacts, or slight misunderstandings of the brief etc. Its about 80% of the way there, but I think that last 20% is going to be a lot more difficult.
I'm sure there are some cool applications for this - maybe if you need a quick and cheap image for your newsletter, personal website or an experimental game for example.
For a serious commercial application I would think it would be easier and safer to pay someone.
I think this understates the market. For every New York Times article, there are millions of newsletters, analyst reports, tshirts, websites, logos, etc. Many of them are abstract eye candy and don't need photorealism.
Actually, even the high-budget articles often have pretty abstract art. I read this Wired article yesterday, and all the illustrations could easily have been generated:
https://www.wired.com/story/tracers-in-the-dark-welcome-to-v...
About usefulness - the CLIP part of the model is a ready made zero shot image classifier. It reduces the amount of work needed for simple image classification tasks to just naming the classes. The generative part is good enough for illustrations. It will make an average web designer have the powers of a graphical artist.
Unfortunately the models are restricted and expensive today. I hope to see a real open AI initiative to train such models and share the weights, but can't hope that from OpenAI.
It really doesn't seem we've hit diminishing returns on LLM model size yet, and that's not accounting for the multi-input types where video, text, audio, and more are being fused together.
That's a microcosm of the effect of automation on the economy as a whole.
We can't have it both ways. We can't pretend like automation will never kill our jobs while simultaneously pursuing the dream of permeant vacation through automation. The very goal of automation contradicts the idea that we would and should always have jobs.
But we are very very far from that point. We have needs, like extending our life, that current technology cannot meet, and if we had excess time, most of us would trade it in exchange for more technological progress toward that goal. Fundamentally that is what's behind the increase in healthcare spending, and the growth in people employed in healthcare.
No they won't. If people work fewer hours, they'll simply be paid less, barring legal limitations like a minimum wage. Meanwhile, costs will remain the same.
My productivity has definitely been increased by Copilot and there's potential for more increases, but I don't see where that would replace the programmer.
Worth considering: https://arxiv.org/abs/2112.04035
Today, if you need a logo made or some clipart for your web page, to do it the correct, legal way, you have to either get lucky with some stock artwork or scout around for an artist, evaluate portfolios, select a few and buy some samples, decide on one and iterate back and forth until you have something you like. Then you have to have legal agreements in place, make sure rights and copyright and royalties and all that shit is decided, be careful about how you use that art (do I have the rights to put it on a billboard too?)
Imagine a far future where any creative work can just be freely generated with a text description, and the output is unencumbered by IP rights. Type something in and get an infinite scroll of outputs, select one, and you're done.
Extend it to all sorts of media: Music! "Two minute upbeat song about lawn care, jazz style." Out pops an infinite scroll of jingles. "Lullaby for 2 year olds about dogs." "20 minute opera in German, about cycling, in the style of Mozart." Movies! "Three part superhero series where the main character walks backwards." "Romantic comedy but with talking turtles." "Sci fi movie about underwater colonies with a shark villain."
This could be the future if intellectual property lawyers don't fuck it up with artificial scarcity and "digital rights" like they fucked up the copying of bits across the Internet.
This isn't real AI and it didn't come up with these images by imagining them. It's a blob of every image on Google Image Search stuck together in a way that's managed to differentiate between them (in the calculus sense).
For job skills to earn a living, you should definitely be paying attention and adjusting in order to future proof. Learn skills that will not become obsolete.
For skills for fun, don’t bemoan the fact that an AI can do it better. There is still inherent value in you learning and enjoying how to paint. Not everyone needs to be the best in the world. It would be silly to say “I love baking cakes for my family, but there’s no point because of British Baking Show.”
Should you not learn math because of the existence of calculators?
This is amazing, I wish I could play with it.
The image to the prompt "A dog looking curiously in the mirror, in digital style" shows a cat looking into a mirror and seeing a dog as itself! Although very creative and "objectively funny" (may cat is like a dog!!!), I think the AI understood "A dog looking curiously (is) in the mirror".
These poems are absolutely hilarious also, it's photo-realistic, but the messages make no sense, though the fonts look so perfect! That's where you know it's clearly AI generated ^^ For now.... Sometime soon it looks like realistic pictures of protests are going to be so easy to generate with arbitrary text
DALL-E 2 failed all the "X looking curiously in the mirror, but the reflection is Y" tests - showing X as the reflection instead, Dave Orr had to do some "hinting/editing" to achieve this :
> Here's one where I edited out the cat in the mirror and changed the prompt to be about a dog, and it did something sensible.
The inability to spell words is because they're using BPE instead of actual text inputs, so it doesn't actually know that words are made out of letters.
I'd love to search through the training data to figure out what is going on here, but apparently that isn't public available either.
[1] https://twitter.com/Merzmensch/status/1513611885576347658
Some observations:
Note that the writing utensil is always in the right hand. It is more evident after the first image that it is neither a pen or a feather or anything like that but a whispy blurry line that goes nowhere.
The book pages are always blank.
All the creatures are green and in the same pose.
The arms are wrong and often disconnected from the hands.
I believe the way to look at DaLL-E 2 output is to break the prompt and the resulting image into distinct concepts/layers.
Each layer is cribbed from some pre-existing image. Hands from here, arms from there, head from Yoda, swap the face. Blur everything together with a style transfer.
Combining these technologies and the cycle of progress will we soon (lets say within 2 decades) be able to:
1. Generate music (e.g https://openai.com/blog/jukebox)
2. Generate CGI 3D Models (e.g https://www.louisbouchard.ai/ganverse3d/)
3. Generate Plots/Fiction (this is probably the most complex unknown, e.g https://medium.com/the-research-nest/interesting-novels-writ...)
4. Generate Animation (e.g. https://getrad.co/ or https://www.deepmotion.com/Animate-3D)
Finally will a next generation Spielberg be able to create a complete film using these tools sometime in the near future..
FUTURE AI: generate the film "E.T. as a comedy with a female protagonist, and the alien should look like a yoda, set in the 1960s" -or maybe we upload a film script and the AI will generate the movie in "Spielberg style, or Scorsese Style etc"...
finally the porn implications are worrying.. how will we ever control this?
I'm sure some gamedev will add DALL-E for visuals to a game eventually.
First off, he fine-tuned the model on text from an online writing/fan fiction site without permission at all. That would in itself be dodgy enough, but if you know anything about these sites, you know there's a lot of pretty dubious sexual stuff there. And it's not as if he didn't know, because he used only a subset of the stories, he must have picked them himself.
But then, when the AI started generating stuff that would be illegal if it had pictures and OpenAI caught wind of it, he blamed his users, threw them to the wolves to stay in good standing with openAI. He would even ban people for what the model did, i.e. even if their prompt had nothing sexual in it, you could get banned for the smut-finetuned AI taking the story in that direction on its own.
Anything based on OpenAI is going to run into situations like these, they are terrified of bad PR.
Or more prosaically, we're MUCH better at noticing something wrong in humans (especially faces !) than in, say, puppies, so that would make DALL-E 2 look bad to an uninformed observer.
You can easily pass the 'look good to an uniformed observer' test with human faces. Remember, faces were something GANs were doing near-flawlessly back as far as ProGAN all the way back in the dark ages of late 2017. (Then StyleGAN did non-photograph faces, like anime - see my ThisWaifuDoesNotExist for a demo of that.) Doing them as part of a larger composition, where the face is a small part of the images and may vary much more than in closeup centered portraits, is harder, but take a look at how Facebook's DALL-E rival Make-A-Scene does it: https://arxiv.org/abs/2203.13131#facebook They specially target faces as part of the training process, with face-specific detectors/losses, and so the faces come out great.
Since you know that training for photo-realistic faces is going to take extra effort, and that you're not going to allow them - you could just spare that effort ! (Or maybe leave it for later, if the network is flexible enough ?)
> I tried to ask for Dall-E by name but that was a content policy violation.
It seems like they are trying to prevent it from visualizing “itself”.
Is there a less exciting explanation?
they wouldn't want Dall-E's name plastered next to some harmful/offensive content. or generate uncontrolled PR from stupid meme articles. wouldn't be a bad idea to chuck all the companies names/brands in there. wouldn't be surprised if "OpenAI" and "GPT-3" are banned too.
based on the comments, the policy violations cover a lot of territory - covid, explosive, nuclear war.
one commenter said that "glass of guidelines" violates the policy, which is a nonsensical statement. suggesting either "guidelines" is a banned word. Or, maybe they have a content policy classifier, in which case a bunch of random stuff will probably trigger it.
_for sure_ they're not worried the "AI might visualize itself", that's not really a thing
If you mean just could such a model architecture be trained to generate furry erotica? Yes. Erotica, and furry art in general, in Tensorfork's experience, tends to be somewhat harder than regular images because of the more chaotic placement of everything such as limbs, but not that much harder. You might need half again as much compute to get equivalent quality results, perhaps, but probably not, like, 10 times as much.
There have to be a lot of limits once you get far enough out of its training data. Even if you could guide it by image prompts, it has to know what that image is meant to represent.
Also what most people mean by "anime style" isn't anime style, of course - they actually want one of Avatar-type cartoons, advertising key visuals, game characters, or pixiv/DeviantArt fanart. None of that is drawn the same way as TV anime.
They get the image embed from another model that can predict image embeds from text embeds to encourage generalizing when sampling.
It does suffer from some of the same issues with binding attributes and counting stuff, as far as I can tell. But that's not necessarily related to its ability to compose concepts.
But if you don't want a universal model but just a specific niche (e.g. furries) then you should be able to train a much smaller model on much less data and get something interesting much cheaper.
it's like a kid learning to write, but still getting the shapes of letters wrong.
You can effectively use it to make it sound like every other side is unbelievable, or just to tire people into giving up, or get past spam filters by generating better than ever nonsense.
And the in-group will take it as a point of pride to believe lies from their own group anyway.
The first images of a news story available could be fabrications.
What if news organizations were able to put pictures of Saddam's WMDs (that remember, never existed) on frontpage stories while the executive was trying to manufacture consent for entering a war?
I feel like the human brain is going to have a hard time being skeptical of fake news when fabricated fantasies can be rendered realistically.
Then the outcome would be... exactly the outcome we had anyway, which was the OP's point. You can already lie with images with Photoshop. It's a little more work, but you can probably still hide it far better than you can pass off a DALL-E generated image as real.
Lying with images isn't the problem. Telling a consistent lie, getting all the details right even when new evidence comes in, that's a problem, and this sort of AI doesn't help with that at all.
https://thispersondoesnotexist.com/
(just keep refreshing)