Limitations of DALL-E
twitter.com
twitter.com
These aspects were much better in the prior version (GLIDE), which conditioned on text tokens instead of just a fixed length vector.
If you want to get started I can recommend the EleutherAI[1] discord. It seems to have a good mix of experience and excitement.
Any pointers on where one has to stand in order to be targeted by said computing resources?
I too wouldn't know how to draw anything related to "ophanim" unaided by a dictionary.
I'd also argue that its ability to come up with symbols that look oddly like letters but aren't actual letters could be a form of creativity.
Don't get me wrong, I'm not a GPT-3/DALL-E fanboy but I don't get why this is interesting. Anyone who has worked with a creative designer/artist/architect/etc. would know that there are so many "wrong answers" when it comes to creating visual imagery. Some of these aren't limitations of DALL-E but rather misinterpretations of our own expectations.
otherwise so what if DALL-E has limitations on some weird obscure prompts. As I am hearing about DALL-E the first time, I am still impressed with the images in this tweet thread.
That it struggles with composition sometimes but not other times might indicate model has learned a particular distribution factorization which is good much of the time but makes it unable to properly account for conditional dependencies in certain queries. The issue of losing track of lots of details may also be related.
Its struggles with negation is also fascinating because that has been a thorn since the early days of GOFAI. Negation is trickier and more involved than what one might naively assume.
For those taking relief from such weaknesses, it's worth remembering that this is as bad as it is going to be and probably not for long. Model size is surprisingly small when compared to recent behemoths. That it can already do all that it does is certainly cause for concern.
Well, ish. Negation (specifically, negation-as-failure, i.e. failure to prove) in logic programming introduces non-monotonicity, but reasoning remains sound and even completeness is guaranteed under conditions.
John McCarthy (father of AI and also LISP) was very enthusiastic about nonmonotonic reasoning because he considered it the obvious way to represent common sense which was, for him, the one thing missing from AI systems. In plain words, nonmonotonic reasoning means that you can change your mind as new information becomes available.
But, you have a point that DALL-E's (and other similar systems') inability to deal with negation is reminiscent of older issues with expressing negation. In particular, it reminds me of datalog, a subset of Prolog that does not allow negation as failure. The motivation for that is to guarantee termination. This guarantee comes at the cost of completeness, so there are programs that cannot be expressed in datalog (but can in Prolog).
There might be something similar going on with DALL-E, or maybe it's just a simpler case of not being able to derive negation from correlations, as I point out in another comment.
It's not actually anything. It's just computer-generated nonsense that kinda looks like something.
DALL-E 2 is very impressive. So impressive that many people have wondered if all commercial artists will be out of a job as soon as DALL-E 2 is made publicly available. These limitations form a strong argument that DALL-E 2 is not ready to completely replace human artists just yet. This is particularly relevant if you're an artist, or if you're in the market for one.
https://www.sefaria.org/Siddur_Ashkenaz%2C_Weekday%2C_Shacha...
The marketing and promotion will, of course, focus on successes. Failures suggest boundaries, limits, and weaknesses.
There's a similar case for humans. Studies of optical and auditory illusions, cognitive biases, and various forms of pathological behaviours, both individual and group, would be examples of this.
Or any equipment or system. The flight envelope of an aircraft is literally the boundary beyond which its behaviour is unpredictable and/or uncontrollable. For software, limits define the extent to which applications or AI algorithms produce useful results.
How much those limits can be pushed is itself an interesting question. The observations of failures of negation (a map of the world without Europe, a bowl of fruit without apples, a man not running) suggest that DALL-E has something of a "now don't think of an elephant" problem. It's also a problem that's been observed in search contexts --- I am recalling a case where Google and/or Amazon repeatedly failed when requested to find shirts without stripes, consistently presenting striped shirts.
Perhaps semantic negation is hard?
"If you please ... draw me a sheep."
http://www.supercoloring.com/sites/default/files/styles/colo...
Okay, so the author had to iterate on a couple of examples. But look at the speed. Most of these examples would take a human a day to a week of labor, for each one. The fact you can rattle off almost any random text prompt like "draw me a walrus in the style of pointillism with a fake looking mustache" and that it can accommodate such requests in one second is absolutely stunning.
More to the point, it doesn't do semantics at all, only syntax, if you will. It has learned correlations between images and words but it has no way to map either the words or the images to their meanings. All it knows is that this word appears often in correlation with that image. So when it sees the word, it gives you the image.
What is the word "no" correlated with? Probably, absolutely anything. Most likely "no" is as probable to appear as the description of an image with apples as it is to appear as the description of an image without apples. But "apples" is likely to describe an image with apples, even if it has "no" in it. For example "An image of apples in a basket with no handle" has "no" and "apples". So- you get apples, "no" or not.
In a sense, nobody has told it that "no" means "no".
(Also, it doesn't know that "apples" means "apples". "Apples" is only the label of some sets of pixels in a space. If you scrambled the labels so that "apples" was correlated with images of elephants... "an image with no apples" would get you an image with elephants).
"No Means No" would make a great title for some future paper on AI NLP making significant progress on this problem.
Of course, I would always expect limitations based on having to "understand" anything beyond word-image associations, like floor plans or solvable mazes, which have meaning well beyond the visual. Those limitations are indeed uninteresting. Might as well try "An image proving the Collatz Conjecture"
If I ask for a red box on top of a blue box on top of a yellow box, and the model gives me an image with two boxes side-by-side (albeit a very high quality one), it's hard to believe there's much "understanding" of (a) box, (b) colour, (c) relative positioning. Likewise if I ask for six penguins and it gives me four penguins, can we confidently say there's any "understanding" of what a penguin is?
Again, that's not to say it's not mind-blowing technology/progress in machine learning. It most certainly is. But I'm not ready to give up on every other angle and go all-in on large diffusion models as the path to AGI.
DALL-E right now doesn't have the facility to ask questions, but iterating on the request statement is a bit of a proxy for it. I'd argue that this is exactly what understanding does in fact look like.
That isn't to say that DALL-E is AGI or conscious or anything like that. It's just to say that the mere presence of occasional misunderstanding is not a marker of un-intelligence. It's precisely when you are able to access the richness of symbolic reasoning that misinterpretation is at greatest risk, e.g. your calculator never misinterprets you.
DALL-E is very good at matching words to pictures and combining the results in novel ways. It isn't intelligence and it isn't art, and never will be. It's a tool that humans can use to accomplish tasks. That has its own merit.
The model has to encode those words into a fixed length vector, an architectural design decision, not an inherent requirement. For example translation systems can translate such phrases perfectly, but they don't pass the representation through a single vector bottleneck.
What if you ask it to draw some Nazi paraphernalia and 3 sims?
https://www.pcgamer.com/bizarre-russian-propaganda-links-sim...
If it learns from humans and produces results like humans would... What's the definition of "limitation" here?
The designer photograph one is complicated and I'm not surprised that DALL-E misunderstand it.
But you're cherry-picking for the apples. It didn't consistently give an picture of five apples when asked for five apples. When asked for a bowl of fruit with no apples, it gives a bowl of only apples. That's obviously strongly sub-human comprehension.
A counterpoint to the maze - it depends what we expect. If you wanted a maze designed, it would be of course a bad example. But if you wanted one painted, it's... good? If it was in a gallery, it could easily be a commentary on life with limitations, dead ends and problems we can't solve.
The computer has no understanding of what it’s doing. And I don’t think you can argue that it’s interrupting a brief for a commission. Is it even producing at???
A human needs to be there at the beginning to either dream up a concept and then describe it in painful detail to the computer.
Or a human needs to be there at the end to sift through the results to pick out the ones that actually match some concept.
In trying to think of a useful application for this: if an artist can lock down the style of the output, and then pump in a script and further description to story board things or get near final output, that could be quite useful.
https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
People draw bicycle-like shapes, that don't quite work - they miss important details, or connect things that shouldn't be connected.
In other words: humans mis-draw things because we don't always do a good job of visually apprehending the world; DALL-E mis-draws things because it doesn't really understand the problem statement.
(I think this would be apparent if participants were asked to draw a bicycle, followed by a bicycle without a wheel. The two drawings would be roughly as "good" in terms of quality, but the humans will correctly obey the prompt and omit one of the two wheels. I suspect DALL-E would struggle with that prompt.)
And note that even if it did (which is definitely architecturally possible given progress in retrieval and demos like WebGPT), it still wouldn't necessarily be 'copying', in the goalpost-moving excuse.
Facebook has a neat diffusion model which uses retrieval on a dataset of images for generation: https://arxiv.org/abs/2204.02849#facebook You can see the exemplars it uses: they both greatly improve the quality of generated samples, but also don't look that much like the generated sample.
Maybe a better comparison would be a trained artist, who likely wouldn't make such egregious mistakes.
No, humans can understand "A red cube, on top of a yellow cube, to the left of a green cube" and draw it.
Dall-E gives you nonsense -- look at the picture.
It fundamentally doesn't understand things that humans do. That is limitation.
Of course it might be great overall and do things that humans can't, but it still has a limitation.
Gary Marcus has been hammering this home at least since 2018, probably earlier. It's more of a memory/association/recall machine than a machine that understands, and understands logic in particular.
"seven red lines, all of them perpendicular, some in green ink, some transparent, in the shape of a kitten"
Then it will respond:
"FU, it cannot be drawn!"
GPT-3 can complete this prompt:
clarify this scene description:
A bowl of fruit containing no apples
With The bowl contains bananas, oranges, and grapes.
So the semantic modeling is certainly doable.Need to write some random background NPC dialogue for a game?
Need to generate a draft Cad design for "Airframe supporting a large turbojet engine with X newtons thrust"?
These are problems where the high cost of inference won't be significant, and there is an expert available to revise and discard the result if its not working.
I'd be curious if there is a good startup model for anything in this space, seems like you would already need to be a vendor of the underlying technical software to make a viable business (CAD,Game Engines,Creative Tools etc.)
One application I'm pining for: I'd love Apple to introduce AI into Logic Pro X (LPX) such that it could take a 4, 6, 8 or 16 bar set of tracks and generate infinite variations, from which I can pick and choose ones to embellish.
Or, similarly, I have innumerable LPX projects that are just a single idea that, in the hands of someone more capable (or with more motivation) than myself, could be turned into full pieces. I'd love to be able to give LPX 2 or more of these projects and it generates transitional pieces between them.
All LPX upgrades are free, but I'd happily pay for this feature.
Besides that, we know it doesn't because models don't recurse because no matter how big they are, they still have constant execution times. And we gave up on systematically parsing text in favor of just vibing anyway (technical term).
That's just as well since most parsers didn't really work in practice; I think I vaguely remember one as kind of working (https://www.link.cs.cmu.edu/link/) but it can't handle bad grammar or misspellings like ML can.
But this version of Dall-E is not autoregressive, they've gone for diffusion models in the generator and for a single vector embedding (no internal spatial grid structure) for the encoder. That explains why it's bad at negations, counting and stacking coloured blocks. I believe these skills were not the top priority of OpenAI this time, they understandably went for pretty images.
If Dall-E took variable length sequences as inputs, its execution time would still be O(N) on the length of the sequence and it would still have a fixed amount of memory to work with. Whereas a computer program of length N can do pretty much anything, and a human given a description can go off and think about it.
If it's simply a lack of data, that should be easy to fix. Given an image caption, every other caption is a relevant negation: 'a photo of a happy dog' is also not 'a photo of a sad cat', and can be turned into the caption 'a photo of a happy dog and not a photo of a sad cat'. So lots of obvious data augmentation and synthetic data approaches to apply there to boost negation.
https://langcog.stanford.edu/papers/NF-cogsci2013.pdf
Fun fact: Young Children also don't understand negation. There are some parenting theories claiming that is why one should use avoid negations in commands.
So much so that it feels like a DALL-E ad :)
Arguably most neural networks don't do symbolic, step-by-step reasoning "natively", and maybe human beings don't either. But they can be told to, and then the same network becomes much smarter: https://arxiv.org/abs/2201.11903
I teach students of the first year of the university, and plenty of them have problems with reading comprehension, in particular to understand sentences with negations.
Yeah, a lot of people seem to have problems interpreting instructions for drawings.
E.g. look at what it outputs for welcome signs. It can do a great job for that. Many of the images have artistic use of focus, perspective and blur. That's because:
https://www.shutterstock.com/search/photo+of+a+welcome+sign
Likewise the "omg ai bias" examples are easily explainable as being due to training on stock photo libraries:
https://www.shutterstock.com/search/photo+of+a+nurse
Every single output they get for "photo of a <job>" looks exactly like what you'd get from stock photos.
This might also be why it struggles with counting.
https://www.shutterstock.com/search/an+apple
Easy.
https://www.shutterstock.com/search/5+apples
Hmmm ..... not so easy. The search is good: the results that actually have apples in them tend to have five, but, the actual labels don't include the counts.
Likewise with negation. How many images are annotated with "photo of a man NOT running"? My guess is, nearly none.
If I had a nickel for every time someone asks this...
Perhaps we should investigate how many HN comments are regurgitated vs. novel?
It still has some ideas how humans look like, but that seems to come from classic paintings, not real world photos of humans.
No longer is it a description of what currently exists, but an expression of what's been queried.
Which may demonstrate something entirely different.
Meanwhile if you do it in reverse, take an exist image and try to recreate it in DALL-E2, you can get extremely closer. Over on Reddit they did a experiment:
https://old.reddit.com/r/dalle2/comments/ub0sfg/dalle_2_imit...
The first 90% of the project takes 90% of the time, the last 10% of the project takes the other 90% of the time.
I don't think it's wise to speculate at this point on what problems are "actually easy" or "actually hard", or how long it'll take to achieve certain milestones. I don't think that even the people at the forefront of this research have any way of knowing for sure.
For some projects, 180% would have been an improvement. I've even had a few projects take ~500% of the time. Arguably I should have cut my losses sooner, but I wasn't very good at recognizing the sunk cost fallacy back then.
We also mustn't underestimate the extent in which the human brain is a fuzzy ensemble of partial experts, and composes the first few salient answers provided rather than trying to find a "best" answer that would be too late. Or how often the human brain creates post hoc rationalizations simply to provide narrative continuity.
Well, the prompt is actually a bit vague - is it a monk sitting together with a skeleton, or is it a single entity that is both a monk and a skeleton? Seems like DALL-E tried to find some middle ground.
My thought exactly. I'd be interested to know whether any attempts were made at explicit disambiguation (eg. "A monk sitting beside a skeleton", etc.).
thanks!
It doesn't get close to the quality of the art pictures produced using DALL-E, but it's nice to play with.