It's not at all surprising that an AI is bad at drawing graphs, and it is also not surprising that even a non-artist human can draw graphs pretty well.
It's not at all surprising that an AI is bad at drawing graphs, and it is also not surprising that even a non-artist human can draw graphs pretty well.
You could equally say "it's not surprising that DALL-E can't draw words"... except that Imagen seems to be pretty good at it.
I think the real reason it's not surprising to you is that you've already seen enough DALL-E results to understand its limitations. It's not surprising that DALL-E can't draw graphs.
If I want to convey happy emotions in the style of Rembrandt, SD or DALL-E will do brilliantly. If I want an apple BELOW a table, or worse, a geometric shape like a triangle, they'll crash-and-burn.
GPT-3 is also really empathetic, but struggles simple logic (and especially mathematics).
Graphs are like the horror case for these.
I can think of ways to make them better at this, but it's not a weekend of hacking.
Why are so many people overthinking this?
From reading the comments here, they're overthinking it because they seem to be taking this as a pre-planned "attack" on AI art generation, rather than just an interesting anecdote on the limitations of these tools.
As someone who has not played with said tools, it was an outcome I found interesting to know: DALL-E et al. can't do specific graphs or even specific logical things very well a lot of the time. That's good to know, and I didn't previously!
Still found the post really interesting as it explores a very realistic use case. A client needs something simple designed for a blog post. Should they use AI or a human designer?
I read somewhere in the comments here that these tools are very bad at counting. Which is an interesting limitation with far reaching implications.
It may not make much difference, but it's not so much that they're bad at counting, as that they don't even try. The way the prompt is parsed and diffused doesn't allow for that sort of logic.
All a "two people" prompt or some such provides, is a hint to push the AI towards that section of latent space where training-set images titled "two people" exists.
That's not "counting", and it would be truly amazing if any sense of math emerged de-novo from this training process. Doesn't mean it can't be done — it means we aren't trying.
It'll be pretty exciting times once we do!
The author could also compare how well DALL·E draws text, but what would be the point of that? Is not being a scientific article a good defense for posting nonsense?