DALL·E 2 vs. $10 Fiverr Commissions
simonberens.me
simonberens.me
90% of my clients couldn't do anything without a human in chat that walks them through all the steps. There's no possible interface simple enough for them to do everything without my help. They can't figure out which files they want and what to do with them once they got it. If there's any possible customisation option - they will use it to make the pre-made template uglier, and then will ask me if I could do something to make it look good again. That's what they are paying me for.
Do you think that people in your line of work, or adjacent lines of work, will use AI to offload brainstorming or to get inspiration?
My guess (as a complete outsider) is that the skill of drawing will remain important, but that there will emerge a new skill: an AI translator, who serves as a midwife for the creation of AI art.
But, stock images have existed for many years. They are considerably cheaper than custom work, are as professional looking and are available immediately. Sounds like an absolute game changer, but in reality the market for custom design work didn't die.
I am not sure that I understand all of the reasons why people pay extra for custom design work in a world where automated stock services exist. Some of my guesses are
- People don't trust their visual taste and want a trusted human to make those decisions for them
- Discovery problem. People are simply unaware of such services and their benefits
- People are willing to pay premium for the knowledge that their design has a human author.
- Last mile problem. Even if the image looks 99% like what you want, you might still need a guy to save it / fix it / crop it / format it because you don't know how to do it yourself.
I am sure that there are more factors. And even if AI images will bridge the quality gap to human-made stock images, all of this will still apply to them. Many services and technologies have been trying to solve those problems for many years. AI will add to that process, but I don't see a reason for a dramatic change in the near future.
That's the next step for AI generation. The AI image will be almost what you want but you will hire someone to fix it
Do people ask for graphs on fiverr like the article? (I can only imagine sort of "must have a powerpoint ready for 9am in Tokyo sort of thing. I know that's a real industry even if that industry always seemed to me like everyone gathering round a fake painting with everyone knowing it's a fake)
Anyway - always interested.
If the client likes the logo, but can't figure out the interface, or want me to apply some changes - he contacts me. I talk to him, do everything that he requests and at the end I sell him the logo + premium for my time and additional custom work.
Usually people ask me for things similar to what they saw in my portfolio. Rarely do I get unusual requests like this graph. If I get a request that I can't do - I will say no or refer them to the graph guy. But if I am feeling creative - I tell them an unreasonably high price. Sometimes they agree and it turns out that I was a graph guy all along
Oct 2022 - "Even after $100 billion, self-driving cars are going nowhere." [1]
People have been historically notoriously bad at predicting how good AI/technology will be in 5-10 years time. If the predictions from 2015 were right, the roads would have been filled with level 4 and 5 autonomous vehicles for years now.
[0] https://www.businessinsider.com/companies-making-driverless-...
[1] https://www.bloomberg.com/news/features/2022-10-06/even-afte...
But NovelAI and Stable Diffusion both have limitations. It's nearly impossible to generate two different specified characters, much less specify two characters interacting in a certain way. For NovelAI, common/popular art styles are available, but you can't use the style of an artist with ~200 pictures. (Understandable, given how the AI works technically, but still a shortcoming from a user's perspective.) Both are awful at anything that requires precision, like a website design or charts (as shown in the article). And, as most people know by now, human hands and feet are more miss than hit.
People are extrapolating the initial, enormous step change as a consistent rate of change of improvement, just like what was done with self-driving cars. People are handwaving SD's current limitations away; "it just needs more training data" or "it just needs different training data." That's what people said about autonomous vehicles; it just needed more training data, and then it would be able to drive in snow and rain, or be able to navigate construction zones. Except $100 billion of training data later, these issues still haven't been resolved.
It'd be awesome if I were wrong and these issues were resolved. Maybe a version of SD or similar that lets me describe multiple characters in a scene performing different actions is right around the corner. But until I actually see it, I'm not assuming that its capabilities are going to move by leaps and bounds.
My partner works in design and her design teams have jumped all in on using Stable Diffusion in their workflows, something that is effectively in "version 1." For concept art especially it is incredibly useful. They can easily generate hundreds to thousands of images per hour and yes, while SD is not great at hands and faces, if you generate hundreds or thousands of images, you get MANY which have perfect hands and faces. Additionally it's possible to chain together Stable Diffusion with other models like GFPGAN and ERSGAN, for up-ressing, fixing faces, etc.
Self driving cars are completely different, no one was using "version 1" of self driving cars within weeks of the software existing. Stable Diffusion and similar models are commercially viable right now and are only getting better in combination with other models and improved training sets.
I think you're shifting the goalposts to what success is here to be quite frank. "The model needs me to be able to specify multiple characters in a scene all performing different actions."
The truth is, if I had to ask art professionals on Fiverr for "beautiful art photography of multiple characters doing different actions", it would be difficult and expensive for them too! And worse, you would get one set of pictures for your money and if you weren't satisfied, you're shit out of luck! On my PC, Stable Diffusion can crank out > 1000 unique pictures per hour until I'm satisfied.
I do agree if you are coming from the angle of "I need concept art of a surreal alien techbase for a sci-fi movie[0]" then SD&co are super useful. I'm not saying they don't have their uses. But those uses are a lot more limited than people seem to appreciate.
> I think you're shifting the goalposts to what success is here to be quite frank. "The model needs me to be able to specify multiple characters in a scene all performing different actions."
Having multiple, different characters in a picture/scene interacting in some way is not an uncommon, unrealistic requirement.
[0] high res, 4k, 8k frostbite engine, by greg rutkowski, by artgerm, incredibly detailed, masterpiece.
It's possible that Stable Diffusion, or minor improvements of, is our peak for the next few decades.
It seems like most would rather wait until autonomous cars are way better than human drivers while not truly acknowledging most human drivers are awful. Sure I dont want people hurt or killed but I think it could have made more progress in prod so to speak.
No, the reason is that for city driving there is no system that is even close to navigating typical driving problems that humans encounter multiple times on a daily basis. There are plenty of videos of self driving cars flummoxed by basic road obstacles.
What people like you call “edge cases” are actually common occurrences.
If you think any non geofenced system is close to average human level competence you are simply deluded.
Do they stop in front of a cardboard box and just stand there for minutes?
Likewise, plenty of people just stop paying attention and read their phones ... idling at intersections much much longer than necessary. Or drive stoned and drive around at ridiculously slow speeds.
I think what we see are CEOs looking to raise funds, and news organizations looking to sell an interesting story that will say "revolutionary tech is just around the corner", but this is motivated reasoning. You're right that this is the same with AI technology, where some people say AGI is just around the corner, whereas some veterans say it may well be decades still, and the truth is we don't know.
So anyway I guess I agree with what you are saying, which is that AI development is difficult to predict and many people make bad predictions. I just wanted to point out that it tends to be people with a motivation to predict rapid growth that tend to produce a lot of these errors. These errors get propagated widely because technology press is one of those groups with this bias. However not everyone makes such bad predictions.
Taking the people who are most incentivized to overhype things to get clicks and/or funding as the consensus view is maybe not the best take here.
If you looked at people in general or engineers in general and looked at the median predicted timeframe, it would've probably been much more conservative.
My wife ended to turning her artistic abilities into a greetings cards / wedding stationery because her social anxiety and low self esteem make it extremely difficult for her to work through the process of figuring out what the customer actually wants and how much she should charge for a commission. The way she describes it, many customers think that they can give you a one-sentence request and get back exactly what's inside their head, except that there is nothing inside their head at all, just a very loose idea. Essentially, they want to flip through an infinite set of mock-ups (that they don't pay for) until they finally stab one with their finger and say "THIS!", but they have no idea in advance what "this" is. When they finally come to payment, they only want to pay for the time it took you to produce the final result, which is "just a simple design!"
In fact, the red-flag customers sound like this: "Hello. I'm looking for the simplest thing in the world and it probably won't take an amazing artist like you 15 minutes to make. It'll be used as a logo at our business so it would be great publicity for you!"
Person doesn't value your skill and will try to low-ball you. Ask them to clarify their one-sentence request and they say "Oh, you know, just a simple logo with something nautical on it". Tell them you'll charge for every set of mock-ups as you slowly figure out what they want, and they disappear.
I think that tools like these could be the first step in your journey with a customer. They have to explain to AI what they want, and refine their statement to the point where it produces "mock-ups" something in the right ballpark. Then you can take their top 3 results and talk through them.
I'm totally with your wife, btw, the attitude of her customers sounds horrible. On the other hand, my experience is that one artist took my $25 and has still not produced what he agreed three months later, and yet asked me if I had more work for him. Another guy offered to do it for free and did it for free in a few days and then refused to accept my money when I explained that I was already paying another guy for the same task so it was only fair that I paid him, too. This was some cover art for a vanity project of mine and I was asking for free contributions but also paid the first artist because he was evidently trying to become a professional. Fat chance of that. Bottom line, if you want good art you have to find the people who are passionate about it.
Oh and image models can't create the art I want, because it's text-based art. Even if they could generate the images I want, they couldn't output them in ASCII or ANSI. In fact I tried and they give me kind of pixelated results, but not recognisably text-character based.
I guess I'd say, the less specified your prompt, the more it seems the AI is able to "read your mind." An interesting little tidbit in the world of human/machine interaction. It's like the results make you say, "Yes, that IS what I was thinking of!" But as soon as you have a really specific idea in mind it kind of stumbles a bit for me.
Am I being engaged right now, was your comment also to generate engagement.hm.
I don't think OP chose graphs because they're "obviously" going to make AI look bad; I think he chose it because it's an incredibly simple image - extremely so. If the AI can't do this, how can you trust it to generate something complex? If it literally can't yet draw basic lines as described, how can it illustrate a story or any form of media where specifics matter?
And I don't think his post title implies that he was going to use some complex art prompt, either. Not in any way.
Its also an entirely different task than the one these AIs were designed to solve. Its like judging a fish by its ability to fly.
I don't see how this isn't the task these AIs are supposed to solve. They are meant to take a text description and output a corresponding visual result. This just demonstrates the narrow limits on the complexity of the input they can take.
If you're saying they're not designed to deal with inputs more complex than one sentence, then sure, I guess I agree. But this post goes to show that if you require specificity in your desired visual output, then you need more than one sentence's worth of complexity, and therefore the current generation of AIs are not yet broadly usable.
It's about illustrating the current limitations. This post is not implying that the technology is a failure or that it isn't enormous progress.
In the blog post, the humans drew it incorrectly as well (although they got closer). If it was as simple as you say it is, i would not expect the humans to err as well.
> If you're saying they're not designed to deal with inputs more complex than one sentence, then sure, I guess I agree.
Indeed. I would further say its not designed for someone to use it as text directed paintbrush. This is not surprising since human graphic artists dont work that way either, or at least get very pissed off when they are micro managed in that fashion.
That said i think its also fairly obvious that these systems are also not replacements for graphic artists in general. The human element is important for a lot of reasons; graphic artists dont just "draw pictures". I dont think people seriously familiar with these systems have ever seriously suggested it was a full replacement for graphic artists, although in fairness random internet commentators certainly have been having a moral panic over it.
Not to mention its entirely possible that an AI more designed for this task would do better.
> But this post goes to show that if you require specificity in your desired visual output, then you need more than one sentence's worth of complexity, and therefore the current generation of AIs are not yet broadly usable.
I don't really agree that this post showed that, but i would agree that these AIs are not the best tools if you have very specific objective requirements.
AIs are tools not magic, there are things they are good at, but they aren't good at all the things and still require to be used with thought.
> It's about illustrating the current limitations. This post is not implying that the technology is a failure or that it isn't enormous progress.
I think the objection is that this article doesn't really demonstrate a meaningful limitation that wasn't obvious. It feels like a strawman. If dall-e or stable diffusion actually succeded at the task, i would be very impressed and consider it much more impressive than most of the pretty pictures everyone shows off.
DALL-E isn't good with symbols like letters and numbers. It can't do even very much logical / mathematical reasoning. So a graph is one of the worst choices.
What it can do is make aesthetically pleasing images that match basic descriptions. So there are more "complex" images that DALL-E can produce than basic graphs.
The image generation models weren't trained on chart images, everyone already knows they're gonna be bad at that. Fiverr artists will obviously be better, though even then, who the hell is paying people on fiverr to draw generic charts?
If you wanted to compare them, it would make more sense to compare them based on how they're actually used (especially in the case of the AI models): to make art.
Though if your title was more specific, ala "DALL-E 2 vs $10 Fiverr Commissions: Who's Better at Charts?" you'd probably get somewhat fewer complaints. Having the title be generic implies that you're gonna be looking at common/primary use cases.
Stable Diffusion was trained on images of charts and graphs. It knows what a powerpoint presentation and even an excel spreadsheet look like.
Here:
It just doesn't know how to generate a graph like the one it's asked to.
> The image generation models weren't trained on chart images, everyone already knows they're gonna be bad at that.
I have no idea how you could even know what was, or wasn't in those models training sets. Yet you posted with conviction as if you were sure you knew. What's the point of that?
Edit - Also, what do you mean "it obviously wasn't the focus"? The focus of what? The focus of training, or the focus of presenting the results on social media?
You could probably find a few driver's ed teachers who taught their students to do doughnuts too, but saying "driver's ed teachers don't teach their students to do doughnuts" would nonetheless be largely accurate.
And don't call me silly just because you used imprecise language to try to make a vague point with great conviction as if you absolutely knew what you're talking about, when you absolutely didn't. Show some respect to the intellect of your interlocutor, will you?
And, seriously, you haven't answered my question: the focus of what? What do you mean by "it obviously wasn't the focus"?
I think you were emboldened by the downvoting of my comment and assumed you don't need to make sense, but I think the downvoters were downvoting something else than what you refuse to answer.
Try getting a landscape in the style of vincent van gogh for 10$ on fiver though. AI will give you that in seconds easily, and that's what's amazing about it.
The big question with systems like those image generation models is to what extent their generation can be controlled, and how much sense it makes. This is exactly the kind of testing that has to be done to answer such questions. Just flooding social media with cherry-picked successes doesn't help answer any questions at all. Because cherry-picking never does.
To be honest, I don't get the defensiveness of the comments in this thread. Half the comments are trying to call foul by invoking some rule they made up on the spot, according to which "that's not how you should use it". The other half pretend they knew all along what the result would be, and yet they're still upset that someone went and tried it, and posted about it. That kind of reaction is not coming from a place of inquisitiveness, or curiosity, that is for sure. It's just some kind of sclerotic reaction to novelty, people throwing their toys because someone went and did something they hadn't thought about.
> Try getting a landscape in the style of vincent van gogh for 10$ on fiver though.
In another comment posted in this thread I tried to get Stable Diffusion to give me a graph with three lines in the style of van Gogh and other famous artists. I'd be very curious to see what that would look like and I can't imagine it easily. I'm left wondering, because Stable Diffusion can't do it. Maybe I should ask someone on fiverr.
Don't know if this is evidence of "framing for more engagement," but this line irks me. The latent diffusion models are pretty powerful, but I don't think there's anyone claiming that today's diffusion models are able to interpret complicated queries better than humans. The interesting part of diffusion models is that they can produce good results at all, not that they are better than humans. We're not in AGI territory. Even text models are still limited in many ways, and latent diffusion is highly reliant on the text model to produce good results. Even simpler queries can run into quite a lot of problems, that's exactly why a lot of people have been trying to figure out the best prompts to improve results.
These image generation tools are being discussed as something that could replace graphic designers (didn’t OpenAI refuse to open source DALLE-2 at least partially due to this concern?). So it is absolutely a reasonable idea to compare image generation vs a human designer.
Saying that, the prompt the author chose to use was hard to parse even to humans, I am not surprised the tools failed so badly.
If they did claim this concern I think we can safely assume that was a lie. Their business model depend on having the models closed so they can more easily charge for access
Why not get the chance to see some failures, too? Isn't it interesting to know what those models are bad at? There's too few examples of that around so that is definitely a very thing to know.
For all the hype around these systems currently, it’s nice to see some places where it doesn’t work.
I think these systems are best understood as similar to the 9 portrait drawings by that guy on LSD from the 1950s. They seem to be able to simulate some forms of consciousness well such as deep sleep and psychosis and be completely oblivious to others
Wouldn’t be surprised if this is already possible with today’s tech, and just waiting to be built.
edit: just tried OP’s prompt with Codex and Colab and generated this image: https://i.imgur.com/OyxJCbz.png
Not quite accurate, but shows the potential for a better language model or some prompt engineering to encourage fidelity to the prompt
It's not at all surprising that an AI is bad at drawing graphs, and it is also not surprising that even a non-artist human can draw graphs pretty well.
You could equally say "it's not surprising that DALL-E can't draw words"... except that Imagen seems to be pretty good at it.
I think the real reason it's not surprising to you is that you've already seen enough DALL-E results to understand its limitations. It's not surprising that DALL-E can't draw graphs.
Why are so many people overthinking this?
From reading the comments here, they're overthinking it because they seem to be taking this as a pre-planned "attack" on AI art generation, rather than just an interesting anecdote on the limitations of these tools.
As someone who has not played with said tools, it was an outcome I found interesting to know: DALL-E et al. can't do specific graphs or even specific logical things very well a lot of the time. That's good to know, and I didn't previously!
Still found the post really interesting as it explores a very realistic use case. A client needs something simple designed for a blog post. Should they use AI or a human designer?
I read somewhere in the comments here that these tools are very bad at counting. Which is an interesting limitation with far reaching implications.
It may not make much difference, but it's not so much that they're bad at counting, as that they don't even try. The way the prompt is parsed and diffused doesn't allow for that sort of logic.
All a "two people" prompt or some such provides, is a hint to push the AI towards that section of latent space where training-set images titled "two people" exists.
That's not "counting", and it would be truly amazing if any sense of math emerged de-novo from this training process. Doesn't mean it can't be done — it means we aren't trying.
It'll be pretty exciting times once we do!
The author could also compare how well DALL·E draws text, but what would be the point of that? Is not being a scientific article a good defense for posting nonsense?
If I want to convey happy emotions in the style of Rembrandt, SD or DALL-E will do brilliantly. If I want an apple BELOW a table, or worse, a geometric shape like a triangle, they'll crash-and-burn.
GPT-3 is also really empathetic, but struggles simple logic (and especially mathematics).
Graphs are like the horror case for these.
I can think of ways to make them better at this, but it's not a weekend of hacking.
Please do this again with a better prompt.
It's funny that what we don't see is a shorter prompt. If you ran this experiment with just "A graph with 3 slightly wavy lines", maybe the difference between AI and human results would be closer. Maybe that's the basis for a legitimate research project, but it's frustrating that the author takes the ball to the 80-yard-line and just gives up.
1) The prompt uses fairly complex grammar which is incompatible with a token-based parser. In particular, symbolic references like "The third […] starts below the second, and generally follows the second" are going to be lost on it.
2) The prompt includes details which a generative network is spectacularly unlikely to be able to handle, like asking for text labels with words like "prosecution" which are unlikely to be present in its training material. (Generally speaking, image generation models can only output short words which they've seen many times, like "STOP" or "PIZZA", and even those can be iffy.)
3) Speaking of training material, most of the training material given to image generation models consists of photographs and artwork. Technical diagrams are much less common, and when they do encounter those images, they're unlikely to be paired with the sorts of detailed descriptions that would be required to produce them on demand.
And it’s like a five minute job in Inkscape where he could’ve just done the paper drawing in that and be done.
A far more interesting blog post would have been looking at Fivr artists vs AI when it came to producing unique character artwork for games, or logos, or almost anything except what was done instead.
It’s just the wrong tool for the job.
Yes, we all have seen badly generated graphs from DALLE-2 before, so it feels like this is an obviously limitation of AI image generation tools. But why should this be such an obvious thing to absolutely everyone?
Yes, this will be the end for some artists but not for others. DALLE2 et al. are merely new tools for new generation of artists. And, we are still figuring out how to use these tools effectively.
In other words: The “AI” is a tool that we humans will use to get things done faster/better etc. Nothing less, but also not much more.
I always thought in my head that this level of creativity would remain our domain for centuries. Even as of like two or three years ago I thought that.
It’s insane to me that today some artists feel they’re going to be replaced soon. The idea of centuries is completely shattered for me and now I don’t know if we’re a year or 50 years away from AI replacing humans entirely in the creative domain. I spent the other day completely in an existential crisis, tbh.
(The main instigator on Twitter is a guy who draws “realistic Pokemon” and hates that an AI may have stolen the art he already stole from The Pokemon Company.)
From what I've seen these networks are rehashing learning set images into something matching some criteria to produce visually pleasing results. Not to belittle the results - it's impressive - but the stuff I'm not seeing here is understanding of generated material - nonsensical z-order, scale/proportions, configuration.
Fantasy images are an easy target because it's all about visually pleasing nonsense.
But then that's sort of a self-limiting factor: it means theres still space for human creativity in creating new things, new styles (as not every style exists yet!) -- at least until said new style gets loaded into the model, I suppose?
Fascinating stuff, really.
This isn't art. It's a graph meant to represent data.
Now, if you're actually Wizards of the Coast, you probably wanna spend the money with real artists anyway, but for any smaller teams, I can see the appeal of just using AI for that kind of use case now.
In fact, my first urge was to ask you to just draw the dang thing already, so I am very glad you included the sketch later!
This might say more about me than your prompt, though, but I thought I'd share the data point.
Perhaps I would have been more successful if I read the instructions with pencil in hand, sketching it out as I went along instead of trying to fit the whole instructions in my head first and then visualize it.
I have a lot of fun treating AI as an absurdist philosophical visualizer. Feeding it very abstract prompts and getting back bizarre results that somehow make sense!
> However, it seems they didn’t catch the part where I said the black line should go between the first and third lines.
To me it seems the Fiverr person did attempt this part, but misinterpreted it. The black line is behind the blue line, but in front of the green and red lines. Does that count as "between" on the z-axis?
The prompt used by the author was hard to parse - I had to re-read it several times. Not surprised both some humans and AI failed.
Bear in mind that the training data for these models has been mostly images and their alt text, scraped off the web. There is a good chance that there's nothing remotely like the examples given here in the training data. (People don't caption their graphs like that.) These models are undoubtably good at doing what they have been trained to do - but I think no-one disagrees that there's plenty of room for improvement.
(And bear in mind that these text2image models only released this year, and that this tech in general has only been invented in the last couple of years, so it's very early days...)
However, once the wow-factor for text to image AI wears off, you start to realize:
a) “AI” doesn’t understand anything about the real world (physics, proportions, shapes, etc.) You curate and pick the image that makes most sense. You start to settle.
b) TTI “AI” can never evolve and become better than the data it’s been trained on. It can combine things for sure, but it will won’t start evolving and become a master artist.
c) Limited applicability - there’s always going to be an edge case in every image that’s going to look off (strange blends, non-perfect circles, coloring…). You need to be a creative and know how to use photoshop to fix this stuff.
Right now it’s cool for creating album art, but that’s easy [1]. Generally I just hope that creative people aren’t too worried about “AI”, otherwise we’ll run out of training data.
[1] https://www.intheknow.com/post/album-cover-challenge-tiktok/
https://blog.barac.at/a-business-experiment-in-data-dignity
* vs the current trend of training diffusion models on 400M images from the Internet (many of them being garbage) with mixed licenses and letting the user take responsibility of the generated images licensing issues.
I agree that they were trained more on artistic images, but I was still surprised with how bad they generalized to a more theoretical(?) context.
https://www.bdaddik.com/en/comics-collectible-postcards/2967...
Modulo s/Lucky Luke/astronaut/g
Note that the image above should be in Dall-E's training set. So it's seen how a horse rides a human. No excuses there.
I bet you could write a few copilot prompts to generate code which would draw a graph like this, though.
Here's the results of three attempts with slightly different prompts:
As you'll see, Stable Diffusion
a) is perfectly capable of drawing graphs, and,
b) completely incapable of drawing a simple graph with three lines as prompted.
It’s just amazing that non design people like me can just conjure up decent looking, and usable stuff with AI. I will definitely use DALL E much more going forward for creative work
The logos are a bit noisy and need redrawing in a proper vector tool but its a great starting point to try out different ideas immediately
(The results: https://twitter.com/dvcrn/status/1578710631838289922)
I've been using Imagen/Parti/Stable Diffusion already as a replacement for "clip art" because it takes ~15 seconds to get results and they are free. Fiverr takes at least 100 times that long and costs $10.
For tasks where the exact content isn't important and you can't invest more than a few seconds or wait a more than a few seconds for results, generative models are already a great solution.
It seems fastest to just draw it yourself; even the pencil drawing was already decent; and you can buy color pens for less than $12.
I needed graphics for my personal never to be published game project. I’m baaaaad at graphics. I was going to use Fiverr but really didn’t want to spend money on what might be a project I never return to. Years later Dalle2 came around and last week I spent an evening using up all my free credits and got all the art I needed.
I’m confident humans can do a better job. But for what I needed, getting about fifty images for about two hours of work was pretty amazing.
But it's still not perfect, a few more examples in a grid, as you can see hands are still a problem:
(My bias: I'm generally for art to be created by artists. I find AI generated images to be a fun game, though. Exploring the minds of those AIs, in a way.)
It's kinda sad that Google desperately wants the cool points for having its own DL models but you can only see it in the form of a store window display.
They probably also remember signing an NDA, and maybe taking some training about how not all of the world's information should be made universally accessible and useful to everyone. For instance, the contents of a user's inbox.
Getting 3 of something 1/6 of the time doesn't really sound like it groks the request.
I would not use these models for graphs yet, but for cool look tarot inspired "clipart" or background images, I think they are already usable.
The faces often look very good and they also have symmetric complexity and individual elements that come in a specific quantity (2 eyes, etc). Lower quality models do generate fly-like multi eye faces, but newer ones are so much more precise!
Here are the results of the prompts "a graph with 3 lines in the style of X" where X in {Rembrandt, jackson pollock, studio ghibli, escher, van gogh}:
So now we're producing art!
Art that's nothing like the prompt.
I get the point the author is trying to make, but I really wish the example felt less contrived.
I can see how it felt contrived, but I hoped to make an apples-to-apples comparison on a real use case. Then to reduce the complexity I tried it on a much simpler prompt.
Please don't make any more graphs not based on data.
You're not making the world a better place.
How long before someone lists "DALL-E Prompt Optimization" as a skill on LinkedIn?
Let’s show all the ways that these AI obliterate $10 Fiverr non-AI commissions, thats what people want to see
The comment by nextaccountic is spot on.
Who says that? Where does it say that Dall-E and Stable Diffusion are only for generating images and art? Why are graphs not images? And why can't they be art?
Those are models that generate images from textual prompts. Aren't you just moving the goalposts by saying they can't generate specific kinds of images?
The homepage defines the process: "starts with a pattern of random dots and gradually alters that pattern towards an image when it recognizes specific aspects of that image"
So based on that description and context clues, I wouldn't expect it to generate a precise schematic.
I had never seen someone try to generate graphs with dalle so I thought it was worth sharing.
Who made up those rules that drive nerds to rage when they're broken? Where are those rules written down? Can you point at them? Or have people just made up those rules in response to that post?
Checkmate, robots.