DALL-E 3 Is So Good It's Stoking an Artist Revolt Against AI Scraping
bloomberg.com
bloomberg.com
Also, for those in the comments saying it's only available via ChatGPT, it's not. You can use it through both Bing Chat AI and Bing Image Creator for free, and although the latter as speed priority tokens that refresh over time, the Bing Chat AI doesn't and you can prompt it to do multiple images over and over as long as it finished commenting and the image is generating.
If anyone wants to see what DALL-E 3 is capable of, they can check these two hashtags on Twitter/X #DALLE3Beta #DALLE3Art.
Microsoft has decided that women are against their content policy.
$20/mo for DALL-E3 is worth every penny.
Expect something soon where you put in a book and get out a graphic novel. Then, put in a screenplay and get out a storyboard. Then, put in a storyboard and get a movie.
I kind of wish they already let you generate this - you want that celeb in this setting, or a mixed race couple, or all Asian characters, etc, go for it vs studios picking and choosing for you.
[1] https://www.insider.com/chinas-flawless-ai-influencers-the-h...
I am less up-to-date on the current state of AI image generation, and I don't feel like I'm fully up-to-date with DALL-E 3's capabilities. But my experience around GPT has been that sometimes articles like this play up a controversy that may not have really changed much or that might be fueled by different causes or social conditions, and they play up that controversy as a way of playing up the capability of a specific model. This is very common with GPT: "our model is so good that it's dangerous." Sometimes worrying publicly and loudly about something is just a marketing technique.
So I don't know that DALL-E 3 is the reason that artists are mad about this, and I'm not completely sure that they are significantly madder than they used to be, and it's not clear to me reading this article that anything beyond the opt-out form is specific to DALL-E 3. Artists are mad about AI image generation period, they are also upset about OpenAI's opt-out process (as one of many processes, Meta's opt-out process is even worse). But are they madder than usual about DALL-E 3's results? Are they actually mad about DALL-E 3 because they think it's crossing some technical milestone?
The title here is an editorialization on behalf of the article writer that does not seem supported by the text or sources. I don't see strong evidence from artists that DALL-E 3 in specific is such a shocking leap forward for AI image generation that they are suddenly now more worried about this than they used to be. Again, I'm not fully up-to-date on everything, so maybe it is a leap forward. But the article doesn't make that case, it just kind of says that DALL-E 3 is amazing and implies that the current conversations are specifically the result of DALL-E 3 being just so good.
I've just seen that play out before; I'm skeptical of any article that takes a general social trend and says, "this trend exists because this specific product is revolutionary." Is there a quote from an artist in this article that I've missed that is saying that DALL-E 3's capabilities have made them more worried about AI generated images?
Last thing I've tried was a Distracted Boyfriend meme, but with people dressed like a cat, a cat tree, and a sofa (cardboard cosplay style). Maybe I simply don't know the right way to explain and make it write a perfect prompt, but it was insurmountable task for DALL-E. It always forgot something - either that there should be only three people, or that there's one cat-costumed person, or that a sofa is a costume and not a piece of furniture, or that I wanted to have the obvious canonical layout of the meme and so on. I just gave up.
With Stable Diffusion at least I can do this using iterative inpainting or ControlNet segmentation. DALL-E simply cannot iterate on existing images.
Overall, DALL-E feels like a dumbed-down version Stable Diffusion without repeatability (can't control the generation, so can't reuse the looking-good seed) and a bunch of extra hoops to jump through.
Also striking: if you average all our creative output and blend it together the result is something most individuals would have real trouble creating.
ChatGPT already doesn't allow you to copy the style of a specific artist.
> To evoke the style of Thomas Kinkade in a chat prompt, one might use descriptions like "radiant warmth," "soft-edged realism," and "gentle natural light." His paintings often conveyed a sense of bucolic perfection and peaceful retreats, featuring cozy cottages, flourishing gardens, and tranquil waters. The scenes are typically bathed in the golden light of dawn or the rosy hues of dusk, with a magical, almost ethereal quality to the light that seems to emanate from within the scene itself. The colors would be described as vibrant yet soft, with a harmonious palette that creates a nostalgic and dreamlike atmosphere.
Using the terms provided, without mentioning the artist, Dalle3 spit out:
https://i.imgur.com/10pr0A9.jpg
https://i.imgur.com/kG4Mapf.jpg
The two prompts I used were [note: I picked a scene / location]:
"generate an image for this prompt: "bucolic perfection, cozy cottages, flourishing gardens, radiant warmth, soft-edged realism, gentle natural light, bathed in rosy hues of dusk, vibrant soft colors"
"generate an image for this prompt: "bucolic perfection, new york city, autumn, flourishing trees, radiant warmth, soft-edged realism, gentle natural light, bathed in rosy hues of dusk, vibrant soft colors"
Seems to get you pretty close to the style of Kinkade with very little prompt effort.
It then failed spectacularly at recreating the style. There may be a way to craft a prompt that does recreate a specific artist's style, but having an LLM do it is not yet in the realm of possibility,
I then took those terms and formed a simple prompt -
generate an image using this prompt please: "tranquil scenic landscape, wet paint style, puffy clouds, serene lake, distant mountain, trees, harmonious, earthy, natural beauty, wilderness, calm gentle realism, peaceful, pristine"
It spit out these images:
https://i.imgur.com/P01OAIh.jpg
https://i.imgur.com/99O8S5J.jpg
Took a couple of minutes. With a very small amount of effort it'll get you close to the style of Bob Ross.
There are unique keywords it associates with Ross that it'll refuse to work with, apparently, "fluffy clouds" perhaps was one (it refused when I used that initially, along with "wet-on-wet paint style"). With a little adjustment you can likely fully get around that.
There's nothing extraordinary about the style of Bob Ross. He isn't famous for producing astounding artwork, he's famous for his speech tone (comforting), instruction (audience being able to quasi follow along at home) and that he produced peaceful natural scenes. There's nothing about his painting style that Dalle3 isn't going to be able to recreate to a high degree. How much time are you willing to invest into the prompting? If you're a big Bob Ross fan, I would think spending a small amount of time to get a highly effective prompt would be worth the ability to endlessly produce new artwork that reminds you of his style.
It's very clearly feasible at this point.
Take the clouds for instance. The closest matches are Waves of Wonder(S15E6), and Summer in the Mountain(S25E5), and that's a stretch. The vast majority of Bob's clouds don't look anything like that. They are mostly diffuse rather than the hard-edged creations DALLE3 generates.
I'd challenge you to find an example among Bob's paintings with which these generations share similar characteristics https://www.twoinchbrush.com/all-paintings?page=1
We haven't figured out micropayments to simple online articles, let alone how much an artist should be compensated for an output from an AI model that includes some weight where the source image was of statistical significance.
My point is that the only chance artists have of getting paid is to charge for the training dataset (the source, not the output). The issue is, a lot of them will post their work online (in worse quality, just as a demo), and that will still be perfectly fine to be in a training dataset. Who will say the AI can't be trained on that? My brain certainly can and nobody is charging me after I see some art online and decide to, say, paint it myself. Are there laws against copying style?
Artists will probably have to come together in a union of sorts to have any power. That has never happened at the scale required to impact AI vendors, I think.
Either that or government steps in and demand AI be trained on some kind of certified dataset with proof of origin, with penalties for using non-certified datasets. Maybe that will be the only way to have people compensated.
We need to advocate for open models. If your work contributed to a model, you should have the right to your own copy of the model.
Viewing a website and the multi-media files displayed on it is not inherently evil or bad and giving it the name "scraping" does not make it so. Let's not adopt this lawyer-ese framing. These people seem to want to have their cake and eat it too.
> Because a trained diffusion model can produce a copy of any of its Training Images—which could number in the billions—the diffusion model can be considered an alternative way of storing a copy of those images. In essence, it’s similar to having a directory on your computer of billions of JPEG image files.
Just like there is "no way" that a 3 Lb hink of gray cells can memorize epic poems or figure out how to design ships to fly to other celestial bodies.
IOW: It doesn't work that way.
The GPTs do not store representations, they store what are effectively generation instructions. A formula ir program can generate nearly unlimited terabytes of data, but we don't need to store the output to recreate it; we need only store the formula or code.
So the "only 3GB" argument does not mean that sufficient instructions have not been stored via ingesting an artist's works to effectively infringe on that artist's works without literally storing copies.
That's exactly what it means. It means no complete representations of the artists work are stored. The definition of fair use. Just like I'm not stealing anything when I go to an art museum and vague memories of some of the exhibits later influence a drawing I make.
That's... not the definition of fair use.
In the AI example, who is the ordinary human who viewed each of these works and is recalling them from memory?
Fair use is not reproducing the same painting in the same style, regenerating almost verbatim large blocks of code or text, etc.
But, while we can argue the boundaries of fair use all day, that is not the key point.
The key point is: the argument that "3GB is too small to hold anything that could qualify as fair use" does not hold water. Just because you or I don't know how it would work, does not not mean that it cannot work. And it has indeed been shown to work in recreating art, text, and code.
That OpenAI are now banning use of directives resembling "in the style of" is good evidence that they consider it a legitimate problem, although that is likely to be ultimately a speed-bump measure.
This doesn't exempt it from information theory. If a 3GiB model was trained on 5 billion images, it fundamentally cannot have more than ~5 bits of information per training image on average, regardless of whether those are instructions for its recreation or a compressed copy of it.
Over-representation of some images (at the cost of the under-representation of others) is an issue, but it certainly cannot "produce a copy of any of its Training Images—which could number in the billions" as is claimed in the lawsuit against Stable Diffusion.
> So the "only 3GB" argument does not mean that sufficient instructions have not been stored via ingesting an artist's works to effectively infringe on that artist's works
Copyright applies to fixed works in a tangible medium of expression - and explicitly not formulae, procedures, or instructions like recipes. I don't think there's a bright line here, but this wouldn't be the direction I'd argue in if I wanted to classify machine learning as infringement.
If you’re using AI art for anything other than personal curiosity you’re probably using it because it costs next to nothing to make a generated stock image for your blog or children’s e-book.
Having taste and knowing what you like doesn't mean you have to know how to draw or knowing the artist.
It's enough if you can use ai for it.
That's what people normally like on art: it's good. You don't see any issue and it speaks to you.
There is no difference in going to art galleries or clicking through the internetz.
It’s as if the embedded rules of photorealism are well-represented enough in the training data as to reliably cross the uncanny valley, but in illustration, where the internal logic of an image is much more self-contained, there’s still too little to go on for any individual style of representation.
New data poisoning tool lets artists fight back against generative AI (https://news.ycombinator.com/item?id=37990750) Oct 2023 - (85 comments)
Nightshade, a tool allowing artists to 'poison' AI models (https://news.ycombinator.com/item?id=38013151) Oct 2023 - (71 comments)
The AI revolution came too quickly for artists to guard themselves against it - to prevent their work from being stolen... yes, stolen. When the music industry took off, sparked by advances in radio and recording technologies, the pioneers of Blues, Jazz, and Rock and Roll were taken advantage of. They weren't compensated with royalties or profit sharing while studios sold millions of their records - still in print to this day. We know better now, yet artists are still being taken advantage of. Now, an entirely new market has been created from their work, with zero of the proceeds going back to artists.
Some good advice for young artists, "Always keep your publishing rights"
https://www.guitarworld.com/artists/steve-vai-frank-zappas-a...
I do like the character consistency, but in terms of quality, I think MidJourney is better.
No API
Are you saying there are parts that aren't documented? For example, the documented API has calls for generating images with DALL-E 2. Is there undocumented API for generating with DALL-E 3? One that uses pay-as-you-go pricing like DALL-E 2, rather than requiring a $20/month purchase of ChatGPT Plus?