How a Stable Diffusion prompt changes its output for the style of 1500 artists
gorgeous.adityashankar.xyz
gorgeous.adityashankar.xyz
For example, I flipped back and forth between Beatrix Potter and Paulus Potter. A rounded white bonnet in one picture becomes a couple of blossoms in the other. The roof of a house becomes some shadowy wall with plants in the other. Two flower pots are very similar, just with slightly different coloring.
It makes it more apparent that the algorithm etches the images out of noise, and if the seed is the same for two images with different prompts, you're likely to see traces of that noise represented differently but recognizable in both images.
Thus the similarities you see would make sense if using also the same seed for the tests.
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
https://wcedmisten.fyi/project/paintingGuesser/
I tried to make something more general, but stable diffusion is fairly inconsistent in how well the output matches the semantics of the input.
Again, if the training data was labeled well enough, confusion about this sort of thing shouldn't happen.
That this happens at all is evidence that the training data hasn't been curated, cleaned, or labeled well enough.
In the reverse direction, you can try:
A horse rides an astronaut
And you will probably generate an astronaut riding a horse. It’s not a poor description of what we want; our assumptions about how grammar should work aren’t being honored.
I'm saying that because it's a little bit tedious to search for the artist you're looking for.
A part from that, I find the idea super-interesting :)
There is a lot of talk in our department as to how we might prepare our students for this technology. It is scary how fast it is growing, and how it is spreading to things like 3D and texturing.
One of my team is already using it in production. It used to take his artists three days to come up with five visual development ideas. Now he can get fifty overnight to choose from.
To my untrained eye, these both look pretty good.
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
https://uploads4.wikiart.org/images/joseph-beuys/how-to-expl...
The Audrey Beardly results are a bit better, but if I wanted Beardly-ish drawings I would likely be disappointed. Beardly's lines were very fine... 'filigree', like a spider crawling over the paper, and he almost never made art in colour. Also, a lot of his work was as sexy as hell. NSFW last link (bottom of page) for relatively mild examples.
https://uploads2.wikiart.org/images/aubrey-beardsley/the-dan...
https://www.theparisreview.org/blog/wp-content/uploads/2015/...
https://www.messynessychic.com/2019/11/13/the-world-wasnt-re...
We think that there may be some room for our students as 'high class' art directors. What will give them unique merit is their deep knowledge of pictorial formalities. Anyone can give the text hint 'flowers in a vase'. But what about...
'Move the camera down to avoid the strong coincidence line between the edge of the vase and the edge of the table. Change the saturation value of the vase to emphasize background/foreground contrast. Increase the amount of negative spaces around the periphery of the flower mass' etc.
Tek like Stable Diffuse may also lead to a resurgence of interest in natural media, like oil paint, water colour and suchlike.
I know someone -- a completely unknown artist -- who used to make a fair portion of their living by drawing D&D characters for people. Unfortunately, orders slowed down, because someone can input one of his images into software like this, and generate endless variations in the same style. Should this be allowed?
Are images created "in the style of" a certain artist completely dependent on images created by that artist? If so, should that artist be compensated? Why or why not?
Even if there are laws against it, the cat's out of the bag.
There's no stopping billions of people all over the world making derivative works at the push of a button.
Yeah, just wait until Disney characters get copied and mixed.
I wouldn't be surprised if their lawyers are preparing to change copyright law ... again ...
A derivative work is an adaptation, translation, or modification of a particular, existing copyrighted work.
If you asked Stable Diffusion for "Vincent van Gogh's Starry Night with a cat looking at the sky", you'd get a derivative work (although Starry Night is in the public domain, so you wouldn't be violating its copyright).
The results vary wildly, even run by run, but I might put them in a few buckets:
1. Similar enough that someone who doesn’t know much about art could be fooled
2. Amateur knockoff but recognizable style
3. Influence is there if you know what to look for
4. Artist probably not in the training data at all
The last one kinda surprised me, for artists whose work is online and who have unusual names. I would have thought those cases would be really good. Maybe they ran out of disk space with all the porn?
Also interesting that it gets much closer for figurative painters than for abstract painters.
It was given a lot of tagged data: 600 million captioned images from LAION-5B. So if you want to know what it might support, you could try any one of the captions from those 600 million images.
The vocabulary is here: https://huggingface.co/openai/clip-vit-base-patch32/resolve/...
It is contextual though, so words in different orders mean different things.
Is there any reason it couldn't be character based, besides the (presumably very large) increase in resources needed to train and run inference? This is all way out of my league, but seems like you could get interesting results from this, since (by my caveman understanding) this hypothetical transformer could make some sense of words it had never seen before, so spelling variants or neologisms and such.
It's actually a byte-pair encoded (BPE is better than character encoding but can do the things you mentioned) list of things that includes words. You can find common English suffixes in it listed separately too.
GPT (and many other modern NLP models) use byte-pair encoding. Your summary of the benefits of this is correct - it can deal novel words much better.
Byte-pair encoding (BPE) is better than character encoding because it can deal with unicode (and emojis).
CLIP uses a BPE encoding of the vocabulary: The transformer operates on a lower-cased byte pair encoding (BPE) representation of the text with a 49,152 vocab size
So strictly this vocabulary is NOT (just) words, it is common sequences of byte pairs. You can see this if you examine the vocabulary - you'll find things like "tive" which isn't a word but is a very common English suffix.
The main difference is that coming up with a word SD doesn't know that's not contrived is really difficult. In an IF game, you are constantly guessing the correct word.
https://github.com/rom1504/img2dataset/blob/main/dataset_exa...
You probably would want to stop after getting the metadata, unless you have 240TB available for the images :)
More details and links to dataset explorers here: https://laion.ai/blog/laion-5b/
It's a bit less than 50k words, but that includes space-padded duplicates.
So in practice asking for art "in the style of <x>" is sort of limiting the denoiser to statistical pathways resembling other images captioned "in the style of <x>". At least, that's my understanding. Still trying to grok ML and diffusion models.
1) Retro, which is essentially attention over large databases, and fast as hell.
2) S4 Layers, explicitly designed for handling long dependencies.
These are orthogonal approaches to memory, and both very effective at what they do.
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
Having the art vocabulary down as well.
In effect, knowing what is present and how it’s tagged so you can « invoke » it more readily in the prompt-result.
Maybe I’m out of my depth. I know the corpus of tagged image used for training is enormous … but I still think that would help the user ( a prompt-crafter )
I will have to check myself.
I tryed with minor success to have stable diffusion draw lesser know non-us personality. It always kind of work, but the palette is limited. For instance I tried charles de gaule and you get something that look like him. But he's depicted talking on a radio or waving his arm around like a politician.
I tried to make him to grocery or play volley ball, it does not really work. While Michael jackson or Dennis Rodman get a way better treatment.
edit : that vocab file is smaller than I thought. "De Gaule" is not in there but neither Einstein or feiman? I think I missing something.
I Can find "obama", trump or Macron. Unclear about Michael Jackson. No beyonce or Denis Rodman. hmm weird, I had great result with all of them. Like.. recognizable details like tatoo or silly glasses.
It's just a very advanced madlibs engine based on a database of a billion of already known images.
If anything, this sort of "AI" only makes human ineffable creative experiences even more valuable, because without it there would be no training set and no "AI".
Do you have a source for that? Somehow I doubt there would be this much interest in a tool that literally just copies existing artwork.
Can't help feeling that this accidentally harms creative types and risks swamping us with visual junk.
The technical achievment is astounding but no-one would seriously claim that crafting an image via a short prompt is creative except in the most cursory way.
I'm probably missing some life changing use-case, but apeing art in random styles can't be it.
The post-art world is here! Just think about how history books will remember this period! The styles that will be borne of necessity, of the need to break down art and find what makes it tick.
Besides having an ai doing the legwork is no much different than Veronese giving large swath of paintings to his novices while focusing on the two/three major parts.
Did they? Because I don't think they did. I think most people were amazed by all these technologies.
Research beats idle speculation.
Google "photography skeptics early history"
> As long as “invention and feeling constitute essential qualities in a work of Art,” the writer argued, “Photography can never assume a higher rank than engraving.”
Ha, dissing engraving at the same time as photography.
Though I wonder if the 'write' meant engraving of a design someone else already produced, or any engraving work at all.
Others are interpreting my original comment as "this is not art", but I'm not really trying to make that argument. Art is entirely subjective and i don't presume to define what is or isn't art.
I guess my point is more specifically "what itch does this scratch"?
It's really cool, and that may well be the answer tbh.
That’s 99% of the people. I mean even to learn prompt engineering will probably be too much for majority of those 99% people, but it’s a huge step forward in user friendliness, compared to, say, photoshop.
”what itch does this scratch"?
How many people post pictures on social media? Many of those pictures are not personal, they show something pretty, cool, or interesting in some way. All of those people can potentially use image generators to achieve the same effect.
These hot takes generally tell you more about the opiner (or the audience they're playing to) than the reality to come. It turns out it's hard to model en entire universe using 3 pounds of meat.
[1] Heinlein listed some of them way back in 1952: https://archive.org/details/galaxymagazine-1952-02/page/n19/...
Also about desktop publishing.
Remember all the printers (ie. people working in the printing industry operating printing machines) that were put out of a job when you could just buy a (electronic) printer for your home computer and just print whatever you wanted yourself?
People were wringing their hands about that too back then... now we take it for granted that we can instantly print whatever we want whenever we want, without having to pay an expensive professional to do it for us (something most people couldn't afford).
Has it resulted in more junk being printed? Absolutely. But it also let people print all sorts of fantastic not to mention useful things that would almost never have seen the light of day without cheap and easy access to home printers.
The xerox copier was similarly revolutionary... as was the printing press itself, which put a lot of scribes out of business.
Photoshop put a lot of airbrush artists out of business, and who does copy and paste with physical glue and paper anymore?
As with photography, printers, copiers and photoshop, artists who embrace this technology will be able to use it to enhance their creativity and speed up their creative process.
There'll be a lot more competition, a lot more junk but also a lot more fantastic art that we can't even dream of yet.
When I was finishing high school in the early 2000s, we still had teachers who made worksheets that way.
I remember one history teacher in particular. She used a photocopier to get sections from books, cut-and-paste them together, and then use the photocopier again to make the final sheets to distribute to the students.
I was very surprised at the time, but also admired the ingenuity. The process is much more physical than using a PC.
I could absolutely see an AI model doing the job of an entire film crew. I have issues with this, but only with respect to the longer term aggregate affects on culture in the broad sense. I cannot honestly believe that much would be lost from the perspective of one project or another.
Most artists spend their lives not refining their brush stroke, but rather their eyes. The way I see it, the impact of curation and artistic direction will matter more and more in the future.
For me it's exciting to use as placeholder art and then have a 'real' artist review it.
I'm biased: I've been working on an image generation app. But the beta users I've had so far will generate fifty or a hundred images in a day. That isn't a use case traditional artists support.
It will; I think the reason we're seeing diffusion models applied to image generation first, is that it's a task that meshes well with the models. But also in general I think people will still be guided by the principle "use the right tool for the job" - this is just another tool. I doubt that the set of paths toward realization for any given needed creative imagery collapses to just "use a model"
I hate arguments like this. Even ignoring how dismisive it is of the achievement at hand, why would you assume ingenuity is transferable like that? Someone who makes a breakthrough in physics is by no means likely to have made an equivalently ground breaking advance in biology if they had decided to study that field instead.
I guess my musing was hypothetical but I was careless in communicating that. I get that we can't centrally plan innovation or human effort - and I certainly wouldn't want to live in a society where this was the case.
Overall: the impression is better with author's popularity. I think if we train the model only with well-tagged filtered dataset - results may be much better, but we effectively will get a 5-year old Prizm app.
I've seen paintings in the style of Donato Giancola that really looked like his style, but in the examples of this site none of the result do. Maybe there needs to be some adjustments to the prompts?
Humans also have a hilariously hard time drawing bicycles, but at least we pretty much always nail the number of appendages.
And even so, too symmetrical faces will look just as un-human as a face that is too asymmetric. You need a face that is just the right amount of symmetrical in order for it to actually look good.
I think you make it sounds simpler than what it is.
It's also not a model that is trained to make as realistic people as possible, it's trained on a lot of different things, so obviously it won't excel at making realistic people. But one can easily imagine that some future models will be heavily trained on making realistic people rather than semi-realistic everything, like Stable Diffusion is trained to do today.
The system used here is actually astoundingly good at producing many artists stylesbecause it's not going for symmetry.
If it doesn't work for such famous painters, hard to trust it for the other ones.
Vincent van Gogh:
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
Paul Gauguin:
https://static.adityashankar.xyz/gorgeous/hair_flowers_full/...
Did some spot checking with some of my favorite artists. Rockwell's paintings are all about storytelling, clearly not present in the work. Their emulation for frazetta doesn't look like frazetta's work at all. HR Giger emulation is a joke. David finch at least gets a penciler's style but misses the use of solid blacks and dynamic posing. Frank Miller doesn't look like miller's work at all. etc etc etc
This list goes on and on. Personally while I understand using the 'in the style of' as a way to change the image results, I think in many many cases the results just don't look like the art of that artist.
Just in case anyone needed to see this spelled out.
The people making fliers have been replaced by AI prompting overnight.
The people doing contemporary fine art with their audience are unaffected.
Hmm, no? Do go and try to make a flyer with any of the AI generators. I’m not saying it can’t happen one day, but the current tech is not there.
That group will never be selling their own work as contemporary fine art but want that prestige, and the one way they had to make table scraps with that skillset is now gone.
A different person is doing their own flyer art with AI and adding words around it themselves, as evidence by my last months worth of fliers that have reached me. Promotion companies have always been up on trendy tools for differentiation.
SD isn't outputting clean graphic designs with sensible content
And haven't there already been people winning contemporary fine art contests with SD? lol
https://www.reddit.com/r/StableDiffusion/comments/x2n0r1/aig...
I haven't seen good examples of that yet but I'm curious how push-button you can make this. Flyers, web design and UI design require the copy, layout, information hierarchy, colours, illustrations, branding etc. to be cohesive so it's a different problem space with way more constraints compared to generating a single image.
If getting the final design requires a lot of rounds of prompting and tweaks, busy people are going to outsource this still (in the hope the prompting and feedback needed to the person doing the work will be less).
It's not a particularly well hidden secret that the contemporary fine art is really not about the art or the artist.
Angel Adams, for instance, clearly wasn’t too present in the dataset.
Edward Hopper was pretty impressive.