I Resurrected “Ugly Sonic” with Stable Diffusion Textual Inversion
minimaxir.com
minimaxir.com
IIUC, OpenAI specifically deny-listed trademarked characters from the training dataset to try and side-step getting sued by companies with enough resources to move the needle on what was allowed and disallowed in terms of digital art creation (since a ruling could disrupt their entire business model).
I assume Emad Mostaque did the risk calculus differently when they open-sourced Stable Diffusion (I don't know for sure, but it smells like his attitude on the question was "I don't care because now that it's open-source nobody can delete it from the Internet anyway").
I had similar issues with VQGAN + CLIP, where CLIP is also trained by OpenAI but otherwise handles copyrighted characters a bit better.
Even if someone had added a blue realistic-looking hedgehog to the dataset, if they referenced Sonic their image would have been deny-listed. So the only way to get there is to add the adjectives one is looking for, such as `a portrait of a blue hedgehog wearing red sneakers and eating a chili dog`... And, indeed, when I try that prompt, I definitely get some Sonic-alikes.
It's like a hedonic treadmill but for AI image generation. I assume this will have the most usage as a once in a while tool for artists to get inspired from, and also a tool for commercial usage such as in a Photoshop or Figma plugin for designers. The lay person who wants to generate images will get bored after a while.
And thought "I assume this will have the most usage as a once in a while tool to decorate my cave with funny animals".
But then drawing animals turned into writing. Turned into printing. Turned into emails and the web and the web 2.0 and now here we are, sending our thoughts across the world with the speed of light. Using software that we wrote to tell computers how to do this for us.
It might be similar with AI driven image generation. That it evolves in ways that are currently hard to foresee. Maybe one day, our thoughts are translated in realtime into images and send across the galaxy with the speed of light. Or faster?
I know the requirements for something like that are extraordinarily high, but I could see it in 30-40 years being viable.
I'm watching this space like a hawk. Though I code for a living it's too different from this work and I can't personally do the stuff I need with it, but that stuff is continually happening (if I wasn't based on M1 Mac I might be a little farther along, but I think I'd still be running into obstacles based on how technical all this is)
I've put out an album, pretty recently, based on my collaboration with a modular synthesizer running systems that let it aleatorically generate chords, key changes etc. and you probably would not know the 'machine' wrote the chords for the whole album.
The place where Stable Diffusion becomes interesting for me is when I can feed my own styles and objects into it. Right now, I don't think that's facilitated for me on my M1 Mac in my relatively nontechnical world: I can run DiffusionBee, and explore the ranges of the dataset's collective visual unconscious. I'm learning how to do much more interesting things than 'trending on Artstation, 8k, etc etc etc'.
I have curated collections of images to which I can apply my own language cues, and a 440 episode webcomic that I could annotate the hell out of (my OWN art, not currently even on the internet), and the ability to feed not just 4-5 images but dozens, hundreds of images into the machine.
At that point, it's a private visual imagination that becomes ME diffusion, and I don't have to tell it 'trending on artstation' anymore. Hell, I could teach my copy a whole set of associations based on just feeding its own output back into it, using my own intuition to associate not-generally-useful concepts like 'cold' or 'loud' or 'disappointed' considered as visual abstractions.
If I can reliably associate 'anticipatory' with visual stimuli, I can begin using it as direction for my own use of SD, telling it that I want this panel more 'anticipatory' or less. If I can feed in a language of panels and borders that has variety and I'm able to associate it with language, I can use SD to generate comic panels with the associations 'unsettling' or 'normal' or 'dramatic' and composite them into final output.
Bear in mind that engaging in this behavior means ME feeding in the associations that are relevant to ME as an artist, effectively making an auxiliary visual subconscious much like I made a modular synth into an auxiliary musical composition subconscious.
No, I'm not bored. If you're bored, maybe you're not an artist? Or maybe you don't have a firm intention and motivation towards which to direct your art?
These things should be like a violin. You can give one to any shmoe off the street, but ability to perform on the thing does not come along with just picking it up and plunking at it. I'm convinced that in order to make visual AI a tool you absolutely must let the artist feed in their own associations, concepts, objects etc. and then direct the output towards their own ends.
personally, I've been waiting a long time for AI to get to the point where I can generate my own animated cartoons. I can see the pipeline for it now: sketch concept art, mess around with inpainting, use textual inversion to build a "dictionary" of assets like character art, scenes, objects, art styles, use that + simple sketches to storyboard, then put the script in and get keyframes, then tween.
and even if the result looks mediocre or a bit weird (for now), the sheer power of solo-animating an entire series is just.. tantalizing.
IMHO, the results of these image generators also tend to be pretty mediocre. Not terrible, but like those Beeple NFT images: something made by someone of middling talent without inspiration, mainly as an excuse to use tools.
Also, when I played with stable diffusion specifically, the stuff it generated frequently had a horror-show quality, because it has no idea about stuff like how many legs people have.
I wonder if these generators will plateau, because the "throw more training data at it" technique will be undermined by mediocre-to-poor AI generated images.
We are in for some interesting times. Whatever the next iteration of Textual Inversion is will be extremely disruptive, especially if the concepts continue to be developed collectively.
I recently trained it on NFTs to generate variants of bored ape and punk style art.
The only problem right now is the variants are not consistent and it's hard to tell stable diffusion to make only slight variant changes with some mask and editing.
You can mask out the top head of the NFT punk and stablediffusion can generate different heads but that's fairly limited in the end result.
I think a cloud service which can automate the training and store textual inversion models would be a really cool startup. Ping me if you want to build this together.
I've been experimenting with it to spit out commercial style illustrations and stock photos. It's a lot of manual work with Google collab and frustrating to try.
I did test it though, which resulted in...this: https://twitter.com/minimaxir/status/1571520043737042946
This one charges but gives pretty good results at 512x512 - and in only a couple of seconds at that resolution. For more logo/cartoony stuff you can generally get a 10 step done in less than a second.
When I signed up a few weeks ago they gave 200 credits free (IIRC). After that it is $10 for another 1,000 credits (a credit gives 5 images at 10 step 512x512 or 1 image at 50 step 512x512).
Only thing to note is that after a while it slows down, but I realised that is because every image generated is going into an array in local storage in the browser. Using the browser's Inspect -> Application area to clear that array every 100 images or so sorted that out.
from diffusers import StableDiffusionPipeline
from torch import autocast
pipe = StableDiffusionPipeline.from_pretrained(
"CompVis/stable-diffusion-v1-4",
revision="fp16",
torch_dtype=torch.float16,
use_auth_token=True
)
pipe = pipe.to("cuda")
pipe.enable_attention_slicing()
prompt = "a photo of an astronaut riding a horse on mars"
with autocast("cuda"):
image = pipe(prompt).images[0]
[1] https://github.com/huggingface/diffusersThat fork released around when SD released, I've been wondering if anyone's integrated those optimizations into some of the more prominent forks that are floating around
You can do it in around 4 hours on Colab, using HuggingFace's notebook. https://colab.research.google.com/github/huggingface/noteboo...
Then you should be able to take the result (a single text token embedding) and use that locally.
(I didn't realise you were talking about textual inversion in your original post.)
It's a bit kludgy getting it working on a 12 GB gpu, I managed it but I'd rather a less hacky solution.
Works with M1 Macs too.
This understanding of plagiarism reminds me of what students tell me when I ask them why they thought they could paraphrase an uncited Wikipedia article as their paper, right before I report them to the dean's office.
I added that line because reproduction of the input dataset into an AI is a valid concern (such as reproducing a Getty Photos watermark), but that isn't happening here.
"Firefox, fast and free"
It feels like going to get a movie, and the title is "Casablanca, super attractive lead actor" or something. It just cheapens the value of the work like it's a low budget film from the adult section.
But what weirds me out is how much of that is needed to get a desired result. "Unreal engine 4k resolution".
What happens when 4k isn't enough? What about twenty years later when all our buzzwords are meaningless, and unreal engine no longer exists? Or what if AI appropriates the word, and that becomes the new definition?
The keyword "4k" doesn't "mean" anything to the model, it maps to the internal space through the language model depending on its training on captions earlier. You could just as well have used any other way of assigning some parts of this 7000-dimensional vector, it just turns out these keywords are usable shortcut, just like specifying a camera model or an artist name in essence acts as a "macro" to the internal space.
If 4k doesn't really "mean" anything to the algorithm, will it still have meaning for us? Or will our usage of the word start to reflect how the algorithm interprets it?
Just like YouTube face. It starts out as us influencing the algorithms, it ends with the algorithms influencing us.
There might be some interesting statistical analyses possible at the end of the day btw, like trying to figure out how many commonly widespread "artistic styles" there are in the world's total of art and photography so far. After all that is what a lot of the prompt engineering is about. Sort of a principal component analysis of the dominant 100 styles or something...
In a 4GB space is embedded potentially millions of images to tokens which correlate with language descriptions. So indicating the desired result somewhat simulates this process of recall and association.
The process lacks any internal emotional motivation to construct any "desired" outcome that a human would, so it needs both a "seed" number and description of the desired recall elements to achieve a specific result.
Or what if it appropriates your name as a style, and puts you out of business?
It's like the old prediction that well-known actors could sell their faces and voices for producers to CGI up new films with, except you don't pay the actor.