Stable Diffusion Textual Inversion
github.com
github.com
I think the main issue here is the computational cost, as - if I understand correctly - you basically have to do training for each concept you want to learn. Are pretrained embeddings available anywhere for common words?
I'm sure they were looking forward to many months of maintaining highly exclusive access and playing "too dangerous to release" games before SD completely upended the table.
The problem is that they will automatically ban accounts that trigger the filter too much, so people would have to burn a whole lot of accounts to assemble an even remotely-complete list.
I'm guessing those teams didn't know in that they're AI researchers, and in my own employment at Google, I've been regularly reminded that being a technologist and being someone obsessed with a technology and pursuing it socially are different things.
Even without knowing the precise individuals that'd do it, I knew in February that by August there would be an open source model challenging state of the art back then, if only because given 6 months _some_ open source team would try scaling to a bigger model.
Another thing to point out is these teams are descendants of open source, Katherine Crawsons open source breakthroughs led to substantial improvements in DallE. Everyone should be saying her name 1000x more often.* She also helped create Stable Diffusion specifically, in substantial ways
* I think. Maybe I misunderstand the technology dramatically. But I think it's just poorly understood how much she's been involved.
Not only that, but OpenAI didn't seem to know their CLIP model could be used to generate images (via Advad's CLIP+VQGAN) at all, otherwise they wouldn't have released it. So they did unintentionally start the "AI art" movement even if they didn't release a trained DALLE.
CLIP isn't the true blocker to entry, the dataset and compute is.
So it seems there actually aren't many barriers to entry at all. There's certainly a lot of legal questions, but if it's this easy to create your own model then it's hard to enforce anything…
The UK recently announced plans to make this completely explicit, to remove any remaining doubt: "For text and data mining, we plan to introduce a new copyright and database exception which allows TDM for any purpose. Rights holders will still have safeguards to protect their content, including a requirement for lawful access."
https://www.gov.uk/government/consultations/artificial-intel...
BLOOM is out there, but not that many individuals with have like 8 3090 to host them, and the inference is still incredibly slow nevertheless
This idea of "technology gatekeeping" sickens me. I'm tired to death of people saying some non-sensical horseshit like, "The technology is too dangerous to be turned over to the hoi polloi!!"
Give me a break... as if someone running StableDiffusion on their home system and creating naked centaur-women out of pictures of Kate Beckinsale and anime waifus out of Ariana Grande photos are going to cause the downfall of the modern era.
StableDiffusion didn't upend the table... StableDiffusion gave the plans to the printing press to every person out there that wants to learn how to make their own print shop... and more power to them all, I say. I've had more fun and learned more about AI models in the past week than in I've had with AI in the past year, and I've been using img2img to feed my own art into SD to create whole new works that I've been able to touchup in Photoshop and upscale to print resolution.
This is truly the kind of computing revolution that I love to see, and that comes around all too infrequently. The good from this will far, far outweigh any negatives.
Hackers built all this technology. There's no way a handful of megacorps are going to take it all for themselves.
Pixel exists (but apparently doesn't count because it's not perfect yet).
Librem exists.
PinePhone exists.
More will exist in the future.
For hardware our world is not there yet and won't be for quite a forseeable future.
It's the difference between free knitting patterns and free cardigans.
I think that misrepresents OpenAI's attitude. As I see it, their claim is closer to "let's discuss whether the stable door should be closed before we find out the hard way what makes the horse bolt".
Given how much trouble we already get from the Gell-Mann amnesia effect, and how many people take spirits and horoscopes seriously, it seems entirely plausible to me that some highly realistic centaur picture could be used as a casus belli for a popular uprising that effectively ends a nation.
(Similar rumours abound even without this tech, c.f. Catherine the Great or Malleus Maleficarum etc.; I suspect arbitrary photorealistic pictures make that kind of drama much more likely to occur and to stick harder when it does, but this suspicion is not strongly held).
Edit:
I want to add that my concerns from tech are less about the general public (most people are basically decent), but from the few percent who hate or fear who now have a much easier time promoting their views (the possibility having always existed is different from it being cheap), and also from those who don't realise the images are generated to fit the text and instead think it's a search engine of existing images (which appears to be a common view judging by the type of complaint certain artists have on any given demonstration of the tech, though public figures complaining about Google search results without knowing they're personalised is also a thing even for actual search).
I just read (skimmed) through the paper.
That's in fact the key idea here: that the training model is untouched. Using the existing, trained model, they use this "inversion" procedure to discover some word that acts as a stable reference for a concept expressed in some images exposed to the model, which the model will understand as a reference to that concept.
There is a pretrained model with those common words, which knows how to do things like, say, "hamburger in the style of Picasso".
Now, without such model having been trained on the works of some artist, or other images, using a few samples (merely five or so), it's possible to uncover a latent word in the model which refers to the concept that those samples have in common. That word is stable in the sense that you can compose it with other words in prompts, and it really seems to denote the concept in those sample images.
In the paper these researches consistently call such a word as the meta-variable S*, and use it in prompts like, "flying monkey in the style of S*".
What I couldn't spot in the paper is a concrete example of what the S* word actually looks like for given examples. I'm guessing that it's some sort of gibberish. According to the concrete usage instructions, the process produces an embeddings.pt file, which you then upload, allowing you to use the pseudo-word * (asterisk) to refer to the concept.
People have been intuitively experimenting with gibberish words in prompts, discovering some stable behavior that seems to correspond to words that the AI "came up" with by itself (like a child, some have noted). This research seems like methodical way of discovering those internal words.
The basic SD model should have all the common words covered, this model's goal is to find a new concept that doesn't exist visually or textually in the dataset, like for example your own face, or a character you designed yourself. Note that this might not be possible to do, the corpus of data or the size of the model might not have held enough information that it can represent certain concepts, or at least represent them in detail. I.e. if you give it pictures of your dog, it might not look quite your dog during generation, even though those details existed in the pictures you gave the model.
If you want personalization that is also highly detailed, you'll have to fine tune the model itself with your own concepts, google has detailed how they did their own fine tuning and called it dreambooth[1].
On a more serious note, this opens up the door to exploring fixed points of txt2img->img2txt and img2txt -> txt2img. It may open the door to more model interpretability.
But for real, plenty of people are going to start rolling their own art and skipping the artist. Not Coca-Cola, but small to medium businesses doing a brochure or PowerPoint? Sure!
But it's not just for art and design, it has uses in brainstorming, planning, and just to visualise your ideas and extend your imagination. It's a bicycle for the mind. People will eat it up, old copyrights and jobs be damned. It's a cyborg moment when we extend our minds with AI and it feels great. By the end of the decade we'll have mature models for all modalities. We'll extend our minds in many ways, and applications will be countless. There's going to be a lot of work created around it.
It might now, but I feel like that will be trained out of it a few more papers down the line
There are plenty of small shops now where somebody knows a little Photoshop and can eek out a design that they otherwise wouldn't be able to using pen and paper.
There are also professionals that use the Adobe suite to enhance their abilities they've cultivated for years.
AI art will simply be a tool that enhances artists but might take away some low hanging fruit jobs similar to how web frameworks pushed people out of the job of webmaster and into more specific roles.
Those were already used in a "let's find something that roughly fits what I want to communicate with this text" way.
Today I created a quick get well soon card using an image from the new Midjourney beta and I have to say the result was exactly as good as if I had used Shutterstock but it took me much less time because the search prompt matched created something I wanted on the third try.
Comparing that to sifting through pages and pages of vaguely relevant images it's a clear win and a lot cheaper
Right now the tech still requires some nuance to be able to slap it all together into what I think most people would want.
While i expect the interface and the like to get a lot better, all good tutorials of this tech so far show many iterations over many different parts of an image to get something "cohesive". Blending those little mini iterations together is VASTLY easier than just making the whole thing, but not just plug and play for something professional.
Still there will be a huge dent in how long it takes to make certain styles of work and that will lower demand considerably, and there's a large market of artists who thrive on casual commissions which this might replace.
If anyone was to be worried I’d think it would be Getty Images.
Getty Images will just get with the program and stock up on bazillions of AI generated images, indexed by the prompts used to generate them.
Someone looking for stock images doesn't want to deal with artists, photographers or feeding prompts after prompt into some AI software, while not quite getting the desired result.
If Getty makes it easier for someone to find some existing AI-generated image than to generate one, they still have something.
A lot of the AI images we see in online blogs and galeries have been curated; people tinkered for hours with the stuff, and cherry-picked the best results. There could be some business model in that, at least for a while.
With stable diffusion, it really creates some nice stuff for what is already available, like a pizza with specific toppings [0]. I've been using it to add pictures to any of the recipes that are added to my wiki site without a picture. I originally tried this with DALL-E with similar prompts and the results are less appetizing.
[0] - https://www.reciped.io/recipes/mushroom-and-onion-pizza/
Piano keyboard: https://mobile.twitter.com/skybrian/status/15629121632697303...
Model train layout: https://mobile.twitter.com/skybrian/status/15629200301441064...
Turtle sand sculpture: https://mobile.twitter.com/skybrian/status/15640856320323174...
If if has any greater meaning, we might all be a little nervous that it’ll come for our jobs next, or some piece of them. First it came for the logo designers, but I was not a logo designer, and so on.
So people were hyped up.
Also replicate links you to github, but in case you cannot run (not have a GPU at home) you can just run some free queries on it. What's to hate about that?
I don't care what you think, You can also check my other post on this thread where I link to the official paper, and has 11 upvotes.
Go bot yourself
Oh, also, very funny to have a 2m old account giving me lectures about what I'm allowed to post or not, you can't even downvote...
This is is what I wanted to share originally https://replicate.com/methexis-inc/img2prompt
Have a nice day to you too, and sorry again for misreading the whole situation!
These models have hundreds of input parameters, not just prompts. There are many ways to configure the different models and various techniques and link up the processing stages.
Getting the best results in a short amount of time requires highly specialized knowledge. The job description isn't "prompt engineer" but something close to that.
Didn't take long for someone to resolve that issue!
It doesn't feel like this tech has hit any hard limits on things that are impossible or very far away yet. Every limitation seems to be getting broken at rapid pace.
It's like how humans create new words for new ideas, use AI to process a visual scene and generate a unique 'word' for it that can be used in future prompts.
What would a 'dictionary' of these variables enable? AI with it's own language with orders of magnitude more words. Will a language be created that interfaces between all these image generation systems? Feels like just he beginning here..
I'm experimenting with generative art with all this new models, and what's coming as output is wild and beautiful.
You can basically make a full AI experience now as a one man show world building with the capabilities only prior to marvel or disney...
This might be a better link> https://textual-inversion.github.io/
Also this reddit tutorial might be useful to you https://www.reddit.com/r/StableDiffusion/comments/wvzr7s/tut...
That link is better in my opinion.
Showing multiple variables, styles Sx in the style of Sy; it's really amazing.
Also, if models start to accept prompts like "in the style of Qinni", surely we're back to the copyright debate. They get away because everyone's art is mixed into a single model, but once distilling someone's artstyle is a feature…
This is not "work out what the prompt is for an image"
Instead it lets you give the model a "thing" (as an image) and then use that thing in your prompts.
So for example they give it a picture of a statue and name it "S", and then can say "Elmo sitting in the same pose as S" and it correctly generates it.
>The objective is similar, but it's: (1) A different approach - they also fine tune the model itself, and they get much much better identity preservation!
From Twitter : >Awesome job! That really extends the applicability of powerful generative models nowadays. Could I ask if you have any timetable for releasing the code please? >We are working on plans for implementation on other open source models
Seems like we have the same thing happening here. The input images are the sensory experiences. The dummy word "S*" is the linguistic symbol that is attached to them.
Imagine entire subreddits consisting of posts, comments, memes, and photos, and 100% of it is pro-[insert authoritarian regime] and it essentially only cost $1M to do it.
Filing charges is pointless, says Ezra. Since two years, she's being harassed on Telegram. It started when she was sixteen: photoshopped nudes with her snapchat account were circulated. They had taken selfies from her social media, and those of her family, and combined them with porn fragments. She doesn't know the perpetrator, but that person takes a lot of trouble to ruin her. "Nowadays, the boys have so many ways to make it look real." source: de Groene Amsterdammer,146/33, p. 21.
The only part that could be improved with money is getting dedicated mobile modems for unique IP addresses, to evade spam detection.
* Train network on thousands of assembly instructions.
* Prompt ‘some bad weapon of this size and material’
* Result simple instructions how to go build it.
Lots of smart people have been trying to get it to have capabilities anywhere close to what you’re describing for years now, to no luck.
We’re safe for now.
The type of generalization necessary to perform what the parent was talking about for instance (synthesizing schematics) is (currently) not possible.
You can invite the bot to your server via https://discord.com/api/oauth2/authorize?client_id=101337304...
Talk to it using the /draw Slash Command.
It's very much a quick weekend hack, so no guarantees whatsoever. Not sure how long I can afford the AWS g4dn instance, so get it while it's hot.
Oh and get your prompt ideas from https://lexica.art if you want good results.
PS: Anyone knows where to host reliable NVIDIA-equipped VMs at a reasonable price?
https://colab.research.google.com/github/pharmapsychotic/cli...
You could do storyboards from a shooting script* now, but generalizing to synthesizing character and camera movement as well as object physics is a ways off.
* A version of the script used mainly by director and cinematographer with details of each different angle to be used covering the scene.
* Do we get lost in it?
* Does today's 'professional' fiction become a lot less lucrative when we can create our own?
* Is there a to leverage this technology the improve the human condition somehow?
We will also adjust to AI generated art like have other creative technologies and the novelty will wear off. We will become good at identifying AI generated art and think of it as cheap.
Still, extremely exciting.
That is one hypothesis for the fermi paradox, Kardashev scale, and the great filter. At some point all civilizations essentially create infinite dreams/thoughts/Matrix style tech in where we all will retreat inward and have an infinite world to play with and essentially become gods in a virtual reality.
Tools like this will absolutely be used by professionals to cut out portions of the workload, but there's still a large gap between something like this and actually making a coherent, cohesive, consistent, paced, well framed and lit story from text alone.
I replaced a Biden meme with Shrek just yesterday using Stable Diffusion image2image in Colab.
I guess, I hope someone reads this and will pick up this. Maybe coupled with a VR set?
Netflix UI of 2033: I want an episode of Senfield with X, Y, Z... Starting brand new (never aired) just generated by the AI episode now.
github with web ui: https://github.com/hlky/stable-diffusion
dev repo (more features, may have bugs): https://github.com/hlky/stable-diffusion-webui
repo with docker: https://github.com/AbdBarho/stable-diffusion-webui-docker
colab repo (new): https://github.com/altryne/sd-webui-colab
can also run it in colab (includes img2img): https://colab.research.google.com/drive/1NfgqublyT_MWtR5Csmr
demo made with gradio: https://github.com/gradio-app/gradio
How were farriers able to ensure their economic status wasn't affected by cars?
How were switchboard operators able to ensure their economic status wasn't affected by automatic switching?
How were travel agents able to ensure their economic status wasn't affected by Expedia, Priceline, and the dozen other sites out there?
No one can know when or how their job is going to go extinct. You just need to pay attention to the changing winds of technology and adjust your ship's sail accordingly.
e.g. GPU compute cycle companies (e.g. nvidia/amd/intel/arm/etc..) or invest in cloud compute companies (Amazon/Microsoft/Google/etc..)
Or learn relevant skills and try to have relevant marketable skills. (Always be learning)
You cannot.
It sucks, but it's inevitable. The same thing has been happening ever since the industrial revolution. Old skills become obsolete, and nobody cares what happens to people who've invested their entire lives into something that is no longer profitable.