1 week of Stable Diffusion
multimodal.art
multimodal.art
What we're seeing now are toy-era seeds for what's possible - e.g. I've been making a completely Midjourney-generated "interactive" film called SALT: https://twitter.com/SALT_VERSE/status/1536799731774537733
That would have been completely impossible just a few months ago. Incredibly exciting to think what we'll be able to do just one year from now..
Still, humans use art to communicate intent, and we still consider AIs to be 'things' , no agency or intent. Being an artist just became a lot harder, because no amount of technical prowess can make you stand out. It s all about the narrative now
For more practical purposes like product design, anyone will tell you that actually drawing stuff is akin to typing in code, it can take a while but it's not the hard part
Do we, though? You and I do, sure. Most people here will, probably. But at least one counter-example was on display a few weeks ago, the guy from Google that told the press that their text completion engine was "alive" and "had agency".
From my friends I talked about this (which are not in IT), most believed him. YMMV, but I seriously believe a good chunk of the population thinks we already have thinking A(G)I. I don't think there is a "we" here, anymore. :/
I don't see the correlation, to be honest. A good chunk (but still a minority) of the people believing something doesn't automatically change the law, anyway, does it? :/
At the current stage i don't think there is any AI that can be punished, or anyone that would credibly claim that an AI must be punished. Its maker will always be punished instead.
Well true but i think that chunk is quite small. It's one thing to nonchalantly say "this is alive" and a very different thing when you have to deal with the consequences.
Yeah, fully agreed! Chatbots and "creative" ML systems are in the weird spot where they can't physically kill or hurt people, like e.g. a self-driving car, and perform tasks that "feel" like they need intelligence.
It's also absolutely quite possible that the "chunk" is way smaller than I think, I'm just blindly extrapolating from my social bubble :D
Other than "you can't trust anything you don't see with your own eyes", what kind of shift is it? People lived like that for literally millennia before photography, audio, and video recording.
At absolute worst, we are only undoing about 150 years of development, and only "kind of" and only in certain scenarios.
Moreover, people were making convincing edits of street signs, etc. literally 20 years ago using just Photoshop. What does this really change at a fundamental level? Okay, so you can mimic voices and generate video, rather than just static images. But people have been making hoax recordings and videos for longer than we've had computers.
I think the effects of this stuff will be: 1) making it easier/cheaper to create certain forms of art and entertainment media, 2) making it easier to create hoaxes, and 3) we will eventually need to contend with challenges to IP law. That's about it. I think it will create a lot of value for a lot of people (sibling comment makes a good point about this being equivalent to CGI), but I don't see the big societal shift you're claiming that this is.
You needed to have tools, skill, resources and time to do such things. You don't need to have that anymore. Anyone can do anything on any scale.
It's something what SpaceX did. Ofc it was possible to launch a rocket before SpaceX, but few really could afford that. Now that prices are low, that opens infinite number of new space exploration possibilities.
I have yet to see an example of "Synthetic AI media" that was both realistic and not immediately recognizable as being synthetically generated.
And if you think being 99% there means we're very close to 100% just remember how long it's taken self driving cars to close the gap (we actually don't know how long since they still haven't succeeded in this).
Would you know if you had?
Man, am I a good photographer or what
I mean, probably if you're familiar enough with squirrels something gives it away, but I'm not.
https://imgur.com/mcJsg0n and https://imgur.com/41eENUO
It also took a number of much worse images to get those ones.
In case anyone is worried, it isn't remotely erotic or otherwise nsfw.
I'm not particularly familiar with squirrels and something about that "photo" looks very off. If you showed it to me in a vacuum I'd just assume someone was trying to make a highly stylized version of something they had a photo reference for, but under no circumstances would I believe that's a real photo.
Three months ago I'd probably have agreed with you. Things have changed.
Downloading a cracked copy of Photoshop and checking out a book from the library on how to edit photos is only somewhat more difficult than learning to use Python and write programs that generate art from some model. And only because learning anything is extremely easy today with so many free resources and help forums.
> Anyone can do anything on any scale.
I'll believe it when I see it.
> Now that prices are low, that opens infinite number of new space exploration possibilities.
Except SpaceX is still in the "crawl" phase of "crawl, walk, run", and they only got even that far because because an eccentric billionaire has staked his reputation on the problem and thrown a huge amount of money at it, without having to worry about things like "reporting to Congress" and "making sure the space program creates jobs in such-and-such voting district". And after all that effort and truly astounding engineering (the rocket lands itself back on the launch pad!!), space launches are still expensive, risky, and complicated, and will remain so into the foreseeable future (~decades).
That's what Bezos did with its Blue Origin. If you check where is Blue Origin in space race, you'll quickly realize it's not enough
Edit: ah, and is space still expensive? If one of universities in my middle-sized country with no space engineering background could afford to launch cubesat via SpaceX, then yes, I think it became cheap.
Most people using those models aren't writing Python code. Check out https://www.reddit.com/r/dalle2/, https://www.reddit.com/r/midjourney/, https://www.reddit.com/r/StableDiffusion/, https://www.reddit.com/r/bigsleep/ etc
I expect that once the technology matures, a smaller and smaller niche of users will be doing any kind of programming
How much practice would one need after doing that, before they're able to match the quality of some of the AI generated art? Not all of the AI generated artwork is perfect, but some of the art would take the average person years of practice to be able to match. Some art requires more than a cracked copy of Photoshop and a weekend of reading a book you borrowed from a library. You may be surprised to find that some people spend years honing their craft.
In practice, "infinite" translates mainly to a handful of hyper-competitive guys checking off "went to space" off their achievements list. There is a very good reason for that: space is very inhospitable, much more inhospitable than Antarctica. Nothing much has happened in Antarctica for 100 years beyond the occasional hyper-competitive athlete and a few research stations. Perhaps a natural resource gold rush might liven up the place for a few decades, until exhaustion and falling back to inhospitable status, dotted with the rare ghost town remains.
Something similar happens in the "creative" space: the Internet unleashed a massive tidal wave of "content", yet the vast vast majority of it is rather trite and devoid of any (spiritual) meaning. Personally, I'm much more inclined to stick with the classics than even 20 years ago, simply because it's not worth my time wading through the deluge of poor quality "content" out there. To wrap up the analogy, I'd rather inhabit a nutritionally rich environment, than getting lost in the the vast, but mostly empty, expanse of the Internet.
I would take today's internet-fuelled media landscape over the landscape of 20 years ago in a heart beat.
The question I often ask myself: is spending time with this content, while entertaining in the short term, perhaps via the novelty factor, also nourishing in the long term? The answer is, sadly, much more frequently NO than in the time of printed books.
The best I can hope is to be able to use Internet as an encyclopedia for laser-focused lookups. Sadly, I am too often caught in browsing random content only loosely related to the original lookup topic.
A lot of the “trite” internet creations have gone on to become absolutely massive songs or artists.
I think we're making the folly of comparing AI generated art to human generated art from the 90s. Humans have "advanced" much further with the advancements now that DALL-E is nowhere close to.
It just feels to me like the internet has arrived in a way that can best be expressed in an Adam Curtis documentary.
This isn't just drawing. It's the start of telling a computer you want something and it making it. Not just pictures, anything. Phyiscal things, computational things,...
Hey computer make me a sandwich, a desk, a computer virus, a paperclip, a gun, a bomb, a little brother.
Consider the gaming world’s concept of “whales”. Customers willing to spend disproportionately enormous amounts of money in game. Can you sell these whales a unique, personalized David Bowie album that is about, I don’t know, maybe the customer’s own life story?
"This episode of Law & Order, but if Jerry Orbach never left the show"
"Final Fantasy VII as an FPS taking place in the Call of Duty universe"
"A 3D printable part that will enable automatic firing mode for {a given firearm}"
I think what we have is a toy and will remain a toy, just like Eliza was 60 years ago. Academically fascinating, and given the constraints of the era, genuinely remarkable, but still a long way from really being useful.
I'm already getting bored of seeing 95% amazing 5% wtf AI generated images, I can't fathom how anyone else remains excited about this stuff so long. My slack is filled with impressive-but-not-quite-right images of all sorts of outrageous scenarios.
But that's the catch. These diffusion models are stuck creating wacky or surreal images because those contexts are essential allowing you to easily ignore how much these generates miss the mark.
Synthetic AI media won't even been as disruptive as photoshop, let alone the creation of written language.
Personally I think a line can now be drawn that starts at the first cave drawing and ends in 2022. Something has fundamentally shifted, a true paradigm shift before our eyes.
I stumbled over Midjourney the other day through these music videos[1][2] generated by Midjourney from the songs lyrics, and I immediately thought we're not far away from this being viable for a cartoon-like film.
Interesting times ahead.
[1] https://www.reddit.com/r/midjourney/comments/x0kv8s/testp_ju...
[2] https://www.reddit.com/r/midjourney/comments/wz1am0/homer_si...
[3] https://www.reddit.com/r/midjourney/comments/x10som/the_amou...
[4] https://www.reddit.com/r/midjourney/comments/x12nqz/robert_d...
Also can run an 8 bit quantized version pretty easily. This takes ~6gb RAM.
The results seem far off from GPT-3 but apparently it can get good results when fine tuned.
Bigger models like OPT 66B can run on cloud machines (or a really big local system)
OPT 175B weights are not open but can be applied for.
175B would require something like 500GB RAM if not quantized. That's a lot, but it's possible to build that locally if you have a couple 10's of thousands of dollars.
Wait a few years and 175B on a GPU will be no problem.
Apparently it doesn't affect the results significantly.
More info:
You can invite the bot to your server via https://discord.com/api/oauth2/authorize?client_id=101337304...
Talk to it using the /draw Slash Command.
It's very much a quick weekend hack, so no guarantees whatsoever. Not sure how long I can afford the AWS g4dn instance, so get it while it's hot.
Oh and get your prompt ideas from https://lexica.art if you want good results.
PS: Anyone knows where to host reliable NVIDIA-equipped VMs at a reasonable price?
The next thing I considered was just buying up a ton of 3060 12gb cards (saw a few new ones for $330) and just hosting a server from my house. This might be a good option if you don't care about speed but care about throughput.
RTX 3090s are also decent in terms of price per iteration of Stable Diffusion. If you want to build a fast service like Dreamstudio I think it's the only option to be able to do it at a reasonable price. If you want to host these in the cloud using consumer RTX cards, you'll have to go with less reputable hosts since Nvidia doesn't allow it. I don't want to name any since I can't vouch for them, but there are some if you search. The cheapest option will be to buy them and host it yourself.
I'm still researching what the best price/performance is for hosting this so if you have any findings please share.
I‘m not really affording this, to be honest — I’m looking forward to switch to a Spot instance tomorrow, which could bring costs down to about $0,20 per hour, but even then I will have to switch it off in a couple of days.
I‘m working on a significant speed improvement — if that works out, and users get a result in under 1 minute if they are first in line, then maybe it‘s possible to make the bot finance itself through a credits system.
I quickly polished things and created a useful README - hopefully it's all correct. If not, let me know!
Plus, there are huge gaps in training. Ask it to draw something simple, like "a penis" and you get nightmare fuel....
https://www.vice.com/en/article/xgygy4/stable-diffusion-stab...
Why'd they "overlook" it? Probably more culturally significant and controversial than any of the others. It's the natural elephant.
"60% of all image generation compute power used for making NSFW material"
I can't yet decide if it's going to be extremely appealing or quickly get [even more] boring and repetitive.
Is this good? maybe... maybe not. Since most of the "normal" stuff already exists, it'll either be something "too extreme" for classic porn studios, or stuff using non porn people to turn into ai-porn stars.
I used to watch porn basically daily. But then after finally deciding to stop watching porn, the idea of porn itself is downright off putting to me. I don’t even quite know how or why. It just is.
And I imagine it will be the same for others with AI generated porn.
The internet tells me that a 2022 honda civic takes about 0.07 liters of gas per km. And it also tells me that that is equivalent to 2394 kilowatt seconds. I.e. 2394 images on a 3 year old GPU per km travelled using a new and fuel efficient model of car...
I'm not worried about this consuming a significant fraction of the power on earth.
Are people
a) Waiting 5 seconds/frame * 60 fps = 5 minutes / second for a video to generate on their personal computer, and doing this constantly enough that it manages to become a problem?
b) Buying computers that can do it real-time, but therefore output vastly more heat, requiring thermal management system akin to a car driving at highway speed?
c) Renting these computers at considerable cost to make these videos?
As long as enough people watch each video (or one person watches it enough times), the energy usage washes out to become negligible compared to the amount of human time invested. I just can't see a world where enough people are managing to consume a kw minute/second producing videos for themselves to watch only once or twice that it becomes an issue.
Personally I'm optimistic that energy/compute is going to continue going down substantially (in which case even real-time video generation might not be an issue). If it doesn't and we don't become substantially better at efficiently synthesizing video, I can't see personalized single use video generation being a thing.
There will be much more compute resources thrown at it to make it render in real time. We're not there yet, but I can see a path to that happening in the next few years.
> b) Buying computers that can do it real-time, but therefore output vastly more heat, requiring thermal management system akin to a car driving at highway speed?
Why not? We already have billions of cars driving around outputting heat. Its an incredible expenditure of energy, sure, but perhaps the value of generated content entertainment will match the value of car transportation.
> c) Renting these computers at considerable cost to make these videos?
I imagine longer term, the opex (i.e energy costs) will dominate the capex (GPU HW). The price of going into a generated world could be similar to going for a drive.
> As long as enough people watch each video (or one person watches it enough times), the energy usage washes out to become negligible compared to the amount of human time invested. I just can't see a world where enough people are managing to consume a kw minute/second producing videos for themselves to watch only once or twice that it becomes an issue.
This is where I strongly disagree. The democratization of skills and tools in creating content will break the one to many media model. You saw this in a large way in what the internet did to content distribution, in how the number of independent people creating content skyrocketed. These models will do the same for content creation. I predict most people will consume content personally generated for themselves or in small groups.
Here's an example: a group of friends puts on their VR headsets for their weekly DnD session. The DM begins describing the scene, which autogenerates around them. Each character can then respond with their own actions / path, and the scenes react dynamically. The hour session costs them $10 in compute/energy.
I'm mostly spitballing. I would imagine that we still have a couple of orders of magnitude reduction in energy costs that can be squeezed out of these models with improvements in specialized HW. But it will be matched against the insatiable demand of consumers for richer interactivity in content.
Or 80% of all NSFW viewing happens at work.
It appears to be a site for AI art, so there's that.
Additionally it is important to note that model was licensed under the OpenRAIL-M LICENSE which is not as permissive as an MIT license and forbids certain outputs to be shared or purposes to be built as apps
> You are accountable for the Output you generate and its subsequent uses. No use of the output can contravene any provision as stated in the License.
Unless I'm misunderstanding you, yes there is, and it was even posted on HN last week:
https://news.ycombinator.com/item?id=32572770
(And yes, many of its results are horrifying)
as far as I am aware, I may have the only semi-tuned nsfw model, though it's not all that great. There are a lot of concepts that SD needs to learn, and they're difficult to teach. It's very likely that by the time I get something usable, it's going to destroy all of the other concepts. @<another_user> also has a fine tuned model, but it's only tuned on generating closeups of female genitals. If that's all you want to generate, then he's produced some really good results. As for distribution, I wouldn't know where to begin, considering the model is 11GB
Some years ago, the pendulum was very much on the other side.
I took a quick look at the subreddit before it was banned and I don't think I saw any real people represented. It was a lot of video game or anime style characters. And one of Shrek with a massive dong.
(I got suspended just for saying "fuck the king of Thailand" - in reference to a politically abused law prohibiting insults to the monarch.)
Obviously if they're made to look like some celebrity, that's problematic.
That’s because you can use it to make porn. Don’t underestimate the motivational power of being able to easily create porn.
The right answer, I'd argue, is that this was Prometheus giving fire to the mortals, and then the mortals quickly discovered everything that could be possible with fire.
Mostly I'm interested in the processing time. Like, using a midrange desktop, what's the average time to expect SD to produce an image from a prompt? Minutes/Tens of minutes/Hours?
It runs pretty well but the most I can get is a 768x512 image, but it's pretty good for stuff like visual novel background art[0] and similar things.
[0] - https://twitter.com/xMorgawr/status/1564271156462440448
(I say “I think” because I’ve uninstalled the nvidia-dkms package again while I’m not using it because having a functional NVIDIA dual-GPU system in Linux is apparently too annoying: Alacritty takes a few seconds to start because it blocks on spinning up the dGPU for a bit for some reason even though it doesn’t use it, wake from sleep takes five or ten seconds instead of under one second, Firefox glyph and icon caches for individual windows occasionally (mostly on wake) get blatted (that’s actually mildly concerning, though so long as the memory corruption is only in GPU memory it’s probably OK), and if the nvidia modules are loaded at boot time Sway requires --unsupported-gpu and my backlight brightness keys break because the device changes in the /sys tree and I end up with an 0644 root:root brightness file instead of the usual 0664 root:video, and I can’t be bothered figuring it out or arranging a setuid wrapper or whatever. Yeah, now I’m remembering why I would have preferred a single-GPU laptop, to say nothing of the added expense of a major component that had gone completely unused until this week. But no one sells what I wanted without a dedicated GPU for some reason.)
Keep in mind to have the batch-size low (equal to 1, probably), that was my main issue when I first installed this.
Then, there's lot's of great forks already which add an interactive repl or web ui [0][1]. They also run with half-precision which saves a few bytes. Additionally, they optionally integrate with upscaling neural networks, which means you can generate 512x512 images with stable diffusion and then scale them up to 1024x1024 easily. Moreover, they optionally integrate with face-fixing neural networks, which can also drastically improve the quality of images.
There's also this ultra-optimized repo, but it's a fair bit slower [2].
[0]: https://github.com/lstein/stable-diffusion
With an RTX 3060, your average image generation time is going to be around 7-11 seconds if I recall correctly. This swings wildly based on how you adjust different settings, but I doubt you'll ever require more than 70 seconds to generate an image.
I quickly threw together a folder structure where I have a md5'd prompt as a folder name, into that goes _promp.txt with the actual text of the prompt and the images i generate in a loop with the seed used and iteration number in the image's file name. That way I can generate like 20 seed-based images for a prompt and if the model bites I let it run with a much higher number of seed-based images. When you have a 1000 to pick from, some of the results are freaking amazing.
I was able to obtain 256x512 images with this card using the standard model, but ran into OOM issues.
I don't mind waiting, so now I am using the "fast" repo:
https://github.com/basujindal/stable-diffusion
With this, it takes 30s to generate a 768x512 image (any larger and I am experiencing OOM issues again). I think you should expect a bit faster at the same resolution with your 3060 because it's a faster card with the same amount of memory.
My first impression is it seems a lot more useful then DALL-E, because you can quickly iterate on prompts, and also generate many batches, picking the best ones. To get something that's actually usable, you'll have to tinker around a bit and give it a few tries. With DALL-E, feedback is slower, and there's reluctance to just hammer prompts because of credits.
This is mind blowing.
I see it as a cheap and fast alternative to paying a concept artist.
But not a revolution. Creating precise and coherent assets is going to be a challenge, at least with the current architecture.
From a research perspective this is, I think, much more than a toy, those models can help us better understand the nature of our minds, especially related to their processing of text, images and abstraction.
Whereas things we associate more with computers, such as hard thinking, mathematics, etc. turn out to be more difficult to copy by a machine, and therefore perhaps more "human".
Did you learn about the "Industrial revolution" or the "agricultural revolution" in class? That didn't take a week, or a year, or a decade to happen. Even the Internet revolution took more than a decade.
This is a revolution. And you're seeing it happen in real time.
Here's some of the stuff I generated: https://imgur.com/a/mfjHNgO
Mostly-automated installer.
Instead of labeling data for what things are, we'll have to label things as being generated or not.
Great write up