Open-source rival for OpenAI’s DALL-E runs on your graphics card
mixed-news.com
mixed-news.com
But, the interesting thing is that while it is so good in producing detailed artworks and matching the styles of popular artists, it's surprisingly weak at other things, like interpreting complex original prompts. We've all seen the meme pictures made in Craiyon (previously Dalle-mini) of photoshop-collage-like visual jokes. Stable Diffusion with all its sophistication is much worse at those and is struggling to interpret a lot of prompts that the free and public Craiyon is great with. The compositions are worse, it misses a lot of requested objects or even misses the idea entirely.
Also as good as it is at complex artistic illustrations, it is as bad at minimalistic and simple ones, like logos and icons. I am a logo designer and I am already using AI a lot to produce sketches and ideas for commercial logos, and right now the free and publicly available Craiyon is head and shoulders better at that then Stable Diffusion.
Maybe in the future we will have a universal winner AI that is the best at any style of pictures that you can imagine. But right now we have an interesting competition when different AI have surprising strengths and weaknesses and there's a lot of reason in trying them all.
The whole "we're all fellow scholars here" bit feels like I'm watching a kid's show about science vulgarization, patting me on the head for being here.
"Look how smart you are, we're doing science!"
I dunno. I like the channel for what it is (a vulgarization newsletter for cool ML developments) but sometimes the author feels really patronizing / full of himself.
In fact Dall-E isn't as impressive for some styles as "older" models (Jax/Latent Diffusion etc)
> 515k steps at resolution 512x512 on "laion-improved-aesthetics" (a subset of laion2B-en, filtered to images with an original size >= 512x512, estimated aesthetics score > 5.0
https://github.com/LAION-AI/laion-datasets/blob/main/laion-a... for more details.
What's remarkable is this: https://github.com/LAION-AI/laion-datasets/blob/main/laion-a...
That aesthetic predictor was apparently trained on only 4000 images. If my thinking is correct, imagine the impact those 4000 ratings have had on all of the output of this model.
You can see samples (some NSFW) of different images from the original training set in different rating buckets here, to get an idea of what was included or not in those training steps. http://3080.rom1504.fr/aesthetic/aesthetic_viz.html
The amount of enjoyment I have derived from playing with Craiyon over the last two months is ridiculous.
Note I think you need 16gb of VRAM to run it.
I started with https://github.com/borisdayma/dalle-mini/blob/main/tools/inf... and pared it down.
I've checked out Disco Diffusion but hadn't heard of Visions of Chaos, thanks. The biggest shortcoming to DD is there's simply not yet a sufficiently trained model to produce stuff to the level of MidJourney or Craiyon
Also the results vary drastically if you have enough VRAM to load more models e.g. a 3090 (24GB) or an A6000 (48GB). I've been saving money and waiting impatiently for 4090s to drop. Check out the Disco Diffusion or VoC Discord - people post their works in there and often you will see results that make you wonder if they're cheating ;)
You can also try disabling use_checkpoints if you have extra VRAM, since it will render a bit faster (but uses more VRAM since it doesn't save intermediary 'checkpoints' to disk).
When/if you get bored, try disabling use_secondary_models which will use a lot more VRAM but can deliver results on a completely different level. You will likely struggle for a few days figuring out which parameters to tweak to get good results (e.g. tv_scale, sat_scale, etc, which are otherwise AFAIK ignored).
In any case I recommend reading A Traveler's Guide to the Latent Space, which I call "The Bible" since it covers so many topics and has links to various studies and will keep you busy reading for months ;)
Also check out the Discord for Disco Diffusion and Visions of Chaos, as you can read endless tips and tricks to getting amazing results.
Have fun! :)
[0] https://www.reddit.com/r/DiscoDiffusion/comments/t7p4bi/seas...
[1] https://sweet-hall-e72.notion.site/A-Traveler-s-Guide-to-the...
I can't use it, but check out this guy's work. incredible detail
> “A percentage of people are simply unpleasant and weird, but that’s humanity,” Mostaque said. “Indeed, it is our belief this technology will be prevalent, and the paternalistic and somewhat condescending attitude of many AI aficionados is misguided in not trusting society.”
Holy shit.
On the one hand, I'm super excited by this technology, and the novel applications that will become possible with these open-source models (stuff that would never be usable if Google and OpenAI had a monopoly on image generation).
On the other hand, I really really really hope Bostrom's urn[0] has no black ball in it, because we as a society seem to be rushing to extract as many balls as possible over increasingly short timescales.
And it depresses me, because well... imagine if other pieces of tech were treated this way. If the internet or crypto or computers or whatever were heavily limited/restricted so the 'wrong people' couldn't use them for bad things. We'd consider it ridiculous, yet it's somehow accepted for these image generation systems.
It sounds riddiculous, yet Photoshop is more dangerous than DALL-E in all regards.
I'm personally in the "AI risk is overstated" camp. But if I'm wrong, all the top-down AI safety in the world is going to be meaningless in the face of a global network of researchers, enthusiasts, and tinkerers.
It's the reason why the rapid amount of progress in this field scares me at times. It feels like gradually being crushed under a wall of the inevitable march towards progress. It could be the case that stopping ourselves before it's too late isn't possible.
Sometimes I get the feeling that the laws of nature will eventually destroy or severely impact any lifeform that gains too much of an understanding about the world. For example, in some other universe it might be possible to survive X more decades if the laws of physics weakened the effects of nuclear war just enough for civilization to recover in a relatively short period of time, but we're stuck with what we have, and that doesn't necessarily map to the long-term survival of a highly intelligent lifeform.
It would be a shame if unbounded curiosity would be humanity's undoing. That curiosity is also a part of me, and my family, and my neighbors down the street, and those in the situation room.
You don’t, so the question is mostly academic.
It's similar to the "we don't sell personal information" claim. Sure, but that's because they make money renting malicious actors access to a black box that contains your personal information. Selling the contents of the box would reduce their overall revenue.
It makes sense for Google to wait for some startup to "go first" in releasing a model largely without controls. That way, some random startup takes the initial heat of "people are using AI for bad things!!" headlines plastering tech blogs. Then Google can do basically the same thing a little bit later, and any attack pieces will sound old hat.
But it's arguably better than living in a world where every single citizen has a nuke.
(Though the potential for harm of diffusion models is far below nuke; it's not "kill millions of people", it's "produce cheap disinformation and very convincing fake evidence to ruin someone's life")
Either way it doesn't matter, you can't control bits like you can enriched uranium. It's just a matter of time. In the grand scheme of things Open AI will be irrelevant.
Yes, it's pretty obvious that Dall-E and similar models won't destroy humanity.
My point isn't that Dall-E is a black ball. My point is we better hope a black ball doesn't exist at all, because the way this is going, if it exists, we are going to pick it, we clearly won't be able to stop ourselves.
(For the sake of dicussion, we can imagine a black ball could be "a ML model running on a laptop that can tell you how to produce an undetectable ultra-transmissible deadly virus from easily purchased items")
I think we’re already past the point where we could have done something about this. In fact, we’ve probably been past that point since humanity was born.
I think it’s probably more valuable if we think about how we’ll deal with it if we do draw something that could be/is a black ball.
That said, so far all evidence points to extreme destruction just being really hard, which leads me to believe that truly black ball technologies may not exist.
Like 'black balling', eg shunning?
Or 'black box', eg poorly understood technology?
> black ball: a technology that invariably or by default destroys the civilization that invents it.
Apart from nuclear scientists I don't know a field where participants are as conscious of the risks as AI research.
Great. Now some of these researchers preceived some risk with this technology. Not human extinction level risk, but risks. So they attempted to control the technology. To be specific: OpenAI is worried about deepfakes so they engineered guard rails into their implementation. OpenAI was worried about misinformation so they did not release the bigger GPT models.
Note: I’m not arguing either way if OpenAI was right, or honest about their motivations just observing that they expressed this opinion and acted on it to guard the risk.
Got this so far? Keep this in mind because I’m going to use this information to answer your question:
> If you know that image generators are not it, then why talk about it here?
Because it is a technology which were deemed risky by some practitioners and they attempted to control it, and those attempts to control the spread of the technology failed. This does not bode well towards our ability to restrain ourselves from picking up a real black ball, if we ever come across one. And that is why it is worth talking about black balls in this context.
Note it is unlikely that a black ball event will completely blindside us. It is unlikely that someone develops a clone of pacman with improved graphics and boom that alone leads to the inevitable death of humanity. It is much more likely that when the new and dangerous tech appears on our horizon there will be people talking about the potential dangers. What remains to be answered: what can we do then? This experience has shown us that if we ever encounter a black ball technology the steps taken by OpenAI doesn’t seem to be enough.
This is why it is worth talking about black ball technologies here. I hope this answers your question?
So I think that horse may have bolted already.
It sounds like you are getting hung up on details of a particular example. That is not useful. We can’t give you exact details for a particular black ball because we haven’t encountered one yet. Sadly the fact that we haven’t encountered one yet doesn’t mean that they don’t exists.
Think about it like this: There are technologies which are easier to stop spreading and there are technologies which are harder to stop spreading.
Example for a technology which is easier to stop: Imagine that a despotic government wants to stop people from space launches. All the known tech to reach orbit is big and heavy and requires a lot of people. It is comperatively easy to send out agents who look at all the big industrial installations and dismantle the ones used for space launches. There will be only a handful of them and they are hard to hide.
Now an example for a technology which is harder to stop: imagine that the fictional despotic government has it in for cryptography. That is a lot harder to stop. One can do it alone in the privacy of their own home! All you need to do is some pen and paper. That can be hidden anywhere! A lot lot harder thing for the agents to find and distrupt.
We talked about how easy to stop the spread of a given technology. Now let’s think about something else. The potential of a given tech to cause harm.
An example for a risky technology: nuclear weapons. If you have them you can level a city. That is a lot of harm in one pile.
An example for a less risky technology: ergonomic tool handles. Those rubbery overmoldings which make it nicer to use the tool long term. There is no risk free technology, but I hope you agree that these are a lot less dangerous than a nuclear bomb.
Do I have you so far? Good. Because this was the easy part. We talked about things which already exists. Now comes the hard part. This requires some imagination: We have seen tech which was easy to control and tech which was harder. We have seen tech which was risky and tech which was less risky. Can these properties come in all combinations? In particular: are there technologies which are both risky and hard to control? Something for example where any able human can accidentally or intentionally level a city or kill all humans? I can’t give you an example, we don’t have technology like that yet.
The example you are asking about is an example for this kind of technology: high risk, hard to control.
Nobody says that you can download software today from github which can help you engineer a deadly virus from household chemicals. This does not exist. It is a stand in for the kind of tech which if it were possible it would mean that we have a high risk, hard to control technology
Does this help explain the context better? Let me know if you still have questions.
> Does this help explain the context better? Let me know if you still have questions.
That's the problem. The thing you're scared about (dangerous technology) has nothing to do with the context (AGI) because there's no reason to think AGI is especially capable of creating any of it or is going to. Humans create general intelligences (babies) all the time and you aren't capable of, nor are you putting any effort into, "aligning" babies or stopping them from existing.
AGI being superintelligent won't give it superhuman creation powers, because creating things involves patience, real-life experimentation and research funds, and while I'll grant you computers have the first they won't have the other two.
Some form of AGI under some circumstances might be black ball tech. There can be other black balls which have nothing to do with AI let alone AGI.
> The thing you're scared about
I’m scarred about many things but black ball tech is not one of them.
One can discuss existential risks without being scared about it.
> they won't have the other two
If you say so? I don’t agree with you on this, but it feels this would mislead the conversation, since AGIs and black ball tech has at most some overlap.
Well, that's what the article's about. (With some generalization, since I don't think anyone expects art AI to be civilization-ending.)
Humanity is not homogeneous, we will always react to new inventions or tools differently , many will use it positively some won’t . Short of weapons of mass destruction I am not sure anything else will destroy civilization itself .
A black ball would be like - suppose nuclear weapons ignited the atmosphere. We test the first nuke, it ignites the atmosphere, a global fire storm consumes all breathable oxygen, kills all plants and everyone on the surface and everything else suffocates shortly after. Plastics aren't even close to this level of harm.
> No, the black ball is a technology that, once invented, humanity cannot survive. Nuclear weapons have been invented and humanity is surviving. Same with plastics.
I don't think that definition is a good one. Technological civilization [1] has survived nuclear weapons for ~80 years, but there's no guarantee it will survive it for another 80 years, let alone forever. It seems like these "black balls" should be though of like time bombs, there are at least two variables: how much destruction it will cause when it goes off AND the delay time before that happens. We shouldn't confuse a dangerous technology with a long delay time for a safe technology. My intuition tells me that there will probably be nuclear war at some point over the next 1,000+ years.
[1] I don't think nuclear weapons can make humanity extinct, so long as there are still little poorly-connected subsistence communities in remote areas. However, if The Market, manages to extend its tentacles into every human community, we're probably fucked.
I think I see the issue, and I think that summary leaves an important facet out that the paper talks about. It would probably be better as "a black ball: a civilization-destroying technology that cannot be regulated, so invariably or by default destroys the civilization that invents it"
Dall-E isn't an x-risk, but an advanced AI might be (though a lot of people have their opinion on that part).
The nuclear weapons deployed during the 60s and 70s were far more powerful than the ones of today. Instead of multi-megaton yields being the default, now most warheads are in the 100-300kt range. (This is largely due to improved accuracy reducing the size of a warhead required to take out a target) That means 2-3x the damage radius of the Hiroshima and Nagasaki bombs at 10-15kt. (radius does not scale linearly with yield)
> I was under the impression the US nuclear arsenal alone can wipe out humanity.
Even if you assume 1 nuke = 1 city, the US only has ~5k warheads. According to [0], city number 3000 has 141k inhabitants. Now, indirect effects are going to kill a LOT more (no more industrialized agriculture or global supply chains), but that'll still only get you to 99% at most. (And that's in this extremely contrived scenario where it only takes one nuke to kill Beijing, and not a few dozen)
Plastics are playing the long game. They have to turn into micro- and nanoplastics first and may then enact undesired, unforeseen biological functions, just like BPA [1].
Not even talking about weaponizing this stuff...
Isn't the point of the discussion started by PoignardAzur about how we deal with such technology after it is pulled out of the bag?
If you define black ball technology as fundamentaly uncontainable, then there is no point in talking about our practices of restricting access to new tech.
You're right that they could always change their minds and that would suck, but so far they seem to be being up front.
This was before DALL-E and way before DALL-E 2.
Not that I think they need to. It’s just weird people defend them as being open just because they release crippled versions of their model.
* They're actually working on a filter right now, but IIRC it's an optional one, for when you don't want to accidentally generate NSFW output.
You may be interested to know that the code for SD has already been released on Github, and they've given the weights to a bunch of researchers in preparation for full release. I've also heard that one of the researchers leaked the weights earlier today, and 4chan has been using them for, uh, stuff.
Stable Diffusion launch announcement - https://news.ycombinator.com/item?id=32414811 - Aug 2022 (37 comments)
Not sure how they compare. DD seems to be quite popular. I'm currently setting up DD locally.
I will otherwise refer you to the "Bible" of latent diffusion: https://sweet-hall-e72.notion.site/A-Traveler-s-Guide-to-the...
Whatever isn't covered in there is probably in the Disco Diffusion cheatsheet: https://botbox.dev/disco-diffusion-cheatsheet/
There are tons of resources out there, and it's a nonstop learning and experimenting process to try to achieve what you want.
"We couldn't process your request at this time. Please try again later. If you are seeing this message repeatedly, please contact Support with the following information:
ip: XXXX
date: Mon Aug 15 2022 XX:XX:XX GMT-0700 (Pacific Daylight Time)
url: https://stability.us18.list-manage.com/subscribe/confirm"
Does any one know how much data download and disk storage does it need?
Larger images will require (much) more than 5.1 GB. In my case, a target resolution of 768x384 (landscape) with a batch size of 1 will max out my 12GB card, an RTX3080Ti.
I've generated some remarkably good-looking print quality images by upscaling 512x512 sources
I've seen a number of 80s/90s VHS recordings of concerts being uploaded to YouTube in 4K (using Topaz) and they look like they were recorded that way, truly amazing. I do hear it can be a bit of work though getting the settings right.
https://colab.research.google.com/github/KaliYuga-ai/Pixel-A...
[0] https://nuwa-infinity.microsoft.com/#/ [1] https://arxiv.org/abs/2207.09814
I actually tried to get Dalle to do this, And it made like three good sprites in the rest were just broken. But it was so strange, because you could see it was still organized as a sprite sheet, it's just the sprites were useless.
I think the practical applications of this technology will be hyper specialized models for specific purposes.
Hopefully we aren't more than a few years away from something practical like this.
Currently waiting to scale-up for improve quality mainly for economic reasons, not quite sure I could recoup the training costs yet. Even more so if I go with cloud training.
NVidia will release the 4090 in september, and ethereum may do "the merge" that will make GPU useless for mining so GPU price could be affordable so I can update my home cluster with affordable 3090s. (But electricity prices are also up).
Also there are new algorithms every month like the stable diffusion, that would obsolete your previous training.
The video generation cost is probably still too expensive compared to just paying a cam girl in a low wage country. But it will probably go down soon.
This is also some sensitive data, as plagued with copyright issues, so it's quite troublesome to legally share training datasets to share costs.
It also has its own challenges with respect to custom dataset creation with text description, so it's probably a better idea to adapt the algorithm to the currently available data to keep the costs low.
Finally once someone releases a model, in the next month there will be at least 3 clones.
There is also the problem to find an adult friendly payment processor.
And the multitude of potential legal issues.
But it's probably inevitable.
So no, your "friend" can't use it for that.
Since we certainly don't want to show generated image of porn or violence to someone that didn't specifically ask for that, the easiest way to ensure that's not going to happen is to just not train on that kind of data in the first place. The worst that can happen with a model trained on "safe" images is that the image is irrelevant or makes no sense, meaning you could deploy systems with no human curator on the other end, and nothing bad is going to happen. You lose that ability as soon as you integrate porn.
Also with techniques like in-painting, the potential for misuse of a model trained on porn/violence would be pretty terrifying.
So the benefits of training on porn seems very small compared to the inconvenience. I don't think it's anything to do with puritanism, it's just that if I am the one putting dollars and time to train such a model I am certainly not going to be taking on the added complexity and implications of dealing with porn to just to make a few people realize their fetishes at the risk of my entire model being undeployable because it's outputting too much porn or violence.
uh have you seen American/European mainstream pornography? it's already pretty violent (ex. face slapping, choking, kicking, extreme bdsm).
I just don't see why this stuff is allowed and protected by the law (if its not recorded and published its illegal) and then we are suddenly concerned about what text can do.
Just one of the many double standards I see in Western society.
That's not at all what I am talking about. What I am saying is that such a model would give everyone the ability to create extremely realistic fake images of someone else within a sexual/violent context, in one click, thanks to inpainting. This can become a hate/blackmail machine very fast.
Even though Dalle-2 is not trained on violence/porn it still forbids inpainting pictures with realistic faces that have been uploaded by users to prevent abuse, so now imagine the potential with a model trained on porn/violence.
Someone is eventually going to do it, but back to your initial question about why it's still not done yet, I believe it's because most people would rather not be that someone.
Someone like Trump would just shrug and say the pee tapes are deepfaked. I don't think its possible for AI to bypass forensics either. So again this narrative that "deepfake blackmail" would be dangerous makes no sense.
Of course, after a while it'll probably stop working, but there will be a period of time where it can be done profitably and a longer period where it will be obnoxious.
And, of course, you could probably always use the tool to scare children, who, even in the future, might not know that everyone would shrug off the generated pictures.
...and the more people will be aware of and stop believing in the "fake reality".
Ensuring this technology is only available to a tiny subset of the population is to essentially give all the power of distorting reality to that tiny group of people.
In fact, I suspect that is precisely the reason.
A porn model should really be it's own thing.
Make no mistake this is an indirectly legal hurdle.
There are attempts to gather porn images and train or fine-tune existing networks on it, here's a recent attempt by an art student mentioned in the article above (NSFW!!): https://www.vice.com/en/article/m7ggqq/this-furry-porn-ai-ge...
However, if it is possible to re-start the training process from the weights of a non-sexually aware model, this finetuning might not take all that long..!
Making money at it should be easy, and places like PornHub wouldn't care about any outrage. The real challenge would be limiting criminal and civil liability, at least to my not-in-the-business thinking.
[0] https://timdettmers.com/2020/09/07/which-gpu-for-deep-learni...
(No affiliation.)
So far, I haven't got a response but it's a Monday
Here's the license
https://github.com/CompVis/stable-diffusion/blob/main/LICENS...
(And if you're insinuating something, just come out and say it so people can engage appropriately)
Still not sure I understand. It's already available in two forms: a Discord bot for a wide group of beta users, and the "researchers-only" source release.
OpenAI only have a paid SaaS version of Dall-E
Would love to see FAANG and SV crash and burn, margins chipped away to nothing.