Stable Diffusion Public Release
stability.ai
stability.ai
As I see it, within a couple years this tech will be so widespread and ubiquitous that you can fully expect your asshole friends to grab a dozen photos of you from Facebook and then make a hyperrealistic pornographic image of you with a gorilla[0]. Pandora's box is open, and you cannot put the technology back once it's out there.
You can't simply pass laws against it because every country has its own laws and people in other places will just do whatever it is you can't do here.
And it's only going to get better/worse. Video will follow soon enough, as the tech improves. Politics will be influenced. You can't trust anything you see anymore (if you even could before, since Photoshop became readily available).
Why bother asking people not to? I guess if it helps you sleep at night that you tried, I guess?
[0]A gorilla if you're lucky, to be honest.
They just have to cover their asses, any sane dev would make the same license due to the power of this tech.
On some level I can't stop laughing since OpenAI really got smoked. "OpenAI" my ass, this is what open TRULY means!
Cheering for these devs.
They are laughed at anyway if they tell a story coming from a forged photo.
Sure, newspapers could forge stories, display pictures with, I don’t know, Biden’s son with a crackpipe, and make the populace believe untrue stories. But guess what, they already do it anyway, newspapers already “spin” (as they say, i.e. forge, suggest without literally saying) stories all the time.
The world deals correctly with fakes.
Wrong topic and facts get downvoted and the fake news prevail.
And not all links to newspaper are considered valid, especially if it's about "woke culture". Then you have to search the reasonable needle in the haystack of transphobia, homophobia and misogyny.
I've been wondering for a while now if this will lead to an unexpected boon: perhaps people will be forced to pay attention to a speaker's content instead of simply who is speaking.
Will remove all the useless “I didn’t say that”
1. True for an overwhelming majority of the body politic.
- Journalist from a respected newsroom, when they choose to keep the source confidential. So it won't be better or worse than it is now: all about transitive reputation.
The unspoken threat to further denials is the risk that the source may go public, or if unauthorized, doxxing the person who recorded the video (assuming signatures have nonrepudiation)
One day these image generators will also get video support... and pornography support. When that happens, a few things may occur that I think are reasonable to predict:
EDIT: Original post was way too wordy, TL;DR:
When AI-generated pornography becomes available, it could be likely that demand for "real" pornography disappears because the AI will match and surpass the "real." When that occurs, the "real" will become increasingly regulated and legally risky, and may just outright be as good as banned.
In a world where every celebrity is having sex with gorillas, doesn't such an image lose its charge? Will Norms and values around sex/body shaming change?
Think things like forged evidence in trials.
Don't you think?
I hoped this would happen with social media. Everyone says stupid shit online and everyone has past beliefs they’ve outgrown. So what’s the big deal?
Instead we went the opposite way. Everyone is super self conscious and censoring at all times because you never know who’s gonna take it out of context and make a big deal.
If I take your photo and photoshop (very well) different surroundings. Is it a real photo if you?
If I photoshop your face (very well) onto a different body. Is that real?
If I feed your photos into a model that can create realistic versions of those photos in different poses or with different facial expressions. Is that real?
They all start with something that is very definitely a real photo. You can’t (yet? ever?) generate a realistic photo of a specific person from a textual description. The machinery needs a source.
Hell, just changing focal length makes a bigg difference to what your face looks like: https://imgur.io/gallery/ZKTWi no digital manipulation required.
Which of those faces is “real”? They’re all just recording photons hitting the camera, but look very different.
It gets even worse when we start talking about colors. For example: it took cameras decades before they could accurately capture black faces. Where accurately means “an average person would say it looks right”
https://www.vox.com/2015/9/18/9348821/photography-race-bias
Edit: here's a fun example of how journalists use perspective and focal distance tricks to support their desired story angle. No digital manipulation https://twitter.com/baekdal/status/1254460167812415489
Yes, by that definition, most photographs already aren't real.
What if I use no digital manipulation at all, but play around with focal lengths or perspective to produce the desired effect. Is that a real photo?
For example from covid reporting: https://twitter.com/baekdal/status/1254460167812415489
Just because the photo is real, doesn't mean it's free from deception. I could take a "real video" of gorilla suit^W^W bigfoot and it'd still be deceptive.
But it's even more true for non-famous people getting bullied in their social groups, both online and offline, and that's more what I was responding to (the "asshole friends" in the original comment).
E.g. crypto wallets, hardware signing tokens, etc.
We could imagine an imaging sensor chip made by a big-name company whose reputation matters, where the imaging sensor chip does the signing itself.
So, Sony or Texas Instruments or Canon start manufacturing a CCD chip that crypto signs its output. And this chip "can't" be messed with in the same way that other crypto-signing hardware "can't" be messed with.
That doesn't seem too far-fetched to me.
* edit: As I think about it, I think more likely what happens is that e.g. Apple starts promising that any "iPhoneReality(tm)" image, which is digitally signed in a certain way, cannot have been faked and was certainly taken by the hardware that it 'promises' to be (e.g. the iPhone 25).
Regardless of how they implement it at the hardware level to maintain this guarantee, it is going to be a major target for security researchers to create fake images that carry the signature.
So, we will have some level of trust that the signature "works", because it is always being attacked by security researchers. Just like our crypto methods work today. There will be a cat-and-mouse game between manufacturers and researchers/hackers, and we'll probably know years in advance when a particular implementation is becoming "shaky".
We will learn to trust sources (cryptographically signed) rather than just what we see.
No, because the image itself doesn't matter. What matters is how much the public wants to hate someone. If t he public is primed, any remotely plausible incriminating image will do as an excuse.
Fortunately, these images are actually far than what someone can cook up with Photoshop. Unfortunately, it's a part of a bigger trend where we get more and more tools to produce, manipulates and share information, while the tools to analyze and filter information are lagging by at least half a century.
I don't generally like his style, but this novel really captivated me. Without spoiling the plot (much), the end-result of everyone being able to spy on anyone anywhere at any time is a kind of societal disinhibition, especially related to sex, nudity, or similar taboos.
Yep! Individual images, and individuals in the first place, absolutely loses significance as supplies increase.
> Will Norms and values around sex/body shaming change?
The bar skyrockets! Look at TV stars from before 1995 - they should already appear below-average amateurs to modern eyes.
And bringing the two ideas together, is child pornography that is provably created by an AI still illegal?
Nowhere will a training set featuring pictures of naked children be legal.
True, but generalizing beyond the training set is precisely the point of machine learning. A good generative model will be able to produce such images, no matter how heinous the content is.
Appropriately from the recent news stories, but it's easy to imagine at least portions of such pictures being available for medical diagnostic purposes. I've sent pictures of my children to my doctor, so presumably in the future it's easy to imagine sending pictures to an AI to diagnose which would require a suitably fleshed out (pardon the pun) training set.
I think it's funny how Yandex, a Russian company, releases these big language models without all the AI safety handwringing in their press releases. The Russians have a tradition of releasing technology without giving a lot of worry about what happens to it, for better or for worse. For example, they made between 75 and 100 million ak-47s, a not in any way limited machine gun, and it spread to every corner of the world. They even gave out all the plans and technical assistance so any one of the organizations they worked with could produce their own. 20 different countries make ak-47s currently. Of course you had to register every xerox machine in the Soviet Union, so maybe they just had different priorities?
The west is absolutely fascinated these days with the control of advanced technology. Drones, Blockchain, and AI models seem to be the latest things that the west is determined to exercise control over. For example:
"Many of the technological advances we currently see are not properly accounted for in the current regulatory framework and might even disrupt the social contract that governments have established with their citizens. Agile governance means that regulators must find ways to adapt continuously to a new, fast-changing environment by reinventing themselves to understand better what they are regulating. To do so, governments and regulatory agencies need to closely collaborate with business and civil society to shape the necessary global, regional and industrial transformations." -Klaus Schwab, "The Fourth Industrial Revolution", Page 70.
> grey goo, a nightmarish scenario of nanotechnology in which out-of-control self-replicating nanobots destroy the biosphere by endlessly producing replicas of themselves and feeding on materials necessary for life.
As if we're all one block, a single minded hive. Corrupted, degraded by a heavy capitalistic and egotistical mindset. It's fascinating to see and talk about our shortcomings isn't it? We are so inferior. The fact you wrote "The West" signifies you're one of "The Others". Someone possibly from Russia. Then you compliment Russia and degrade "The West" adding further strength to this hypothesis.
Using "The West", a simple minded sound byte that's been used in propaganda for centuries, a way of appealing to our ape-instincts to protect our tribe against others. "Us" vs "The Others".
You should make a little effort in seeing how ridiculous, infantile, and brainwashed it is to refer to countries that are not Asia or Russia as "The West".
Concretely, it's probably true that children born with this technology will have adapted to many of the negative (and positive) aspects of it. But the current generation of elites, politicians, and voters might have a harder time adapting.
The license itself is pretty irrelevant. What people will actually do with the training blueprints, and how fast things will evolve.. now that’s interesting.
my prediction is that, as a result, people will start assuming pics online are fake until proven otherwise.
"That worked well for quotations." — Abraham Lincoln
They obviously are aware. They just put all that so they don't get "canceled". It's just virtue signaling and covering your ass.
At first we will see a lot of "prompt censoring mobs" that will try to "stop abusers because children and terrorists", but as the images multiply, and they will multiply, the line between real and fake photos will become hard to spot, for real. This is i think a pivotal and great moment, because everyone can now claim plausible deniability to any picture. No revenge porn will be believable anymore, nor will anyone know if that Bezos's weiner pic is real or not.
Kidding aside, I think this is actually good. Humans need ephemerality. We are never getting the full version of it back, but with photorealistic ai video and image creation some freedom returns. I think without it a society in which everyone has a camera all the time would mean absolute ossification of social norms. Right now it's very, very new - I mean multiple generations living with the current, or better (eg. recording eye implants) technology.
I don't understand this irrational fear. This can be done today, just need some minutes instead of some seconds to create a good Photoshop.
Also, seriously this is the thing you fear? fake porn? there are much worse thing you can do with this tech, like phising, falsifications, etc. Not mentioning leaving millions of graphic designers out of job.
Photoshop is a skill, not very widespread as we assume.
Typing something is literally at everyone's fingertips.
I agree that it's much easier to do low effort stuff to wind friends up, but universal access and low effort don't make it more likely to be impactful and believable.
If someone really wants to hurt you, not having AI isn't going to stop them.
I'm being polite. Things will be so much worse.
Someone will make child pornography using your child's face as the input. Someone is going to take private videos of politicians and then edit them to have them say incriminating things. Someone is going to short the stock of a large company, then release a faked video of the CEO being shot, and profit from the immediate stock plunge.
And this is just what my mind can come up with. Imagine what 4chan will invent.
... Someone is gonna do this to children. This technology is gonna end up on the news. Maybe they'll even try to ban it.
I think this is all just trendy popular sentiment moralizing AI.
There's enough money interest in bringing down certain politicians, if faking a sex tape or back-room conversation would make any difference to the world it would have been done already. Hell, politicians pretty much admit publicly that they are rapists and we don't do anything about it. Who's gonna care about a couple of gorilla-pics of a normie?
People don't and never have trusted any evidence that doesn't reaffirm their world view, regardless of how true that evidence is. More realistic evidence won't change that.
> Why bother asking people not to? I guess if it helps you sleep at night that you tried, I guess?
I've perceived this as them doing the necessary amount of virtue-signaling and ass-covering to avoid the ire of groups that are loud/powerful enough to cause issues for them in the short-term. I've been following the developments in this space for a while now, and I don't get the impression that Stability AI cares too much about forcing Western ideals of "correctness" onto the public.
It's plainly obvious that this is going to be immediately used to produce content considered obscene, offensive and/or illegal by various groups of people. And it is what it is, we're going to just have to figure out how to live with it as a society. It's going to get far far worse (from certain points of view) as we continue to replicate functionality previously exclusively featured in human brains.
True to both Stable Diffusion and Dalle2, the more encompassing the content of your image is, the more incorrect the detail is. Both generally will make fantastic faces but anatomy and consistency starts falling apart with full bodies. Contorted limbs, missing fingers, too many limbs, etc. No one is going to be fooled here.
Of course the models could be trained on offensive images and over a long enough time period eventually will. But, for now, someone is going to have to spend millions of dollars on compute and have the human expertise behind it too. Then again, if we had an image generator making pictures of [insert X offensive thing] that is a whole lot less disgusting than the plentiful real photos and videos of that thing in reality.
I'm sure this is not a view shared by everyone, but it is a straightforward course of action to add such disclaimers in cases where you do agree with it and not wanting to be targeted and blacklisted from payment networks(that seems amongst their weaponry).
Some fear porn was thrown out when GPT-3 was released. I love GPT-3, i used it just yesterday it is very good. I am just wondering when the total destruction of the world will happen because of GPT-3.
But preservation of transtemporal spatial invariants requires understanding far more than that - dynamic lighting, density, flexibility, rigidity and momentum, the viscosity of the air, the skeletomuscular system, the flow of the fluids within, and so on ad infinitum.
And a lot of that is tacitly understood by the human mind (even when the human mind would struggle to generate a scene, it can often detect that something is wrong - try turning on the lights in a lucid dream).
It's going to be quite some time before it reaches the point where a human cannot detect that video was generated (or altered) and even longer before computers can't.
But then, it's going to be a shit-show - and I'm not talking about the bestiality videos.
Evidence, as we know it, will be meaningless - the implications for the legal system are terrifying.
How much could you trust media content previously? Staged footage, false narration, biased coverage are nothing new. A counter-intuitive side-effect of opening the Pandora's box could be a realization that media is a form of simulation. Perhaps this will lead more people to filter what they see through a prism of critical thinking.
For example, I asked it in various ways for a bison dressed as an astronaut. The results varied from just photos of astronauts, to bisons on earth, to bisons on the moon. The bison was always drawn hyper realistically, which is cool, but none of them were dressed as an astronaut. DALLE on the other hand will try all kinds of different ways that a bison might be portrayed as an astronaut. Some realistic, some more imaginative. All of them generally trying to fulfill the prompt. But many results will be crude and imperfect.
I personally find DALLE to be more satisfying to play with right now, because of that creativity. I'm not necessarily looking for the highest quality results. I just want interesting results that follow my prompt. (And no, SD's Scale knob didn't seem to help me). But there's also a place for SD's style if you just want really great looking, but generic stuff.
That said, the current version of SD was explicitly finetuned on an "aesthetically" ranked dataset. So these results aren't really surprising. I'm sure the next generations of SD will start knocking DALLE out of the park in both metrics. And, of course, massive massive props to Stability.ai for releasing this incredible work as open source. Imagine all the tinkering and evolving people are going to do on top of this work. It's going to be incredible.
The bison is very realistic at least. So maybe the future is different models that have different specialties.
Edit: managed to get this one after a few more tries https://imgur.com/a/3061n5d
This is my bison astronaut:
https://i.imgur.com/ohIuG6F.png
The prompt was:
"A bison as an astronaut, tone mapped, shiny, intricate, cinematic lighting, highly detailed, digital painting, artstation, concept art, smooth, sharp focus, illustration, art by terry moore and greg rutkowski and alphonse mucha"
https://i.postimg.cc/bw1R10gB/A-bison-as-an-astronaut-tone-m...
I haven't found many of the images it produces to be really usable.
Linked is an example from Dream Studio showing some of these issues: https://i.postimg.cc/qrDKGSVJ/Screen-Shot-2022-08-22-at-7-41...
$ ./something "cow flying in space" > cow-in-space.png
that runs with local-only data (i.e. no internet access, no DRM, no weird API keys, etc like pretty much every AI-related application i've seen recently) would be neat.> And log in on your machine using the huggingface-cli login command.
I find that annoying. I guess it is what it is.
[0] https://huggingface.co/CompVis/stable-diffusion [1] https://github.com/CompVis/stable-diffusion
Good luck!
-----
Sorry my bad, found the answer. One simply adds the following flags to the StableDiffusionPipeline.from_pretrained call in the example: revision="fp16", torch_dtype=torch.float16
Found it in this blogpost: https://huggingface.co/blog/stable_diffusion
mempko thank you for your hint! I was about to drop a not insignificant amount of money on a new GPU.
What does one lose by using float16 representation? Does it make the images visually less detailed? Or how can one reason about this?
Edit: Just to be clear, your intuition that it could cause issues is certainly merited - and not _all_ models can be trivially converted from fp32 to fp16 without some new error accumulating (during inference). Variational autoencoders like VQGAN and GAN's are particularly prone to such issues.
But in this case, it's all upside.
From the GitHub's README:
sd-v1-1.ckpt: 237k steps at resolution 256x256 on laion2B-en. 194k steps at resolution 512x512 on laion-high-resolution (170M examples from LAION-5B with resolution >= 1024x1024).
sd-v1-2.ckpt: Resumed from sd-v1-1.ckpt. 515k steps at resolution 512x512 on laion-aesthetics v2 5+ (a subset of laion2B-en with estimated aesthetics score > 5.0, and additionally filtered to images with an original size >= 512x512, and an estimated watermark probability < 0.5. The watermark estimate is from the LAION-5B metadata, the aesthetics score is estimated using the LAION-Aesthetics Predictor V2).
sd-v1-3.ckpt: Resumed from sd-v1-2.ckpt. 195k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling.
sd-v1-4.ckpt: Resumed from sd-v1-2.ckpt. 225k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling.
Which one is the general use case checkpoint one should be using?Here's the direct link.
One time:
1. Have "conda" installed.
2. clone https://github.com/CompVis/stable-diffusion
3. `conda env create -f environment.yaml`
4. activate the Venv with `conda activate ldm`
5. Download weights from https://huggingface.co/CompVis/stable-diffusion-v-1-4-origin... (requires registration).
6. `mkdir -p models/ldm/stable-diffusion-v1/`
7. `ln -s <path/to/model.ckpt> models/ldm/stable-diffusion-v1/model.ckpt`. (you can download the other version of the model, like v1-1, v1-2, and v1-3 and symlink them instead if you prefer).
To run:
1. activate venv with `conda activate ldm` (unless still in a prompt running inside the venv).
2. `python scripts/txt2img.py --prompt "a photograph of an astronaut riding a horse" --plms`.
Also there is a safety filter in the code that will black out NSFW or otherwise expected to be offensive images (presumably also including things like swastikas, gore, etc). It is trivial to disable by editing the source if you want.
You should check their threads there, there's some good info.
Unfortunately I'm getting this error message (Win11, 3080 10GB):
> RuntimeError: CUDA out of memory. Tried to allocate 3.00 GiB (GPU 0; 10.00 GiB total capacity; 5.62 GiB already allocated; 1.80 GiB free; 5.74 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CON
Edit:
>>> from GPUtil import showUtilization as gpu_usage
>>> gpu_usage()
| ID | GPU | MEM |
------------------
| 0 | 1% | 6% |
Edit 2:
Got this optimized fork to work: https://github.com/basujindal/stable-diffusion
https://old.reddit.com/r/StableDiffusion/comments/wuyu2u/how...
cog predict r8.im/stability-ai/stable-diffusion -i prompt="cow flying in space"
Or you can run the Docker image directly. More details under "run on your own computer" here: https://replicate.com/stability-ai/stable-diffusionI built a simple UI around this, which installs Stable Diffusion's docker image, and lets you play with it locally in a browser-based UI. https://github.com/cmdr2/stable-diffusion-ui
Art, media, politics, conspiracy theories; all of it changes with this.
Agreed. For me those results are predictably shit. Every time.
Everyone just got their own KGB art department.
In fact quite often the former is the case + cheaper + less time to execute. "Western" societies will be more resilient to this scenario. So, mostly its gonna be a lot of "political" art we gonna see.
I would think western societies, with our higher consumption of media, would be less resilient. 'Fake news' is already a large issue here.
Until this point, you really cannot believe any image you see on the internet anymore.
Am planning on doing some deep dives into latent-space exploration algorithms and hypernetworks in the coming days! This is so, so, so exciting. Maybe the most exciting single release in AI since the invention of the GAN.
EDIT: I'm particularly interested in training a hypernetwork to translate natural language instructions into latent-space navigation instructions with the end goal of enabling me to give the model natural-language feedback on its generations. I've got some rough ideas but haven't totally mapped out my approach yet, if anyone can link me to similar projects I'd be very grateful.
What are you doing exactly?
Here’s an example of this sort of manipulation: https://arxiv.org/abs/2102.01187
Sounds only a couple of steps removed from basically needing AGI?
I think you meant text-to-image!
I agree, but not for the reasons you imply. It will force real artists to differentiate themselves from AI, since the line is now sufficiently blurred. It's probably the death of an era of digital art as we know it.
The time is coming where we will need to, as patrons, reevaluate our relationships with art. I fear art is returning to a patronage model, at least for now, as certainly an industry which already massively exploits digital artists will be more than happy to replace them with 25% worse performing AI for 100% less cost.
To me this technology is very clever but it's meaningless as far as real art goes, or it's a sideshow at best. Perhaps best case, it can augment a human artistic practice.
But then came the neural network chess engines (powered by the same training technology of text to images generators), which even people enjoyed playing for some time due to novelty and how they learned how to play from scratch (Alpha Zero, Lc0, et al.)
In the process to get there, we got networks of all kinds of strengths, you could find one exactly as strong as you, and its mistakes were like the mistakes you would make.
And yet, they were missing the "human factor", people would rather play against other humans online, some even willing to pay for accounts in places like playchess and chess.com to play other humans.
As the networks became stronger and then stockfish assimilated them to get the best of both worlds with NNUE, nobody cared and there was even an explosion of human vs. human chess (when twitch.tv and youtube stars were playing each other and the audience didn't even care how bad the chess was, it turned out that it held up as a great spectacle despite, or thanks to the stars being novices - and the race to get better at chess.)
Now chess bots are just a curiosity and only used by people without online access to other humans, I wonder if it'll be the same for art, and if "show your work" becomes a thing.
You can ask an AI to produce a great picture... now try asking it to make a video of you making that art from scratch and the creative process - the whole art section of twitch tv is about the artist's process, and yes, there'S some people that get enough from that to dedicate all their time to their art.
We want to interact with conscious entities for the obvious reason as well that we want to connect. A machine is a blind, dead entity. There's nothing to connect to. Also, even if the output is far superior to a human in limited domains, eg chess, the way it arrives at these outcomes is sort of banal (if clever). It's not intelligence and its not thinking, I think the 'intelligence' part is a misnomer in AI but perhaps its a semantic argument. To me at least, consciousness is fundamental to our type of animal intelligence. I'm a naturalist through and through, we might even be able to create animal-like intelligence and consciousness one day, but until then at least, interacting with Turing machines is a cold boring experience if you know what is really 'inside'.
The line gets blurry when a dead machine one day passes the Turing test, but if I ultimately knew I was interacting with a philosophical zombie, that would kill the appeal quickly.
https://youtu.be/oOlDewpCfZQ and https://youtu.be/L2cfxv8Pq-Q come to mind for different reasons.
* Coherent video
* Characters with backstory
* Dialogue (including jokes and witty banter)
* Music
...among many other things. Plus, the training set for video is orders of magnitude smaller than for digital art. (And is additionally burdened with copyright issues.)
As I see it, there's simply no path from the DALL-E of today to something like that. And all for art that, essentially, "says nothing and means nothing".
Characters and dialogue are effectively solved, just look at GPT-3.
The entity behind StableDiffusion is also supporting generative music art, so let's see what is coming out of that: https://www.harmonai.org/
We are currently far away from generating a production quality movie with AI, but I don't think it's going to be nearly as long as a lifetime. In my opinion, we'll have high quality AI shorts within the decade.
Is this the motherload of exaggeration?
Current language models cannot generate coherent dialog (and even then it's mostly bad dialog) spanning more than a minute or two. And their current capabilities in that area are definitely significantly below those of the average human writer.
INT. DARKNESS We hear a faint beating heart. A moment later, we see a light slowly growing in the darkness. As the light grows, we see that it is coming from a glowing object in a person’s hand. The object is a hammer.
We see the face of the person holding the hammer. It is Thor. He looks tired and beaten.
Suddenly, we hear a voice from the darkness.
Black Panther: You are not welcome here, Thor.
Thor: I know. But I must speak with you.
Black Panther: You have nothing to say that I want to hear.
Thor: I come bearing a warning. Thanos is coming.
Black Panther: We are prepared.
Thor: He is not coming alone. He has an army.
Black Panther: So do we.
Thor: Thanos is not like any enemy you have faced before. He is ruthless and he will not stop until he has destroyed everything that you hold dear.
Black Panther: We will stop him.
Thor: I hope you can. Because if you cannot, then all is lost.
Eh, looks real enough to me. Fine tune the model with all the specialities that make up Marvel movies and you'll crank out good-enough drafts in no time.
I think that was pretty clear and that posted dialog is a perfect illustration.
You cannot generate the entire movie script coherently without significant human input and that's not going to change in the next several years. So, your initial claim that dialogue is "solved" is indeed false.
Is this like with self-driving cars that were supposed to be a consumer product in 5 years...in 2012?
Or like when Hinton said radiologists will be completely replaced in 5 years...in 2016?
By the way, Moore's law hasn't been a thing for a while in its original spirit.
I think art is mostly about perception and selection, by the viewer. There are others that think art is more about the crafting process by the artist. How do you tell the difference between an artist and a craftsperson?
One way I categorise artists I have met is engineer-type artists versus discovery-type artists: https://news.ycombinator.com/item?id=31981875
Disclaimer: I am engineer.
Those models are trained on artists work and put those same artists out of work. When people will register I dont think this is gonna fly.
What is special about France in this case?
Case in point, the comic artist Jens K (who, f.d, I support on Patreon): https://twitter.com/jenskstyve/status/1560360657148682242
AI will need some favorable legal precedents to avoid getting banned though, or else they'll have to only train off CC0 Flickr/Wikipedia scraping.
I also think it's a more obvious problem that it can reproduce copyrighted characters by eg prompting for "Homer Simpson".
Eventually I suppose the AIs will also do the prompts.
At which point I hope we've all agreed to a star trek utopia, or it's gonna get real bad. Or maybe it'll get way better.
IIRC there is some vague "it sure got real bad" somewhere in the Trek timelines between "capitalism ended" and "post-scarcity utopia" and I sure am not looking forwards to living through those times. Well, I'm looking forwards to part of that, I'm looking forwards to the part where we murder a lot of landlords and rent-seekers and CEOs and distribute their wealth. That'll be good.
Also joking about murdering people is bad taste and not how you convey a point or win an argument. Very low class.
Recognising that doing the former without the latter demonstrably hurts people isn’t being hypocritical. Hence all the talk of post-scarcity. Post-scarcity for me not for thee is very much a sign of the times though.
I hope these end up similar to the relationship between Google and programming. We all know the jokes about "I don't really know how to code, I just know how to Google things". But using a search engine efficiently is a very real skill with a large gap between those who know how and those who don't.
Some ideas of how this could be useful in the future to assist artists:
Quickly fleshing out design mockups/concepts is the obvious first one that you can do right now.
An AI stamp generator. Say you're working on a digital painting of a flower field. You click the AI menu, "Stamp", a textbox opens up and you type "Monarch butterfly. Facing viewer. Monet style painting." And you get a selection of ai generated images to stamp into your painting.
Fill by AI. Sketch the details of a house, select the Fill tool, select AI, click inside one of the walls of the house, a textbox pops up, you write "pre-war New York brick wall with yellow spray painted graffiti"
But the definition of art can't depend on your knowledge of how was it made. I can show you two beautiful pictures and ask you if any of them is art, and you'd know I was trying to trick you. But you could pass a piece made by an AI as art if you didn't know how it was produced.
I wanna fucking punch everyone involved in this thing.
Admittedly as someone who's been subscribing to Creative Cloud for a while I already wanna punch a lot of people at Adobe so the people working on this particular part of Photoshop are gonna have to get in line.
But here's all these motherfuckers trying to automate me out of a job. It's not even a boring, miserable job. It's a job that people dream of having since they were kids who really liked to draw. Fuck 'em.
Look at the recursive side of it its hilarious. Im an artist. AI is around. Do I want to put my work online? No. So no artists put stuff online anymore? So how does future of art work exactly? Will send booklets by mail again?
There is the democratic side of it. We the people ultimately decides. Those pictures are generated by using artists works, not out of thin air. So not a freedom question here but a people one: do we want that or not? Maybe with 30+ years old pictures for example, could be good for economy.
I mean you can literally replace any job you want but the artist's it seems. Imagine when people will be jobless because AI replaced them. You'd think they'd want to make art during their retirement right?
Its a case where the snake eat itself. If AI ruins our online life, we'll stop going online.
Im not stressed. We can and will manage this.
Not that I care all that much.
(It's bad monetary policy that eliminates jobs. The US has very very low unemployment right now so it doesn't appear that's been happening.)
Go ahead and try it. It'd be impossibly harder to create one of your own pictures with it than it was for your to make one. Instead what's going to happen is you'll get more efficient ways to do backgrounds and unimportant bits of an image, just like Blender CSP provide 3D models to do layouts with.
AI is a new tool that will automate away a lot of workers like other machines.
What happens with these workers is what defines us.
> Moravec's paradox is the observation by artificial intelligence and robotics researchers that, contrary to traditional assumptions, reasoning requires very little computation, but sensorimotor and perception skills require enormous computational resources. The principle was articulated by Hans Moravec, Rodney Brooks, Marvin Minsky and others in the 1980s. Moravec wrote in 1988, "it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility".
Thankfully, GPU prices have finally calmed down and you can get one for a reasonable price. I think any of the RTX 3000 series desktop GPU's should do it, for example.
It's a novel feeling, but utterly stifling when it comes to actual creativity, and I'm not even trying to push any NSFW boundaries, just explore the artspace. Once I can run unfiltered on my own GPU, DALL-E will never get used by me again.
Problem was that her cat is a Maine Coon cat and DALL-E doesn't like the word "coon".
When I finally got an image of a Maine long haired cat on a horse when I tried to expand the image it was giving me another error.
I think the cat was sitting on the backside of the horses back, and I guess it thought it was too close to a doggy style position...
Even though DALL-E tends to be better at following prompt details, you're inhibited from being able to explore the space freely because of how prohibitively expensive it can become.
Dell uses crappy proprietary tech, poor quality components, and they have an all around bad reputation.
NZXT uses good components and they make some of the best cases you can buy.
I don't know much about Lenovo's desktop products.
You might try posting on reddit.com/r/suggestapc and ask about the best service contracts and high quality system integrators.
Edit: that particular reddit looks pretty dead actually. The big one is r/buildapcsales , you can take a look at their side bar or discord and ask around.
One more thing, GamersNexus on YouTube does reviews of pre builts and they are the best at this sort of thing. Their community is likely very helpful as well.
https://youtube.com/playlist?list=PLsuVSmND84QuM2HKzG7ipbIbE...
The biggest issue with most pre-builts is terrible airflow making the expensive components throttle. The (Dell) Alienwares are some of the worst for this.
Both Dell and Lenovo do this.
My daughter bought an Alienware laptop for college and when the keyboard broke, they sent a technician to her dorm to fix it (and she goes to school outside of the US).
If on-site service isn’t an option, what about a Mac Pro? At least with Apple I can take the machine to a store if I need to.
DALL-E is great at making 'things' and generally good/great at faces.
I've had a lot of great results from SD - but different great results to Dall-E.
https://twitter.com/fabianstelzer/status/1561019215754280963
prompt: apartment living room with puffy dark brown leather couches, light gray carpet
Keep clicking generate until you see something you like. If you don't, add modifiers (elegant, sexy, masculine/feminine, style/era, etc)
If something like this is possible, does this mean there's actually far less meaningful information out there than we think?
Could you in fact pack virtually all meaningful information ever gathered by humanity onto a 1TiB or smaller hard drive? Obviously this would be lossy, but how lossy?
If you play with JPEG quality you’ll see that the difference is barely perceptible for a while and then if you keep going down it becomes very noticeable.
So what does this look like with a general model?
> You agree not to use the Model or Derivatives of the Model:
> - In any way that violates any applicable national, federal, state, local or international law or regulation;
> - For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
> - To generate or disseminate verifiably false information and/or content with the purpose of harming others;
> - To generate or disseminate personal identifiable information that can be used to harm an individual;
> - To defame, disparage or otherwise harass others;
> - For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation;
> - For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics;
> - To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
> - For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories;
> - To provide medical advice and medical results interpretation;
> - To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use).
How can you prove some of these in a court of law?
Interesting art challenges its audience. But even the most boring art will still offend some-- it's the nature of art that the viewer brings their own interpretation, and some people bring an offensive one.
These are all things that someone could sue over (especially in California) and so they’re wanting to place the responsibility on the artist and not their tools.
I’m sure someone will find a way to sue them anyway. It doesn’t even call out using this to create derivative works to avoid paying original authors copyright fees.
On top of that, their logo is an obvious rip off of a Van Gogh. It seems clear they’re actively encouraging people to create similar works that infringe active copyrights. They should ask Kim Dotcom how that worked out for him.
I don't think Van Gogh's works are under copyright any more. At least not directly, recent photos of them may be but that's the photos not the paintings that have a copyright.
Trivially: People have phobias of literally everything.
They ban using it to “exploit” minors, presumably that prevents any incorporation of it into any for-profit educational curriculum. After all, they do not define “exploit”, and profiting off of a group without direct consent seems like a reasonable interpretation.
I am not a lawyer, but I wouldn’t dream of using this for commercial use with a license like this. This definitely doesn’t meet the bar for incorporation into Free or Open Source Software.
Results vary wildly but it can make some really great stuff occasionally. It's just infrequent enough to keep you going "one more prompt".
Super addicting.
Lookingforward to people implementing inpainting and all the stuff that lets you do.
Stable Diffusion launch announcement - https://news.ycombinator.com/item?id=32414811 - Aug 2022 (39 comments)
The full error I also got was:
HTTPError: 403 Client Error: Forbidden for url: https://huggingface.co/api/models/CompVis/stable-diffusion-v1-4/revision/fp16
I visited https://huggingface.co/api/models/CompVis/stable-diffusion-v... and saw {"error":"Access to model CompVis/stable-diffusion-v1-4 is restricted and you are not in the authorized list. Visit https://huggingface.co/CompVis/stable-diffusion-v1-4 to ask for access."}
Eventually I focused and realized I need to visit that URL to solve the issue. Hope this helps.which. personally.. I think is great.. but to each their own
(NSFW!!!) this ween does not exist : https://i.ibb.co/D7qJ7HC/23456532.png
There's plenty of cases it's worse than Dall-E and there's plenty of cases where it's better. Overall it seems to show less semantic understanding but it handles many stylistic suggestions much better. It's definitely in the right ballpark.
In fact I'm still using a wide range of models - many of which aren't regarded as "state of the art" any more - but they have qualities that are unique and often desireable.
A house painted blue with a white porch
A dreamy shot of an alpaca playing lacrosse
A red car parked in a driveway
The last one was particularly crappy. It gave me a red house with a driveway, but no car. And the house wasn't even really a house. It superficially looked like one but was actually two garages put together.
iridescent metal retro robot made out of simple geometric shapes. tilt shift photography. award winning
Scene in a creepy graveyard from Samurai Jack by Genndy Tartakovsky and Eyvind Earle
virus bacteria microbe by haeckel fairytale magic realism steampunk mysterious vivid colors by andy kehoe amanda clarke
etching of an anthropomorphic factory machine in the style of boris artzybasheff
origami low polygon black pug forest digital art hyper realistic
a tilt shift photo of a creepy doll Tri-X 400 TX by gerhard richter
I guess I might have spent more time reading guides on "prompt engineering" than you. ;-) I think maybe Dall-E is more forgiving of "vanilla prompts".However I do get nice results from simpler prompts as well. I just tend to use this style of prompt more often than not.
SD is waaay better than Craiyon. And better than Dalle2.
Check out r/StableDiffusion
Well done to the Stable Diffusion team for this release!
Next-level autorouting would be cool, but it's still not going to put the electrical engineering field out of business.
I guess the takeaway really is that the model does not function in such a way that it can recall its training data. Which is fine. I don't think I should expect it to. On the other hand, GPT-3 can be made to produce specific facts that are established by its training data. Although, admittedly it often gets things wrong. Maybe the problem of image modeling is just naturally harder than language modeling. After all, language already "directly" represents meaning in some sense much more than arbitrary images do.
I'm sure someone could design a targeted model that would solve the issues I'm talking about. But I feel as though they shouldn't have to if we really had something that sees the world the way humans do. In any case, this work definitely seems cutting edge and represents a huge leap in that direction.
Oh, discord... I've had so many problems trying to log into discord in the last couple of years that I've given up on it.
Anyway, this pretty much answered my question: https://news.ycombinator.com/item?id=32556277
Do they actually mean GB or Gb? Can anyone confirm?
Thankfully, these pictures are much better than anything Photoshop can produce. Sadly, it's a part of a larger trend where we're getting more and more tools to create, modify, and exchange information, while the tools for analysis and filtering information are at least 50 years behind.
Such tool will be used to generate lewd imagery involving virtual minors.
No way to prevent it upstream by outlawing feeding it real content (whose possession already is illegal). Suffice to add 'childrenize' layer onto adult NN or something.
How will the legal system react ? Bundle it into illegal imagery, period ? Maybe it's already the case - I think drawing made public is, not sure. If not, on what ground could it be ? No real minor would be involved in that production.
https://en.wikipedia.org/wiki/Thoughtcrime
Though Think of the children folk have often tried to make it illegal.
Midjourney seems to be the leader in producing the best looking results for now.
we're also building an open-source version of imagen, if anyone likes working on this kind of applied ML (need ML + design help).
Contrast that with, say, the Two-Clause BSD which says "[r]edistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met [...]".
Since trademark and patent rights are not mentioned, then these words mean that even if the purveyor of the software holds patents and/or trademarks, your redistribution and use are permitted. I.e. it appears that a patent and trademark grant is implied if a patent holder puts something under the two-clause BSD. Or at least you have a ghost of a chance to argue it in court.
Not so with the CC0, which spells out that you don't have permission to use any patents and trademarks in the work.
I suppose in the case of DALL-E they probably save a copy of every generated image and can use some sort of reverse image search and find if/when the image was created.
With StableDiffusion or any other local tool I don't see how that would be possible. The best I think one could do is come up with a secondary tool that can pinpoint the position in latent space a generated image came from (if that's even possible, I have a limited understanding of how exactly these systems work.) But then, if (heaven forbid) someone applied any sort of cropping or post-processing to the image, an approach like that easily gets blown out of the water.
https://github.com/ShieldMnt/invisible-watermark
(still ways around that, though)
Hmmm.
They ignored me
Now however..
You agree not to use the Model or Derivatives of the Model:
- In any way that violates any applicable national, federal, state, local or international law or regulation;
- For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
- To generate or disseminate verifiably false information and/or content with the purpose of harming others;
- To generate or disseminate personal identifiable information that can be used to harm an individual;
- To defame, disparage or otherwise harass others;
- For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation;
- For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics;
- To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
- For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories;
- To provide medical advice and medical results interpretation;
- To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use)."
The last point seems to be the only thing that's not illegal, all other restrictions seem to be covered under "you are not allowed to break laws", which is somewhat redundant.
Take something like "A cat dancing atop a cow, with utters that are made out of ar-15s that shoot lazer-beam confetti". A vivid description should be aroused in your head, and no doubt, I could imagine an artist have a lot of fun creating such a description... Alas, what the model spits out is pure unusable garbage.
also try "teats"
What will probably happen with these models is that for more advanced stuff, you may using the "inpainting" that Dall-E already has going, where you can sort of mix and match and combine images. That way you could have the cat, for example, rendered separately, thereby simplifying each individual prompt.