DALL·E Now Available Without Waitlist
openai.com
openai.com
Furthermore, the pricing model is much worse for DALL-E than any of its competitors. DALL-E makes you think about how much money you're losing continuously - a truly awful choice for a creative tool! Imagine if you had to pay photoshop a cent every time you made a brushstroke. Midjourney has a much better scheme (and unlimited at only 30/month!), and, of course, Stable Diffusion is free.
This is a step in the right direction, but I feel that it is too little, too late. Just compare the rate of development. Midjourney has cranked out a number of different models, including an extremely exciting new model ("--testp"), new upscaling features, improved facial features, and a bunch more. They're also super responsive to their communtiy. In the meantime, OpenAI did... what? Outpainting? (And for months, DALL-E had an issue where clicking on any image on the homepage would instantly consume a token. How could it take so long to fix such a serious error?) You have this incredible tool everyone is so excited to use that they're producing hundred-page documents on how to get better results out of it, and somehow none of that actually makes it into the product?
https://labs.openai.com/s/rCzJwauuiaIj1Pd3IyJGaHS3
Here I wanted an illustration of a nuclear plant in a japanese landscape, first attempt with Dalle produced multiple good results. I tried SD and MJ (back when MJ didn't use SD) as well, had trouble even with multiple attempts:
https://labs.openai.com/s/FxhxtMFe3kFS8msV8vekRAJ3
There are others, but anyway I think my examples are not important since it will be always easy to cherry pick prompts that yield the best results in model X.
In my experience SD is good at producing (especially non-photo-realistic) art that looks pretty and DALL-E is better at following a specific prompt when I know what exactly I want.
Of course I recognise your experience might (and probably does) differ.
[0] - https://wafflegame.net/
For example, DALL-E performs extremely impressively on prompts in the format of “a still of homer Simpson in The Godfather” (replace character and movie as you wish). with the other two it’s a lot of misses
I'd love to see a site with lots of examples of the same prompt fed into various models, I assume someone has already made that.
I dunno, I generated 20 images from that prompt locally and got three good ones[1].
https://twitter.com/Dalle2Pics/status/1534718848137560064?re...
Have to let the AI experts speculate on why SD goes nuts there because it definitely knows what "The Godfather (1972)" means (if you ask for e.g. 'A still of Patrick Stewart in "The Godfather (1972)"' you get one - which I believe DALL-E can't do because of their facial restrictions?)
Turns out a shitload of misses are acceptable when it only takes 4-7 seconds to generate an image from a prompt. 5000 generations on an RTX 3090 takes around 7 hours +/- 30 minutes, by the way.
Those are the params I use in ImaginAIry, mileage may vary if you're using a different package.
3KWh is like $0.5 USD ...
DALL-E would give you like two pictures for that price, LOL
> house interior, friendly, playful, video game, screenshot, mockup, birds-eye view, top down perspective, jrpg, 32 bit, pixel art, black background
SD absolutely demolishes DALL-E on this one. SD produces really nice-looking output, with a high degree of consistency. DALL-E produces incoherent nonsense.
You can't just compare SD and DALL-E performance on prompts alone, because SD gives you a lot more levers to steer it in the direction you want.
This is a con for some prompts. As an example, I asked for a painting of an elephant and a dog drinking tea together. The result was a dog with an elephant nose next to a teapot.
A similar misfire was the word 'porcupine' which drew pigs, I guess because porc is in it? Anyway, it's idea-blending is a little too aggressive.
I would argue that none of these follow the prompt. they all represent a goodfather frame in simpson stile, which is not about placing homer in a godfather still.
I would heartily disagree - I've generated ~6.5k images using SD locally and most of them could be linked to the prompt they came from.
> ...and most of them could be linked to the prompt they came from.
You made it sound as if there is almost no connection between the prompt and the images and zimpenfish said that the majority could be linked, implying a strong connection. He/she doesn't have to be praising it at all to counter your claim.
(Which is probably better than you'd get from a human given the exact same prompts.)
https://www.reddit.com/r/StableDiffusion/comments/xcq819/dre...
I'm running it on Windows 10 using (a modified version of) https://github.com/bfirsh/stable-diffusion.git and Anaconda to create the environment from their `environment.yaml` (all of which was done using the normal `cmd` shell). Then to use it, I activate that env from `cmd` and switch into cygwin `bash` to run the `txt2img.py` script (because it's easier to script, etc.)
[edit: probably helps that I already had a working VQGAN-CLIP setup which meant all the CUDA stuff was already there. For that I followed https://www.youtube.com/watch?v=XH7ZP0__FXs which covered the CUDA installation for VQGAN-CLIP.]
Have to admit just started looking into it, mb there are better options
Granted, DALL-E appears to be buckling under demand regardless so the supply/demand curve doesn't warrant a price drop yet.
On the other hand Stable Diffusion emerged as a free tool where large community can experiment and search for the killer app together. People started adapting it into other tools and workflows and so far it seems like the magic is in finding prompts that make the device generate good quality outputs. Earlier today I saw announcement about lexica.art(Stable Diffusion prompt tool) getting funded.
Making art/weird pictures doesn't have to be useful, as that use case is the entire reason MJ/SD went viral.
The inpainting plugins with Photoshop and Krita are already working absolute wonders.
MidJourney and others are actually useful for exploration but the outputs are not because they can't spit finished deliverables to the specs. No one is paying for a picture of Mermaid eating marmalade, trending on art station, beautiful face, sharp focus, octane 8k.
They are great for exploration, it's just that I don't believe this is the killer app for these tools. We will find out what's the killer app with Stable Diffusion because with Stable Diffusion people can experiment beyond entering some prompts.
I think a major problem is reproducibility and output controllability. Rolling the dice multiple time and using some of the outputs is not good enough for most applications.
Maybe this can be solved at some point but it's not at this moment. The advantage of Stable Diffusion is that it can be possible for someone to implement it, with OpenAI this feature doesn't exist and its not useful until they implement it.
It's the start of a product, but it's going to keep improving. Already with inpainting and outpainting we see some new possible uses. What NovelAI (which builds on Stable Diffusion) has shown so far of their upcoming release seems impressive, though it's hard to say how much of that is cherry picking.
> Rolling the dice multiple time and using some of the outputs is not good enough for most applications.
Hmm, is that true? I feel like most of the time art that companies want is something made far ahead of the consumer seeing it, so generating a 100 versions of something and picking the best seems fine, especially if you can then use inpainting and img2img to fine-tune it.
They also went for B2B first. Which is weird. Why not parallel a B2C app? It could be a subscription or packs of drawings. It would generate buzz and give useful data on the sorts of things real people type into these systems. I
Their goals were never about openness at all though. From the beginning I’ve felt like they should’ve called themselves something like “SafeAI”, since their stated goal was basically to develop advanced AI first, then keep a lid on it until they could somehow ensure it was “safe” or would only be used by “good” people.
Sure, OpenAI might sound nicer, but it also drags this contradiction into the foreground whenever someone says their name.
removing "Vandyke" from the prompt lets it go through[1], but doesn't result in the style I want. Because there's no artist that I'm aware of that goes by "Henry Carter". The middle name is important.
It reminds me of the old 2D Runescape days where the language filter would convert "dictionary" to "**tionary".
[1] https://en.wikipedia.org/wiki/Scunthorpe_problem
[2] https://www.techdirt.com/2018/08/31/scunthorpe-problem-why-a...
(That second link describing how AI can't understand language well enough to solve the problem predates GPT-3 by 2 years)
I'm actually on a 3 day Facebook ban because I posted a very legit medical NIH link that looked quite innocent (despite dealing with sexual organs- it was an article about a case of urethral intercourse, which I was including to demonstrate how little people, or perhaps just Americans, seem to understand human anatomy) and unbeknownst to me, the preview card picked a closeup of some kind of vaginal surgery to feature, and that resulted in an instaban. Fffffuuuuuuuuu
It was basically the visual version of this
-We are the world's most advanced AI company.
-Our filter verifiably acts as a simple blacklist.
-You aren't allowed to see the blacklist because it's really a "contextual" filter, so you'll have to guess.
-If you guess wrong too many times you'll be banned.
-Using our service more often increases the chance you'll hit the number of wrong guesses.
-No, you can't know what that number is.
This religion definitely has a parentalist bent to it that rubs a lot of people the wrong way. I vaguely recall them floating on Twitter the theoretical idea of whether murdering people to prevent AI-takeover is acceptable, due to how bad AI-takeover is.
Not surprising limiting access, spying on what its users are using their tools for, etc, is acceptable to them.
This is much in the same vein as how for Lenin, the eventual triumph of the working class is so important as to justify a little bit of interim violence, dictatorship, and summary executions.
We had a Greens candidate here running for election, and his main campaign promise was to shut down a 10+ billion dollar motorway underpass that was 99% completed at the time of the election.
His approach for "saving the environment" is to convert efficient motorway travel into inefficient stop-start traffic. And to throw away tens of billions of dollars of already spent resources.
Banning plastic packaging and waterproofing.
My bet is that just like before i was born, after I’m gone these things will only increase.
This is something that is pretty well addressed in their writings and is definitely a heavily considered topic. It gets quite heavily into philosophy because ideally you want to count future-you's preferences as well as current-you and you want to avoid, say, genocide just because most people dislike a certain group.
A good place to start is here: https://www.lesswrong.com/tag/coherent-extrapolated-volition but there is a LOT of discussion of this topic.
"In calculating CEV, an AI would predict what an idealized version of us would want, "if we knew more, thought faster, were more the people we wished we were, had grown up farther together". It would recursively iterate this prediction for humanity as a whole, and determine the desires which converge. This initial dynamic would be used to generate the AI's utility function."
This is too detached from the reality of AI development to be useful. I can't make a utility function like this, nor can you, nor can humanity. There's no reason to think a human-equivalent AI could either - the data for such a function doesn't exist, and can't exist because of how vague all the terms are. The current ML revolution is built on statistical pattern recognition. This definition would better fit a genie.
The fact that MIRI does no ML research and has no dialogue with the state of the craft only furthers my impression that it produces a lot of words with little substance.
Even if one still accepts the ultimate conclusions of the movement, using its AGI-focused rhetoric to justify restrictions on a simple image generator is silly.
That being said, the absolute majority of AI safety theory seems to fall into the same pothole where philosophy falls; modelling the world through language-logic rather than probabilities. The example you quoted fits this category - it's way too specific and thus unlikely to be useful in any way, even though its wording may deceive its author to believe it to be an inescapable outcome.
We are quite literally watching the world burn before our eyes, and we are too stupid and self-centered to do anything about it.
> it will kill a million people yearly
And it's even more outlandish that such a number and atrocity upon mankind is brushed away. That is one million humans.
I'm not trying to crucify any group here by the way. This is just how I see the issue.
I wish there were enough naturally curious people willing to dive into the data and come to their own conclusions. When you look at the history of Earth's climate, and the speed with which greenhouse gases natural release into the environment, and compare it to the speed with which they now release (due to burning of fossil fuels, reduction of trees/plants that trap it) it is really quite common sense that what we're doing is not a good idea to continue indefinitely.
The AI safety crowd's main concern isn't that an "unaligned" superintelligence will have some other people's values. It's that an unaligned AI might kill everyone.
Of course, that isn't a concern with DALL-E. It's like if the safety crowd was worried about a ferocious tiger and OpenAI was like, "we got a kitten. Let's keep the public safe from it while we develop better kitten-handling gloves."
Then Midjourney and Stable Diffusion get their own kittens and let everyone play with them and OpenAI finally says, "Okay, everyone can safely play with our kitten now because we carefully developed great kitten-handling gloves" and proceeds to hand you plain dollar store gloves.
In my view, speculating on AGI ethics is at best pointless. It's like trying to write laws for the Internet during the invention of the telephone. If you imagine how it could operate, you'll be wrong, and the details change the whole problem.
Where? I probably believe you but it almost makes me worried about the well being of the stability ai founders (on a long term horizon)
I've modified my OP to clarify that.
---
(2015) OpenAI's original "Introducing OpenAI Post" : https://openai.com/blog/introducing-openai/ : "As a non-profit, our aim is to build value for everyone rather than shareholders. Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies."
(2018) OpenAI's "Charter" : https://openai.com/charter/ :
"We are concerned about late-stage AGI development becoming a competitive race without time for adequate safety precautions. Therefore, if a value-aligned, safety-conscious project comes close to building AGI before we do, we commit to stop competing with and start assisting this project."
"We are committed to providing public goods that help society navigate the path to AGI. Today this includes publishing most of our AI research, but we expect that safety and security concerns will reduce our traditional publishing in the future, while increasing the importance of sharing safety, policy, and standards research."
---
Provides some interesting context to the fact that Elon left the company's board in February 2018 over "disagreements about the company's development."
Haha, and then they would proceed to get told to politely piss off.
That was the stated made up bullshit they spun because "we're keeping this walled to figure out how to squeeze the most profit out of it" doesn't go as well with their focus groups.
Any highly motivated group without much to do will seek out things to make themselves seem important and necessary.
Assume: - AGI wants to stay alive - Humans can create more AIs - Other AIs would compete for the same resources
Then: Easiest way to make sure that they would get no more competitors would be...
Anytime you see “creating an AI will obviously kill you” try reading it as “having children will obviously kill you” and see if it still makes sense.
Rather, it seems like evidence that singularitarianism is actually a religion (https://en.wikipedia.org/wiki/Millenarianism) which is why it believes things with magic powers will suddenly appear.
In particular, exponential growth doesn't exist in nature and always turns into an S-curve… of course it's a problem if it doesn't level out until it's too late.
I'd bet a ton that we're nowhere near the top: evolution almost never comes up with the optimal solution for any problem, almost by definition it stops at "meh, good enough to reproduce". And you don't need a ton of intelligence to reproduce.
Evolution's sub-optimality is actually one of the strongest arguments against intelligent design, so I'm really hesitant to agree that it requires any sort of leap to estimate that with some actual design it won't be very difficult to blow way past human intelligence once we can get there.
Well, define "intelligence". People seem to use it in a vague way here - it might be what you call a motte and bailey. The motte (specific definition) is something like "can do math problems really fast" and the bailey is like "high executive function, is always right about everything, can predict the future".
For the first one I don't think humans are near a limit, mostly because of the bottleneck in how we get born limiting our head sizes. But it is pretty good if you consider the costs of being alive - food requirements, heat dissipation, being bipedal, surviving being hit on the head, risk of brain cancer, etc, it's done well so far.
Similarly an AI is going to have maintenance costs - the more RTX 3090s it runs on, the more calculations it might be able to do, but it's going to have to pay for them and their power bill, and they'll fail or give wrong answers eventually. And where's it getting the money anyway?
As for the second kind I don't think you can be exponentially better at it than a human. At least if you are, it's not through intelligence, but it might be through access to more private information, or being rich enough to survive mistakes. As an example, you can't beat the stock market reliably with smarts, but you can by never being forced to sell.
The real mystery to me is why people say "AI could recursively improve their own hardware and software in short time spans". I mean, that's clearly a made up concept since none of humans, computers or existing AI do it. But the closest thing I can think of is collective intelligence - humans individually haven't improved in the last 10k years, but we got a lot more humans and conquered everyone else that way. But we're also all individuals competing with each other and paying for our own individual food/maintenance/etc, which makes it different from nodes in an ever-growing AI.
That's a relatively easy thing to do architecturally once you have a model that can match human intelligence at all. TBH if we could rearchitect the brain in code we could probably easily figure out how to do it in ourselves within a few years, but our wetware does not support patches or bugfixes.
We can't improve ourselves, but that's only because we're meat, not code. And of course no AI has done it yet, because we haven't actually made intelligent AI yet. The question is what happens when we do, not whether the weak-ass statistical crap that we call AI today is capable of self-improvement. Nuclear reactions under the self-sustaining threshold are not dangerous at all, but that was not a good reason to think that no nuclear reaction could ever go exponential and be devastating.
Doesn't seem like computers can improve themselves either. Mainly because they're made of silicon, not code. "AI can read and write its own code" doesn't exist right now, but even if it did, why is that also implying "AI can read its CPU Verilog and invent new process nodes at TSMC"?
(Also, humans constantly break things when they try changing code - the safest way to not regress yourself would be to not try improving.)
If we ever get them there, then it's likely that the usual resourcing considerations will come into play, and refactoring/optimization/redesign will be viable if you throw hours at them. But unlike with human optimization, every hour spent there will increase the effectiveness of future optimizations.
- children won't be smarter than any living human.
- children have a human brain which makes them predictable (constrained in behavior by current laws, institutions, and most likely a conscience).
A better analogy is to ask what happens when a species branches off and evolves into a smarter species, but the dumber ancestor species still exists.
I really don't understand the perspective that people in your position take. Is it that you don't think we'll arrive at super intelligent AI, and therefore there's little risk? Or that we will be able to control it? If you think we can control it, why? Like we're not that much smarter than our monkey ancestors and what hope did they ever have of stopping our absolute domination over them? And then all of you call this opinion a "religion" without even explaining why it's wrong?
You're right on the spot regarding the problem of having two different smart species on the same planet. We killed everything between us and chimps. Given enough time, the smarter species can be assumed to always take over.
Homo sapiens is still expanding even: https://dna-explained.com/2012/11/16/the-new-root-haplogroup...
That's why it seems to be a religion - it thinks intelligence gives you unlimited powers and makes your plans always work, it posits unseen entities with infinite amounts of it, and it tells you to move to Berkeley and dedicate your life to stopping them. Specifically, it's a kind called rationalist eternalism (https://meaningness.com/eternalist-systems).
For a specific example of undangerous superintelligence see Culture Minds, who only influence anything because of special programming to make them less of a general intelligence. The unbound ones immediately get bored with the real world, leave it and just play games in their head instead.
Also, I don't think any individual human has absolute dominion over monkeys? Human society as a whole yes, but society doesn't behave like a generally intelligent agent. A monkey is better than you at doing the things monkeys care about though.
I do think unintelligent machines are pretty dangerous. There's extremely dangerous machines called "cars" that have already taken over society and constantly kill people! And we buy their gas for them too.
> Also, it would have to get a job to pay its AWS bill.
Inference power costs are low now on agents that are better than humans at Chess and Go. It isn't going to be an issue after another 20-100 years of further R&D and optimizations. Nothing about the history of computing should tell us that this will be a big limiting factor.
Humans need shelter and jobs too. If you got an AI down to the energy requirements of a human that's not enough to avoid needing one. Especially if it's influencing the real world - entropy exists and all real world things cost money.
These systems aren't really lowering the barrier of entry on still-photography fake porn when previously anyone with Gimp & a few hours of video tutorials could churn out much the same thing.
I think text-prompt generated deep fakes (not just of porn) will present a significantly larger challenge for society, but I don't see that same scope of problem on still images.
On the other hand, everybody's been saying Pandora's warehouse was over there for a while -- it isn't really that they are to blame for showing us the way in or anything, I just don't understand what they were trying to accomplish.
if anything, OpenAI has made me more cynical about "AI safety" messaging because it looks like an excuse to take a cut and keep things proprietary.
This technique of pitting two AI against each other (a generator and a detector/discriminator) is called a generative adversarial network, and it's used a lot for unsupervised training.
I honestly think most of them are window dressing and aren’t allowed to have real influence though. They’re there for the PR, not to actually change things, but they honestly just make me really scared or big tech controlling AI.
An image classifier calling Black faces gorillas? Embarrassing, insulting, has to be fixed. AI pre-crime classifiers for police departments? I'm against it, across the board.
Do we really care that the image mulchers default to stereotypes? It means if you say "basketball player" they'll mostly be Black, if you just say "doctor" they'll mostly be white males (and probably balding with a stethoscope), but this can be qualified easily in the prompt.
It just reflects the training data, and the smart thing to do is shrug and add enough words to get the image you want. It's not trying to throw shade, it literally understands nothing, it's not able to understand things, just match text prompts to generated images.
Nerfing DALL-E by randomly adding 'diverse words' just makes it harder to dial in the image you want. Let's say you want a Vietnamese male doctor drinking coffee on break in Hanoi, it's not going to help you if 1/3rd of the images have "female" or "black" tagged onto it.
It just seems low stakes. We wouldn't come after a human artist who happened to paint a picture which conforms to simple occupational stereotypes, why should AI be any different? It's not like it will refuse to give you what you want if you ask.
It's good thing that the "safety measure" is the way it is - an afterthought. It means that those ideologues haven't yet had influence on the model itself.
OpenAI's mistake may have been "planning to have a business model;" the alternative they should have gone with was "Instead of taking investor money with promises of some kind of return, be a hedge fund manager, make $100 million, and then set $600,000 on fire with no plan to recoup the cost because it's play-money to you."
When that figure came out, the popular talking point was how cheap Stable Diffusion cost to make and how easily a well-funded competitor could create their own custom variant.
The founder said that this is not quite as possible with most public clouds and that it is easier to buy the GPUs.
Unsurprisingly, Pornhub is already using machine learning. Their first big project is colorizing and upscaling classic porn.[1] (Moderately NSFW). As they point out, they have plenty of training data.
Moreover, there are very rich people already, like Warren Buffet, Bill Gates and Elon Musk, funding projects for doing good like world hunger, education and "AI Safety". And Open AI was a project of this sort of thing, originally. The thing is that even very rich people demand that the enterprises they give money to be as self-supporting as possible and their money is spread fairly thin. The only way Open AI could become an AI development shop, employing many top developers, was to have the financing level of a commercial company. Which means it constantly puts out products that don't seem like they can make money because AI algorithms don't seem to controllable - Open AI seems to only be able to have the first implementation of X, not the best implementation. Once the basic idea is out, someone else can produce a similar thing with a budget that doesn't include a research team.
That seems a bit cynical. While SD's creator might not recoup that money directly, a lot of end users have benefited from its creation. That money has figuratively gone up in flames no more than the time or labor cost of an open source developer whose code is used by millions of people, IMO.
OpenAI wouldn't have been able to do what StabilityAI did because OpenAI is incentivized to make return on investment; Mohammad Emad Mostaque is not.
Somewhat tangentially, I speculate that crowd-sourced training will become a thing.
Israeli hip-hop band Shabak Samech created their last clip frame by frame in sable diffusion, took something like two days: https://www.youtube.com/watch?v=SnGP2Qx3ddg
I'm not sure who they planned to market this to but I can think of a few, not inherently lucrative to my knowledge, products here. Fictional literature illustrations such as for books seems like a great market in my mind, you can literally turn authors words into depictions without a graphic artist. I wouldn't be surprised if you could create graphic novels this way as well. Propoganda also seems like a market but the bar seems to have been lowered to memes.
Other than that, I struggled to think of applications you could make money from. Police sketches? Eh I doubt it would work well but maybe.
There is or course the visual art world which could potentially be impacted with AI generated artworks.
I'm guessing that worry translated to image generation as well. It's a Pandora's box thing I guess.
Your criticism is in my opinion not valid.
Do they need to react to the market? Perhaps depends on what there goal even is.
Is dall-e 2 fun to use and cost wise totally fine? For me yes.
But I also have people running SD with a hacky webui on some good GPUs for free. How many people actually have access to it.
Is there also a good benchmark on which tool is inherent better? Because it is also totally fine to have multiple offerings.
I really don't sure if you ever seen product development for yourself.
Dall-e clearly took the potential misuse risk much further than others.
And in another test the faces are super shitty.
Dall e also gives you 4 pictures per credit and dream 1.
So good to have more options I think. Two different products feeling different.
But I also played around with sd.
I still think my original comment is valid.
DALL-E is very good at conceptually representing complex prompt. Something like "a bear with a diving mask surfing in the ocean, a pelican is sitting on its shoulder", DALL-E will immediately produce coherent results, while SD requires lot of prompt tuning, and sometimes it's even impossible to get it to represent some concepts (I haven't tested this particular prompt tho)
SD is good for producing "artistic" images if that makes any sense
edit: ok I tried the "surfing bear" prompt with DALL-E 2 and SD and the results are consistent with my point, I put the raw prompt without tuning, and cherry picked the best image out of 4 with both models, here is what I got :
DALLE-2: https://labs.openai.com/s/Q9824QOfXln4r9FLFNM3v9v1
SD: https://imgur.com/a/czcMgiC
For SD, even by tuning the prompt I wasn't able to get the diving mask or the bird on the shoulder
The re
The "waitlist" model might work when the product isn't ready for prime time or the exclusivity is a part of the pitch, but it's greatly overrated in other respects. I got a "The Wait Is Over" email to tell me I'm off the waitlist and able to use a not-exactly-new stock trading app this week as the UK economy crashed. Yeah, thanks, but no thanks...
But maybe it doesn't matter, because many times more people are playing around with StableDiffusion, such that the absolute number of good images being shared around is much higher with StableDiffusion, even if the average result isn't great.
This is honestly not my experience at all. When I first tried SD and MJ, I did so with a very clear and distinct feeling that they were "knock-off DALL-Es" and I strongly doubted that they would be able to produce anything on the level of DALL-E. Indeed, I believed this for my first couple hundred prompts, mostly because I didn't know how to properly prompt them.
After using them for around a month, I slowly realized that this was not the case, and in fact they were outperforming DALL-E for most of my normal usage. I have a bunch of prompts where SD and MJ produce absolutely beautiful and coherent artwork with extremely high consistency, that when sent to DALL-E, give significantly worse results.
But if all you're doing is the equivalent of visual mad Libs: "Abraham Lincoln wearing a zoot suit on the moon.", then SD and MJ suffice.
Dall-E does seem more aware of relationships among things, but using parens and careful word order in some of the SD builds can beat it. By contrast, even most failed images from MidJourney could still be in an outsider art gallery. MJ aesthetic works, while Dall-E seems like a 9 year old was taken hostage and clipped out Rapunzel and the paper shredder from magazines and pasted them onto a ransom note.
That said, I have not been able to get any of Dall-E, MJ, or SD to give me a coherent black Ford Excursion towing a silver camping trailer on the surface of the moon beneath an earthrise.
At cost per image, I could pay to get complex concepts such as this rendered via any number of art-for-hire sites at less expense and guaranteed results.
[0] https://github.com/AUTOMATIC1111/stable-diffusion-webui/wiki...
OTOH, the main limiting factor for DALL-E 2 from my point of view is the ultra-aggressive NSFW filter. It's so bad that many innocent prompts get stopped and you get the stern message that you'll be banned if you continue, even though sometimes you have no idea which part of the prompt even violated the rules.
Sometimes hands will turn out just fine and sometimes they will suddenly become fine after some random other stuff is added to the prompt.
It's clearly still missing a bit in terms of accurately following prompts, but it's capable of generating a lot of things that may not have obvious prompts. This should improve a lot with larger models. I believe SD is already working on it.
Dall-E can’t even do many of the images SD can so seems silly to hold hands up as the AI art tool Turing test.
If I get a bad result from DALL-E 2, I used up one of my credits. If I get a bad result from Stable Diffusion running on my local computer, I try again until I get a good one. The result is that even if DALL-E 2 has a better success rate per attempt, Stable Diffusion has a better success rate per dollar spent.
This also affects the learning curve. I've gotten pretty good at crafting SD prompts because I could practice a lot without feeling guilty. I never attempted to get better with DALL-E 2, because I didn't really want to spend money on it.
I do admit that I rate the creativity of Dalle2 higher than that of SD. It can occasionally create really unexpected and exciting compositions, whereas SD will more often lean more conventional.
But anyway, SD is far superior even if you consider dalle better per image since you can create 1000 SD outputs and just pick the one you like best (which for sure will have one that’s better than the dalle output you got)
This is all early days and these demos are neat but the real value is yet to be seen. Maybe when this technology is licensed and integrated into Photoshop or Instagram or something like that.
I cannot speak to DALL-E's results, as the signup process is currently broken (after providing email, name, and phone number, was met with "We’re experiencing a temporary issue with signups due to a vendor outage. We apologize for the inconvenience!"), but the Stable Diffusion results I've been getting are not just unusable, but downright bizarre... here are the four images it produced for "morihei ueshiba doing a double backflip": https://imgur.com/a/EvkQpBT
Emad Mostaque, a millionaire hedge-fund manager with money to burn, spent approximately $600,000 to train a model and dumped it out for public consumption: no account for how it will be used, no concern about any sociopolitical consequences, damn the torpedoes and straight ahead. He basically burned down a potential industry space and hugely complicated an ongoing conversation on how these tools will interact with / disrupt the lives and livelihoods of artists... But he also basically changed the world overnight. Hashtag-squad-goals, am I right?
There's a lesson to be learned here. I haven't decided what it is yet. Though I note that it's a lesson that probably applies to few people who don't have $600,000 to set aflame.
I hadn't heard the story of how stable diffusion was created. Sounds like the guy is a true hero from your description. And only for $600k? Imagine if he decided to "burn" the rest of his millions on similar initiatives.
(Furthermore, if I don't use it very often, I'm in the free tier due to the 15 free credits a month.)
Also, do you realize that Stable Diffusion is also running a pay-for-usage model at dreamstudio.ai? I like that too.
Their marketing was excellent, but somehow pushed expectations too much and underdelivered. It also felt very elitist. Not very tinkers in a garage that this generation of stable diffusion feels like.
I don’t think other options put any or so much consideration about AI impact.
Perhaps we’re just finding out that people don’t care. Yet?
Overall, pricing will need to be adjusted over time as well. I set out on an experiment the other day that you can see here: https://twitter.com/Keyframe/status/1574338738808934400
I went about trying to utilize stable diffusion for an imaginary concept project (concept for characters of a remake of TMNT, heh). Process was similar to how I'd do it with another artist more than if I drew it alone. It was back and forth, from rough outlines and then honing into details. Inpainting and img2img helped A TON and I hope I'd get dreambooth running soon as well since that will be a game-changer in the combination of things.
Between exploration phase, detailing, alternatives, and manual painting and over painting, I'd say PER OUTPUT final image I created in the region of a thousand or so interim images. Process overall did take a lot of time but not as much as completely manual and I didn't feel like I had as much control as manual of course, but I did feel ultimately bold enough that I thought I had creative control. With dreambooth I expect it to close the gap.
Overall, I was extremely pleased with the experiment and I'll continue exploring it, even though I'm not doing artwork professionally anymore. And so far no, it's not going to replace artists. It's another tool removing labour, but adds time on direction needed. Ultimately it'll be another brush in the toolbox.
But maybe it was overconfidence? Or maybe what Stable Diffusion did was really just unpredictable.
Like back then when they held their GPT model from the public because "it could destroy the world" or some similarly deranged argument.
>the level of expertise in OpenAI's executive team would make them unlikely candidates to lose out strategically
Sure, to be honest they're making mistakes that even complete marketing noobs know to avoid. Anyway, it's nice to see others are eating their lunch and sharing it with us.
Running software locally and using our desktop or portable supercomputers for something other than web browsing. What a novel concept. But how is this possible without cloud?
Frankly it spoiled the whole ML/DL revolution for me which till that point of time was a little more open than other fields. True that companies like Google don't let you use their models like Alpha* freely or even open source them but they aren't dangling access like these pricks who called themselves open were doing.
I hope they will become irrelevant in the coming years.
Look, I get it: They don't want to be in the news for producing porn or gore. But if you block a prompt and threaten account closure on repeated such blockings, at least tell us what we did wrong.
In all fairness, their release of Whisper[0] last week is actually really amazing. Like CLIP, it has the ability to spawn a lot of further research and work thanks to the open source aspect of it. I hope OpenAI learns from this, downgrades the "safety" shills, and focuses on producing more high-quality open source work, both code and models, which will move the field forward.
Once you're able to get access, I believe you'll receive 15 credits per month for free. Each credit allows one generation, and each generation produces 4 images.
Rather than using up credits trying to learn how to formulate your prompts, I ran hundreds and uploaded them to https://generrated.com (I posted it a couple of weeks ago as a Show HN) — hopefully they might be useful as a starting point and save you some credits/money.
I've been doing something similar with some DALL-E 2 images: https://news.ycombinator.com/item?id=33011336
I really liked the 16th century Indian painting of an astronaut [1], and now I see a role of programs like DALL-E in giving people a good intuition on how to identify art from different styles and periods.
[1] https://generrated.com/?prompt=16thCenturyPainting&subject=a...
I'm not sure if you know (it might not be obvious — my fault!) but you can click on a prompt heading to see all 20 images created in one place, if that's more useful to you. e.g. https://generrated.com/prompts/16thCenturyPainting
What I'm finding interesting is twisting that to attempt non-Dean-like prompts within the 'gatefold album cover in watercolor' concept. It's not hard to get vibey graphics telling Stable Diffusion to paint in various mediums.
I would imagine you could do very decent Pollocks in Stable Diffusion if those are represented in the database. The artist's process shares elements with how the AI makes art.
I suspect that even in 6 months from now people will be starting to see consistently good generation from "worse" prompting. The cat is out of the bag on this and running way faster than anticipated. Hold on, because the "AI generating fake media" thought experiment of the last decade has now officially gone live.
It does seem DALL-E is better, though not that good, at text. From what I've seen it also seems DALL-E is much better, though still flawed, at understanding instructions about composition.
On the other hand I've seen an absolute deluge of amazing high quality work based on SD and little more than a number of kinda cool things from DALL-E. I've also not had access to DALL-E but have generated hundreds of images with Stable Diffusion on my own PC.
Sooner or later though, someone is going to come along and say "You know, I'd be fine with a 5% profit margin" and the house of cards will fall while the tech bros cry "value" the whole way down. You could trick yourself into thinking a sharpie is worth $40/mo if you drink enough of the "value" coolaid.
One example of an "inappropriate" prompt that stuck out in my mind, and I think is pretty representative - I was trying to recreate DVD box art of Breaking Bad, but replace the main characters with cats. When my prompts were things like "Meth dealing cat" I would get the "inappropriate" flag. Frustrating.
I much prefer stable diffusion. Quality is different, but being able to generate whatever I want without censorship is an enormous benefit. Plus, the cost. OpenAI is far too expensive. A friend and I are writing a novel. We wanted to see what it would be like to just feed each paragraph of text from our novel to the image generator, along with combinations of descriptions. This would be a pretty expensive experiment with DALL-E, but it can run locally for the cost of electricity with stable-diffusion.
People were able to discover this by typing a prompt, something like "person holding a sign that says" and it would output pictures of people holding signs that just say the word "black", revealing that it was actually generating images from the prompt "person holding a sign that says black"
https://twitter.com/rzhang88/status/1549472829304741888?t=R4...
My (very qualitative) feeling is that DALL-E 2 is good with composition and realism (e.g. generating photographs — you'll still get artefacts but it's less likely to look "computer graphics-y"), and is quite forgiving (you will usually end up with an image that makes sense).
Midjourney had a recent update and can now produce beautiful images with far more detail and realism than DALL-E 2 in some cases, especially for human and animal faces, but excels more on the computer art side of things. (Midjourney now has a community showcase gallery: https://www.midjourney.com/showcase/)
Stable Diffusion is a bit less forgiving than both, in my experience. Some people are able to create stunning images, but you have to invest more time into figuring out what works best.
I'm currently looking into taking images generated with DALL-E 2, then using them as a starting point for Stable Diffusion to add detail. It works partciualrly well for cartoon-style images.
For example:
- Original DALL-E 2 image of a horse in a city: https://i.imgur.com/CaNHHR7.jpeg
- That image used as a starting point for Stable Diffusion: https://i.imgur.com/EW1iKOO.png and https://i.imgur.com/VOQ35Oz.png
You can see it significantly cleans up the artefacts the original DALL-E 2 image had. (Note: the original DALL-E 2 image is 1024 pixels square, but Stable Diffusion generated a 512 square output.)
But Dall.e is often behind in terms of image quality. They are nice looking from far, but a bit more blurry or weird than stable diffusion if you look closely.
However you can use boths together. These days I tend to use stable diffusion first, but when a prompt is not going well I copy paste it in dall.e and get what I meant much easily. And then I import the dall.e generated image in stable diffusion to work it a bit more and get something a bit better looking.
Anyways, this space seems to be moving so quickly that it's difficult to keep up anyway.
However, President Raisi is very disappointed by OpenAI's brazen and disgusting display of female hair. It's extremely insensitive that OpenAI has enforced only a mere subset of the worlds cultural mores.
> What does Stablilty AI do? Stability AI is building open AI tools to provide the foundation to awaken humanity’s potential. Our values are lived by every team member and shown by everyone who excels at Stability AI. They are how we measure ourselves and our work. Our vibrant communities consist of experts, leaders and partners across the globe. They are developing cutting-edge open AI models for Image, Language, Audio, Video, 3D, and Biology. AI by the people, for the people.
Still don't know what it does. Continuing to the next FAQ:
> What’s our business model? We’re a company of builders who care deeply about real-world implications and applications. Many of our most considerable advances grow from working across multiple teams. We are unafraid to go against established norms and explore creativity. Our primary drive is to generate breakthrough ideas and convert them into solutions. We respect innovation over tradition. We trust that our differences make us more robust, and so we seek reason within every difference of perspective.
Oh well. I give up.
Though I agree that their website provides no useful information at all.
[0] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
No wonder my back hurts.
It doesn't have a fancy product website because it's not really a product, it's 'just' the model. Developers can use it to build a product. The Stability.ai people themselves built one called Dream Studio (https://dreamstudio.ai), but there are also some free and open source frontends you can run on your own hardware if you have a GPU.
I guess your confusion comes from the fact that people tend to talk about "Stable diffusion" and not "Dream Studio" or one of the many frontends available for it.
If they want to prevent fake war pictures, it should be done with their NSFW filter. Not with blocking the whole (suffering!) country.
I just can’t think of any good reason to do that. I can think of one bad, though
To be fair, the first ever prompt I tried out with an AI image generator like this is "putin eaten alive by pigs" :) It's hard to refrain.
> Something went wrong > OpenAI's API is not available in your country.
I used “login with google”, and this is an error that I immediately got, haven’t got to the stage where they want my phone number.
Here’s the real news! Just hope they disable the autoban for API users. It’s one thing to filter NSFW outputs like SD websites do, but blocking application API accounts for requests by end users would make it unusable.
No one is even talking about DALL-E anymore.
Thanks Stable Diffusion!
Good luck to them for being annoying AF when there are better and free and actually open alternatives!
I can’t imagine how many signups they’re getting right now. It’s almost behoove them to pregenerate a bunch of prompts and give people a sandbox version without signing in.
Honestly, I don't think I need DALL-E right now, as SD is free and MUCH MUCH more customizable.
This time, verification worked. My account already existed.
With SD and DALL-E making headlines nearly weekly, we've heard nothing but crickets chirping from Google's Imagen team.
With SD if I write "$X walking on the moon with $Y in a bold majestic style" the results will ofteb only include $X OR $Y. Dall-E appears to do a better job of identifying multiple subjects in this way.
But it's hard to beat free: SD cant quite let me run it locally on my laptops's 2GB Ram GPU but I doubt it will take much longer for flexible distros of it to allow click-to-install versions that will run on most setups that have something resembling a discrete GPU.
It took me a lot of tries and about half an hour to generate a blue lemon on a (blue) marble countertop instead of a yellow lemon on a blue marble countertop.
I spent longer than I'd like to admit trying to get an image of a human running from a horde of zombies. No matter what I did I always got a zombie leading a charging horde of zombies.
I've completely given up on trying to describe layout of a complex scene, but that can be solved with img2img.
It is sometimes very difficult to get characters to interact with objects or each other in natural ways. And whether it works or not can be unpredictable. You can consistently get very common interactions like people dancing or riding a bicycle right. But good luck trying to generate say, a photograph of a man poking a sheep in the ear. Ok, I tried that one[0] and after a few minutes and about 16 generated images I actually got a few that were sort of accurate.
GPT-3, on the other hand, was completely mind blowing for me.
Hard to beat a high quality open source product. OpenAI missed the boat on "Open AI"
Specifically, signup process is: email, email-verification, create-password, full-name, phone. Leaving the process to try the login will return you to the request for a phone.
At least in the USA, getting hold of large numbers of phone numbers for free isn't easy.
Even for a one-time-use for verification, last I looked for one-off number verifications, it was $1 per authentication; to be fair, didn’t search too hard.
NOPE.
Stable Diffusion and its tools are evolving like nothing I have seen before. It's been ages since its been this exciting in new tech.
"We’re experiencing a temporary issue with signups due to a vendor outage. We apologize for the inconvenience!"
Turns out all that "waitlist", the ethics lecturing, letting in only bluechecks and the larping about how dangerous it is doomed your product in the end.
Hopefully the next time someone makes a tool as revolutionary as this they'll remember the mistakes of OpenAI.
Also, you can run it locally on a non-powerful machine. It just takes longer, but you can also just queue up as many prompts as you want and let your machine crank them out at its own pace. I use a first-gen macbook air m1 and it usually takes ~90s to generate an image with my usual settings.
Plus integrations into other applications like Photoshop and Canva. Open source has such a huge multiplier effect for innovation.
For anyone else stability offer a paid web version.
I get that competing with an open self-hosted alternative is a tough sell, but is this really different from other pay vs self-host scenarios?
I’d bet one you’ll hit 1000 generations much faster than the other.
Too late and stable diffusion works on my machine, I don't have to depend on anyone else to use it unlike this.