DALL·E 3
openai.com
openai.com
Ah - this isn't out yet. That puts this in the "announcement of an announcement" category (https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...).
Let's have a thread once the actual thing is there to be discussed. There's no harm in waiting (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...).
I wanted a way to experiment/see what DALL•E 2 could create and share with others as some sort of inspiration/a starting point.
This was before the API was available, so I had to generate and save them all manually. And it was rather expensive! But fun.
Looks like I'll have to update them all for DALL•E 3 when I get access.
Because there's no way to control the seed, a direct comparison (using a before/after slider, for example) probably wouldn't make sense. But I could put the group of 4 images from each version above/below each other as a general comparison, perhaps?
Even if it was the same seed, from my understanding Dalle3 would have to be just a further trained version of the same checkpoint to even resemble Dalle2's image. Like stable diffusion 1.4 vs 1.5 and 2.0 and 2.1 will make identifiably similar images, but 1.5 vs 2.1 vs SDXL won't look remotely similar.
Even more so because I'd wager they changed their encoder and/or decoder too.
* I think that if they generated something like a controlnet for guidance the same way in both models then they might be comparable but from my understanding Dalle2 doesn't work that way at all.
Comparisons would still be interesting though!
EDIT: it's only living artists actually that you can't prompt (hopefully, the article says so at least)
> ... in the style of a famous Spanish artist who was born in 1881 and passed away in 1973 [and a bunch of other shit about Pablo Picasso]
(I also notice that this is more verbose than just "in the style of Pablo Picasso", which probably helps OpenAI's bottom line given costs associated with token counts. I doubt that's their intention with the change, just something of note. And, of course, a living example would be more applicable for copyright issues but the idea is still demonstrated.)
But otherwise you're right — some might not work.
It's an interesting problem. Like, what's the point of conception for a work of art?
Looking at the images it's particularly interesting how it seems to have never once gotten the text correct, always just being a little bit off. Well sometimes way off, but mostly quite close.
Even then it's not perfect since I'm getting info off of the command you send, which may have fallen into whatever the defaults were at the time, and so when interpreted today, not easily possible to reconstruct the version/seed/etc. from that point in the past, if you didn't include it in the prompt. But still, I just like having a folder of 30k images that I can never lose, with at least the prompt, so I can go through and re-run them later (even manually) to get comparisons over time.
- ChatGPT integration is absolutely huge (ChatGPT Plus and enterprise integrations coming in October). This may severely thwart Midjourney and a whole bunch other text-to-image SaaS companies, leaving them only available to focus on NSFW use cases. - Quality looks comparable to Midjourney - but Midjourney has other useful features like upscaling, creating multiple variations, etc. Will DallE3 keep up, UX wise? - I absolutely prefer ChatGPT over Discord as the UI, so UI-wise I prefer this.
Currently with Midjourney/SD you sometimes get an amazing image, sometimes not. It feels like a casino. SD you can mask and try again but it’s fiddly and time consuming.
But if you could say ‘that image is great but I wanted there to be just one monkey, and can you make the sky green’ and have it take the original image modify it. Then that is a frikkin game changer and everyone else is dust.
This _probably_ isn’t the way it’s going to work. But I hope it is!
Generate an image, maybe even a low quality one. Fix the seed and then start iterating on that
As far as the OP goes, they claim you don't need to prompt engineer anymore, but they just moved prompt engineering to chatgpt, with all of the fun caveats that comes with.
Since gpt4 natively understands images, there's the potential for it to look at the image, and understand what about it you want to change
So long as it actually can create an image virtually the same but changed only how I want it. That would just blow everything else away.
I mean I guess it’s. Or that ridiculous. Generative fill in photoshop is kind of this, but the ability to understand from a text prompt what I want to ‘mask’ - if that’s even how it would work - would be very clever
Anything else that is 'open source' AI and allows on-device AI systems eventually brings the cost to $0.
Especially with DALL-E. Honestly I'd be more excited if MidJourney released something new. DALL-E was the first but, in my experience, the lower-quality option. It felt like a toy, MidJourney felt like a top-tier product akin to Photoshop Express on mobile, still limited but amazing results every time, and Stable Diffusion feels like photoshop allowing endless possibilities locally without restrictions except it's FREE!
Probably not. Bing Chat (which uses GPT-4 internally) already has integration of Bing Image Creator (which uses Dall-E ~2.5 internally), and it isn't good. It just writes image prompts for you, when you could simply write them yourself. It's a useless game of telephone.
I'm not going to post links, but there are several active projects & companies already doing that.
> Creators can now also opt their images out from training of our future image generation models.
So this version was again trained (without permission) on copyrighted work. And they try to shift the burden onto artists to manually opt out.
Aren't they afraid some court might, at some point, force them to pay each artist back a fee for each generated image?
So am I!
You also have the right to have your eyes open in a public dressing room, but you’re not allowed to turn on your camera and film.
You’re also allowed to have fireworks but not recreational C4s. And so on.
I'd say they're banking on the horse having bolted by the time such a thing might happen (i.e. courts would need to force 1000s of very large powerful companies to pay millions of people - an insurmountable legal effort).
So far there is no solid proof of that. They didn't disclose the sources or the methodology. Except for 'trained' and 'copyrighted' the rest is questionable. Otherwise they would be already paying royalties.
They could have used the output from prev version 2 with prompts generated by GPT, and then corrected by humans based on the produced image. Also they could use CV to analyze new/old images. I.e. if there is a new feature in the image add it to prompt and train again.
They are certainly striking under-the-table deals with big IP holders like Disney to not poke the bears, but leave all smaller actors defenseless (or rather penniless, more so than they already are).
Sam Altman has already empirically proven himself to be rich enough to be above the courts with the whole WorldCoin thing, why should he assume it would suddenly be different now?
https://www.google.com/amp/s/www.coindesk.com/policy/2023/08...
Which means there are countless (free) amazing tools around SD.
StableDiffusion is threatened by exactly nothing.
(others have mentioned that SD shall happily generate porn: I don't care about that... But I care about SD being the actual "open AI").
Images generated, with the exact same settings (including seed), on m1 laptop are not the same as the images from my nvidia GPU desktop with the SD-webui.
Well put. I was wondering how that aspect was going to be acknowledged & phrased here.
“so for everything else, there’s Midjourney”
Also, it's basically free if you own a Mac or an iPhone.
Just take a loot at civitai to see the kind of finetuned models that are out there.
They can't do that with MidJourney and DallE.
EDIT: for what it's worth, I'm not making NSFW stuff with MidJourney. I'm talking about things like being unable to use the word "cutting" or "slicing" because they could be used to make gore but I wanted "A stock photo of a person cutting cheese on a counter"
They mention you can disallow GPTBot on your site, sure, but even if you do, what happens if the Bot already scraped your image? In any case, probably other people would just publish your picture in some other website that does not disallow GPTBot anyway.
A quick glance at /r/Midjourney or even the images featured in DALLE link above shows how boring the “default result” is when using a generator. While it may be easier to create images, you still need some artistic sense and skills to figure out which ones are appealing. In the bigger picture I think this basically means that illustration-type art will become more of a curatorial activity, in which being able to filter through masses of images becomes the predominant skill needed.
ControlNet is a obvious counterexample. If you think "diffusion is just collaging", upload a control image using this space that cannot exist in the source dataset (e.g. a personal sketch) and generate your own image: https://huggingface.co/spaces/AP123/IllusionDiffusion
"DALL·E 3 is designed to decline requests that ask for an image in the style of a living artist. Creators can now also opt their images out from training of our future image generation models."
Very carefully-worded statement. So...still relying on people's hard work, but on the upside, you get to opt out of having your work be fodder for DALL·E4." </s>
* using living artists work to train models = good
* generating living artists work using said models = bad
Good ethical consistency from the OpenAI crew
This even applies if the AI copies an artists art style (in the same vein as a human looking at one artists art over a weekend and then being commissioned to paint something in the same style, which is completely legal since you can't copyright an art style; although Adobe would love that[0]).
(I would also argue that it learns and generates images in ways that are non-human, just based on speed and scale alone)
Also another thing that's been on my mind is I wonder if all this AI generation stuff could cause a Games Industry style crash where due to such a over saturation of highly advertised but meaningless/worthless AI content consumers lose interest and stop spending money in different respective industries (books, ganes, films, digital art, music, etc.) and then they crash.
https://investor.shutterstock.com/news-releases/news-release...
More transparency about the training data, as always, would be greatly appreciated.
You can't do that. It's copyright-maximalist copyright infringement.
or attribution even
Some folks seem to have some strong ire towards OpenAI (maybe a bit less recently), but for one, they seem to do a really, really, _really_ good job at making themselves "the benchmark to beat" for certain things, and in doing that, I think they really seem to push the field quite far forward. <3 :'))))
Yes, they should. OpenAI IS Microsoft, never forget this. Any old timer like myself remembers the crap Microsoft pulled in 90's. And nowadays they still would do the same (and in background sometime they still do it) if they would lead in those areas. I have no love towards FB/Zucky boi, but the move to make LLAMA free is a good one. Hopefully another leak comes from inside OpenAI and we get access to everything.
I think that might be a bit "enemy of my enemy". Remember "commoditize your complement"? Not that I'm averse to the tech giants forcing each other into a race to the bottom.
What is the meaning of this? Why is it part of your post?
Perhaps the Dall-E 2 unintentionally got that better.
In addition, having access to a library of prompts, and being able to produce, create, and store images within the web interface will unlock this type of generative ai for images to many more people.
Compare this to the midjourney way, in which users must not only sign up, they have to use a discord bot (not saying this is hard, but more so, a larger barrier to entry).
Native integration will mean instant adoption by millions on day 1.
Why is the spoon writing on the back of a clipboard, for example?
This is about DALL-E 3, which is just announced. Nobody's played with it yet so we don't know if it's a lot better or not.
> DALL·E 3 [...] will be available to ChatGPT Plus and Enterprise customers in October via the API, and in Labs later this fall.
The open source community is pushing forward SD forward far faster than Midjourney is improving.
You can see it at https://cosmictrip.space
(and coming soon is an adventure game series backed by DALL-E and GPT-4)
If you want to self-host, check out comfyui. It is a breeze to install, and offers an api for headless interactions. On my 5 year old i7 NUC it produces a 1024x1024 image (cpu only! no gpu needed) in around 20 mins using the SDXL model.
Wonder if DALL-E itself counts as a forbidden living artist or if soon we will „generate x in the style of DALL-E 2“
However...I'm a very visual thinker, when I'm thinking or speaking, I often see images in my head of what I'm trying to convey.
I wonder if aligning ChatGPT and DALL-E or something similar so that I can "see" an image of what the computer is saying as well as the text might be a great next step towards making me feel more engaged with the technology.
That and native speech to text would be nice so I can just talk at it and have it just sitting on the side as a helper-bot instead of being the main point of my focus while I work on or do other things.
You might know Thai already but have you tried using ControlNet to make images with stable diffusion? It allows you to input an image that it follows along with the prompt, and there’s a bunch of ways to influence it. You can even give it a hand drawn sketch. Or have it pick out position of limbs, or use the edges of objects and keep those.
If you have something specific you want to create then it’s amazingly helpful.
The only easy to use site I know of that offers it is happyaccidents.ai (or you can run yourself if you have SD installed)
This conversation reminds me of the old Star Trek:The Next Generation episodes in the Holodeck, only they're talking to "the computer" to iterate on a holodeck scenario design.
I think this might just be inherent in the domain - the state space for images is so much larger than it is for text, so there's just a lot more ways to interpret a text prompt. Sane "defaults" help, but it might just be inherently true that it takes a longer prompt to get close to what you're seeing in your head.
Too bad, since there might be some interesting advances (the way the model follows the prompts better for instance), but OpenAI is continuing to advance the tech behind closed doors.
They honestly could have just waited to announce this until it was actually released!
Would love to be proven wrong!
If you are actually a visual artist, I think the leader of the pack right now is controlnet, because you can exactly determine the visual structure of your image. While MJ or Dall-e may be better at "imagining" concepts, or have a more aesthetic sensibility (with Loras and custom-trained models, I'm not even sure about that) with controlnet you can very precisely specify how your image should be structured.
This is closer to how traditional artists work. They don't go for the details (color, texture, shading) first. They do a sketch: what are the big forms in this image? How do they fit together? Then they begin filling in intermediate details. What is the color palette? Where are light sources? Which areas have contrast? Which do not? Only after they have done all of that preliminary work will they actually implement the frills on the dress, or the twirls of the mustache.
If you are just a person who wants to make some pretty pictures, Midjourney (and, Dalle3, now) is probably your best bet. If you are an artist who wants to use an actual tool, you are using StableDiffusion. I think it's unlikely that the centralized "plug and play" Midjourney or OpenAI will ever be able to or interested in replicating the complex interface of stablediffusion. But there is a tremendous opportunity for a startup that can improve the UX of the complex workflows that are being developed by "ai artists."
That's also why I am convinced that MidJourney / Dalle will not replace artists. You simply cannot, with a single prompt, replace the work of a true visual artist.
They already are, because employers don't care about "true visual art," they care about cost and productivity, and getting an intern or someone outsourced to write prompts is both cheaper and faster than paying an actual artist, and capitalism dictates the path of least resistance is the path all competitive business must take. Companies are already replacing their creative staff with AI, or are planning to. AI generated art is already everywhere in advertising. And yes, they contain obvious errors that wouldn't exist with a real artist. And no, companies do not give a damn.
> MJ is already better than SDXL if you don’t need any of the [elements of the rich ecosystem beyond the base models]
Yeah, I think you kind of missed the point there. (Also, not convinced you are right even there, MJ seems to be much worse at prompt-following than base SDXL model, and on other qualities in the range where subjective opinions are going to vary considerably on which is better, judging from the head-to-head comparisons with prompts I’ve seen, largely from people claiming that MJ is better so presumably not trying to subtly favor SDXL in the construction of the comparisons. Because of the ecosystem, its been a long time since I found MJ more useful than even SD 1.5-based toolsets.)
1984, the year the Mac was introduced (and when the more customizable, nitty-gritty Apple II series was their main seller) was also Apple’s all time peak in share of the personal computer market in terms of units sold.
So…yeah, I think the comparison is apt, but maybe not in the way you think it is.
Also, cherry picked examples from Dall-e 3 may not be representative of the average output. Like some SD 1.5 models may look amazing on civitai or reddit, but you soon realise that they are terrible on average and overfitted to very specific kind of pictures and characters.
There's one image with "Explore Venus", and in the video, the hedgehog has a mailbox with Larry on it. Both of those look good, but obviously super cherry picked.
There can be only one!
This seems weird to me, but I admit I'm about as far from an artistic person as it is possible to be. I understand why it was done (people kept asking for pictures in the style of that one dude and they were better and he hated it) but it still just seems strange.
The content that the people themselves posted to the Internet? I'm pretty sure that OpenAI didn't break into any artist's studios and steal anything.
Do art students steal when they tour the museum? Please, I really need to know... what sort of dingbat philosophy is it that thinks that this even slightly resembles stealing, in either of the "copyright infringement is stealing" or the "actual theft" meanings.
I hope the DSM VI includes copyright maximalism in its list of mental illnesses.
The path way to an OpenAI monopoly is quite clear, especially with the controlling stake from Microsoft. So I won't be surprised to see OpenAI continuously attempt and revive their regulatory capture using licences [1] against actual 'open' AI companies who release their papers, code, models, etc.
[0] https://news.ycombinator.com/item?id=35960125
[1] https://www.reuters.com/technology/openai-chief-goes-before-...
Does this mean that ChatGPT plus price will increase? Otherwise the value you will get for the subscription is crazy!
“An expressive oil painting of a basketball player dunking, depicted as an explosion of a nebula.”
> When prompted with an idea, ChatGPT will automatically generate tailored, detailed prompts for DALL-E 3
Then there’s a video of someone typing into ChatGPT and it responding with images
Plus users will be able to directly create images within the ChatGPT web interface.
$ http https://openai.com/dall-e-3
HTTP/1.1 404 Not Found
Cache-Control: no-cache
Connection: keep-alive
Content-Encoding: gzip
But others seem returning normal 200.https://github.com/lllyasviel/Fooocus
Other stable diffusion UIs have this as an option too.
I'd love to see some basic usage included in my GPT+ sub
edit: Welll... they say "Try ChatGPT (DALL·E 3 coming soon!)" ... if they are including some DALLE usage in GPT i'm going to resub instantly. So huge.
Really accelerating AI development, with humans in the loop for now.
Nice! I teach ChatGPT on midjourney’s documentation and it gives me great prompts on my loose ideas
This follows that concept
What does this mean? How do you use ChatGPT Plus through the API?
Anyone have any tips on getting beta access to these models earlier? I've been in GPT3 beta since June 2021, and I was only ever able to get Codex early.
How generous.
who at OpenAI decides the right and wrong??
Also at first this objection really resonated with me. I think the meaning of OpenAI has spread pretty well now, and that it's getting to the point where raising this objection is tiresome.
There is an important point to be made about how it got popularized as being open and then they went and closed it while keeping the momentum, but that should be made instead of just saying, "wait, it's called OpenAI but isn't open?!?!"
However, stuff gets started as an open play all the time and gets closed, without open in the brand name, for instance https://ghuntley.com/fracture/ - hence the term https://en.wikipedia.org/wiki/Openwashing
I can actually appreciate “open” as “open to access” or “open to actually having a product iso posturing about being so far ahead but never releasing anything worthwhile” (looking at Google).
Edit:
OpenAI is a corporation and their stakeholders include Microsoft, Peter Thiel, and Infosys.
Yeah, not really surprising that they are not "Open".
Don't get me wrong, I love open-source, open-weights research, but the elephant in the room is that people aren't willing to do that on dirt-poor postdoc salaries anymore, for good reasons, especially when greedy landlords are now charging upwards of $4000/month just to have a reasonable living space, and the government takes close to 50% of your salary.
Years on, it's still a little hard to fully grasp the imminent, momentous impact this (has yet to have?) on commercial artists. I fear it will become pretty much impossible to make any sort of living off art in the next decade.
I mean: the outputs on the page are just awesome. Leagues ahead of the stuff we have now. And I'm already seeing old-gen generated AI images on corporate blog posts—everyone will jump on DALL-E 3.
Being in the profession right now must be very discouraging indeed. My heart goes out to those artists who will eventually be replaced by cheap, intuitive text prompts.
Commercialising art for corps was one of the last ways to exist as an artist in today's economy and get by. I fear the extinction of the profession will have a big impact on our cultural capital.
Software developers have been having technology completely eat their work out from under them since the dawn of the industry. But jobs aren’t being lost over it; there’s more demand for software developers than ever. When software advancements reduce the work needed to produce output, the world has moved on by taking that as the new baseline which software developers build upon and demanding more software to be built with bigger and better features.
Back when I first started working, it was somebody’s entire job to take pages and pages of content and transcribe them into HTML manually so that they could be published on a website. People used to do that all day, every day. Then CMSs came along and completely eliminated that work. Now sure, if a developer decided that all they wanted to do was write static HTML and refused to adapt their skills, they would be out of a job. But we all used the new technology to build more dynamic websites that provided more value. The new technology didn’t take away our jobs, it provided an opportunity to do a better job.
It’s the same with this. These aren’t tools to replace artists – these are tools artists can use to do more and better work. These aren’t tools that will reduce demand – when everybody can get on-demand, totally custom artwork, people will want more of it, not less.
What you're talking about is illustration, which will indeed have a difficult time in the future.
It would be a terrible outcome if authorship and illustration are mostly reduced to editing and touching up errors in AI-generated statistically likely art.
(Although oddly, for programming, I'm really looking forward to that outcome).
I can easily imagine a future where a single dedicated individual or at least a very small independent team can make a full-length movie without leaving their apartment on a tribal budget that rivals a big Hollywood production costing hundreds of mullions to produce today.
Right now you've got the writers striking, worried that the studios are going to replace them with AI, I think this is totally backwards, it's the studios who should be worried because the barriers to entry that protect them today are about to come crashing down.
I am already seeing a few people make short films using these tools. Right now anything "AI generated" has a certain novelty factor, like back in the day when 3D CGI was new and we were all rendering chrome 3d spheres and shiny red cubes and cylinders on black and white checkerboards, this phase is going to pass soon enough. Perhaps that surreal Midjourney glow will be part of 20's nostalgia in the decades to come. There's a whole new set of skills a new generation of artists are going to master and do things largely unimaginable a year or two ago. They're going to make art that expresses their own perspectives and ideas and not just what's currently allowed by the current consensus, just as the artists before them did.
Making a living from art is something that the majority of artists even before GenAI did not reach.
I agree that it will change things a lot, but fearing the "exctinction of art" is a bit dramatic IMO.
It is all just impetus to go and find the prior form of "art" we have had all along. Art need not be about the Artist and her Works, or about creating something for some kind of vague consumption. Art can be about collectivity and shared truths. Art used to be about speaking to God or whatever, and whoever actually placed the pieces of glass were very secondary. At my most optimistic, I feel like some turn like this (but probably more secular) is inevitable. People need to get stuff out, and if this pressure is not relieved by our current ideas of labor and such, it will find a way nonetheless like water through stone.
Does a new form of “art” evolve that makes use of these seemingly omnipotent brushes?
I think the market for affordable prints and copies will definitely suffer, because I can create prints of Midjourney art that looks as good as anything I see in our small local galleries and gift shops.
Also, it would be nice if there was something akin to a watermark that was perhaps invisible but upon inspection using a certain tool will reveal if it was created using the Generative AI model (similar to how real currency can be inspected and differentiated from fake currency notes)
I also think the output isn’t yet good enough to be usable without some human intervention (fixing hands, spot treatments, etc.)
And so continues the trend of "progressive" AI companies deliberately handicapping their models for no real good reason.
https://www.pcmag.com/how-to/pornhub-blocks-access-utah-here...
In the end, this is about liability, not morals.
How is blocking porn a "progressive" thing? aren't conservatives the ones blocking porn these days?
for example: https://www.rollingstone.com/culture/culture-features/republ...