DALL·E: Introducing Outpainting
openai.com
openai.com
https://old.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
It seems more like they were trying to accomplish the opposite.
https://www.technologyreview.com/2020/09/23/1008729/openai-i...
And being so open they first lock the model up and charge a fee, so anyone can pay. Just spreading democracy through paid API calls. /s
I was a bit mean, they did kick the field in the butt and pushed us ahead even with all the stubbornness and secrecy. But now they are just holding us back.
You might think that the nuclear bomb is a good analogy to use here, but it is not, because once the field has advanced to the point in which one group can develop AGI, it is now possible for other groups to develop it with relative ease, unless you actively take over the world first and deny those other groups the compute resources necessary to train/run AGI.
The point is, once these algorithms are upon us, you must be willing to accept what impacts they will have, even if it destroys entire industries. The alternative being that you destroy the industry slowly rather than quickly, while simultaneously widening the gap between the elites and everyone else.
The mistake is thinking that people can't adapt to the times, which is only true if you are actively holding them back.
If someone developed AGI today, the best thing to do would be to instantly throw up a torrent of it and spread it as fast as possible, because if a sole entity is able to get it first and kick the ladder away, we are most likely screwed.
I still don't understand what this would mean. Where are all of the terrible things that were supposed to happen, now that Stable Diffusion is available?
We've been able to create completely photorealistic fiction for decades now. See any movie with CGI for an example of whole worlds, and people, that don't exist. The bar has gradually been lowering (see the amazing CGI that YouTubers do these days), and now maybe there is a bit of a step function down, but being able to make things that aren't real isn't remotely new. I don't understand the fear.
The danger for society is in not already knowing that is the case, since it's relatively trivial, without AI, to make fake content.
We are way past that. If anything, trivialising and democratising the generation of fake content, will educate the masses. Perhaps, bring back the trust for science, which we need a lot.
Obligatory XKCD:https://xkcd.com/2650/
I'm pretty sure one of the primary arguments for Dall-E and Stable Diffusion existing is that there are lots of people who can't draw anything.
On one side there's no problem to like CP if it's AI generated and no real kids are harmed.
On the other side it may desensitize the public to this type of content.
> I still don't understand what this would mean. Where are all of the terrible things that were supposed to happen, now that Stable Diffusion is available?
Mainly people making porn (e.g. stuff like deepnudes). It seems like a lot of work has gone into into preventing that (e.g. filtering porn out of training data, having porn-detection models to block porny output). There's also been a lot of talk about political fakes, etc, but I'm not sure how likely that is to actually happen at this point. I think one of the "selling points" of limiting access to DALL*E was that they could revoke access to people who they deemed to be misusing it.
I wonder is it possible to distribute the task for volunteer's idle GPU (like BOINC)
Also you’re probably not seeing all the pedo content that people are already generating for themselves.
I think that's a good thing. A slow creep is dangerous, because people may not notice, and would be tricked. An explosion of insanity makes everyone realize they can't believe everything they see on the internet.
We'll get more Q-Anon than Wikipedia.
Any favorite examples?
The former is being used as justification for the latter.
Paternalistic moralizing as a method to discriminate who gets access to models. Everyone else gets these cloud-service table scraps. That's why Stable Diffusion is so awesome -- YOU have the model!
>leveraging it for money
Funny thing is that people didn't stand still and invented diffusion and other CLIP guided image synthesis methods, and DALL-E 2 copied the method, completely changing from the first architecture.
Their arrogance is that they think they can ride the dragon. They want to be the ones to discover, advance it, and control it. But everyone else doesn't have time for that shit.
Of course maybe it's just because text is less visually impressive than images.
BLOOM incorporates some new ideas like ALiBi which might make it better in a more general sense. They haven’t released official evaluation numbers yet though so we’ll have to see.
and it shouldn't be difficult to pick which of these is actually true.
If they hadn't called it OpenAI, the conversion from open to corporate would be much easier, where's now it's a significant pain point in OpenAI's reputation. It's kinda nice to see the original vision still somewhat having an effect despite no one left to propagate it.
Sam hasn't been at YC in years and (based on anything I've seen) isn't active in YC at all. As for "YC encourages good CEO-community relations", I have no idea what that means* but it has nothing to do with HN. We encourage good content-community relations and that's it.
You have a long history of posting dark insinuations about YC/HN, not to mention nagging the mods about how bad we are and how much better you yourself have done the job in the past. I mostly let the latter go, but when you start with the ethical insinuations, that gets my dander up. It's time you stopped smearing people's reputations on HN. If you have evidence of wrongdoing, post it—I'm sure the community will be extremely interested. If you have no evidence, please stop from now on.
(Edit: I realize it probably sounds like I'm over-reacting to the parent comment, but this has been a longstanding pattern. We can cut people slack for years, but not infinitely.)
OpenAI stuff and Stable Diffusion stuff (and DeepMind stuff for that matter) are all popular on HN because the community is super interested—that's literally it. We're not pulling strings or playing favorites (we don't even have favorites in that horserace, at least I don't). As a matter of fact, the last thing I did before randomly running across your comment was downweight the current thread because of the complaints at https://news.ycombinator.com/item?id=32665587.
* unless you mean that we advise founders about how to write content that actually interests the community—that we do, and not only YC founders but non-YC founders, open source programmers, bloggers, and anyone else. That's all a consequence of wanting HN to have good content and seeking to avoid the boring stuff. By the way, I'm working on an essay about how to write good for HN and avoid boring stuff too; if anyone would like to read it, email me at hn@ycombinator.com and I'll send you a copy.
Your w0t m8?
That's my username, and I'm also named Daniel. This is a conspiracy.
EDIT: just realized.. what if dang was an advanced AGI the entire time and HN has been one giant turing test to see if any of us would notice
"Use Them or we will lose Them!"
I used Twixt (An abbreviation of Betwixt) to replace inbetween in a submission.
Edit: to fix formatting and spelling
https://www.reddit.com/r/StableDiffusion/comments/wzj8kk/a_c...
Since DALLEmini and DALLE2 are more "creative" (since they use a better text transformer) you can use them to generate the input and SD to refine it for more fine detail.
https://www.reddit.com/r/StableDiffusion/comments/x209sb/pre...
Various (subjectively judged) SOTAs on at least some subset of at least this family of tasks is changing somewhere between daily and hourly right now. I've been watching this stuff closely since fairly early ImageNet days and I've never seen a Cambrian explosion of "how the hell did that do that?" events at anything like this cadence.
If researchers have great products or findings to show off, it's easier than ever to simply publish them somewhere online and let impressed audiences spread the word. Gatekeepers have been irrelevant to truly great science for a long time.
It's mediocre science that needs gatekeepers to distinguish it from not-even-mediocre, truly substandard research.
They’re just the only ones calling themselves “OpenAI”.
Even if you don't want to take the 30 seconds to set it up in a free Google Colab environment, the paid DreamStudio version is still half the price of Dalle.
Amusingly it's more open in every way.
https://github.com/hlky/stable-diffusion is another popular and good tool.
I can think of ways to fix Stable Diffusion since it's open-source. I think I could bridge the gaps as I see them in about a weekend of hacking. I'm not sure when I'll get that weekend.
(Footnote: What I want to do is not something I can explain without a technical blog-post-length document or a zoom call; it's about the same level of complexity as the other major SD hacks we've seen)
That said, for me, I agree that dalle does much better pencil sketches.
According to Google's Imagen paper this is literally because the pixels get multiplied by the cfg scale and start clipping; they have a technique called dynamic thresholding that replaces it. Not sure if SD uses this, but I saw Emad hinting they were training an Imagen model…
Anyways I think it’s way too early to start taking sides. I enjoy using all these system.
I have no examples but imagine things like at the top of his game are immensely problematic, albeit not very visual to begin with.
Sure, both SD and Midjourney produce absolutely beautiful artworks most of the time. But if you want something specific and out of the ordinary it takes a lot of attempts and promptcrafting (and sometimes you are unable to accomplish what you want at all).
However, my experience is that these prompts (which SD/MJ struggles with) often produce good results in Dalle2 even on the first try.
Of course, OpenAI has very limiting content policy. But if I have something very specific in mind and it passes their rules I currently chose Dalle-2. Even though I've spent much more time with SD.
I expect most people aren't going to be generating images all day, so using a cloud-based service for occasional use will still make a lot of sense.
Stable Diffusion offers a paid service to do this too, and there's nothing wrong with that business model. Prices will probably come down, though.
I'd worry more about, say, keeping your photos on Google and losing your account somehow.
I didn't investigate any further because I'm not actively trying to generate porn. It was pretty annoying to have my results blanked out because they were (apparently) NSFW - so either the filter was triggering a lot of false positives or the model was generating NSFW content for non-NSFW prompts.
To remove the filter, you just have to comment out the two lines of code which call it. Instructions are here: https://reddit.com/r/StableDiffusion/comments/wv2nw0/tutoria...
I don't want to be a doomer and have have surely unknowingly misused terms as well, but its definitely noticable how these originally clearly defined terms are getting used in entirely new ways.
And it's not just with technical terms like this, it also applies to originally obvious terms such as racism, sexism etc which have lost their original meaning entirely
For example, the word "term" comes from the original latin "terminus" that means "end" or "boundary". It only got the meaning you used it for centuries after it was first used in English. See: https://www.etymonline.com/word/term
I was just pondering that our originally clearly defined terms are rapidly getting used in very confusing manner, which increases the difficulty of a discussion, as participants interpret words very differently.
I dont think that people look up the actual definition of terms in a thesaurus anymore. They hear it in some context and create their own personal definition. It wasn't as obvious before the internet i think, but nowadays everyone is bombarded with technical terms all the time, which likely contributes massively to this increasingly fluid terminology
The race to the bottom refers to a competitive situation where a company, state, or nation attempts to undercut the competition's prices by sacrificing quality standards or worker safety (often defying regulation), or reducing labor costs.
This appears to the the exact opposite, a race to provide more services and more features for a lower price, based on optimising and/or lower profit margins. AKA Capitalism actually working for a change.
I have been thinking the same thing, it's sad Steve will not be able to see it
That is still the best you can do... which happens to be about the best I can do! Just like my introduction to the computer at a young age has atrophied my handwriting quality.
There does seem to be a lot of vague angst about how this will affect the nascent "Prompt Engineer" career track, but I hope most are comfortable letting the open innovation play out a bit before trying to personally monetize it..
In this context it's a good race. This software seems to have caught fire and tons of people are playing with it and providing tons of crazy new tools for cheap or free.
It's a race to the top for us.
You're forgetting (or not knowinng?) NVIDIA Canvas that came out one year ago, give or take. Literally turns stick figure complexity drawing into photorealistic stuff.
Reference, for those who haven’t seen the original joke to which my joke was referring: https://www.reddit.com/r/pics/comments/d3zhx/how_to_draw_an_...
(See also: https://knowyourmeme.com/memes/how-to-draw-an-owl)
It's going to get so cool and interesting, I think.
A lot of the conversation around art may focus more on composition and objectives of the artist in the new prompt engineering world, with less bias from factors such as rendition quality etc. creeping in since it's so incidental.
New forms of art will emerge and/or gain popularity that focus on trying things the tools aren't good at yet. The human artist of the gaps. The niches will constantly be shifting.
I wonder if we'll learn to recognize the output of certain popular models and perceive them as instruments. "Made by xy on z" instead of "xz on guitar", so to speak. I remember the 90s/early 00s internet when it was always easy to tell when something had been done on Flash, just because of its line anti-aliasing rendition style being so distinct and familiar.
The novelty will wear off, and we'll all start to feel a bit disappointed that the average human's imagination is pretty limited and novel/original ideas remain somewhat rare as the patterns and tropes in all the generated art emerge. It's great you can put the space needle where you want it and get a good-looking city and space ship, but how many variations of a cyberpunky skyline with a space ship do you need? And then we'll celebrate the novel stuff that does happen, as always. I suppose the tropes will evolve faster as the throughput goes up.
>rendition quality... [is] so incidental.
There's this thing in painting called 'mark making' and it can be the difference between an all-time-great painting and a throwaway portrait. Mark making speaks to every momentary choice of physical process a painter employs and reveals their thought process. For some of the greatest painters, it reveals their genius.
Do not discount execution. Overlooking "basics" and "mechanics" is what results in disappointing work.
There's a lovely documentary called "Tim's Vermeer" about Tim Jenison's - one of the founders of NewTek, the people behind Video Toaster and LightWave, incidentally both tools that made hard visual art tasks accessible to wider audiences - hobby side project to prove that Vermeer used sophisticated optical tools to capture and copy his scenes from physical sets, rather than e.g. paint his famous grasp on lighting purely from his own mind. He builds such tools himself and then proceeds to successfully create his own Vermeer-alike painting, despite possessing very artistic skill himself.
It's full of good ruminations (and good at sparking more) on tools-vs-artistry but also execution-vs-method, and whether designing and adopting innovative tools and the tedious process to use them made Vermeer less of a genius, or just a genius of a different kind than otherwise presumed.
It's very accessible and doesn't require knowing anything in particular from the art world.
Hockney, one of the featured 'expert painters' is a hack who doesn't actually know how to paint* and therefore claims that certain gradations are certainly impossible without some sort of additional lens device. Meanwhile there are 19 year olds at the Grand Central Atelier pulling off just that.
I own a camera lucida. It got in the way more than it helped. At best, it's a novelty, now collecting lots of dust. Vermeer probably had one but it's altogether way more likely that he just had a well-trained observational/representational faculty. There are some killer painters using cameras now (Will St. John for one) but they typically have a decade or more of very rigorous direct observation to lean on.
I suggest following Ramon Alex Hurtado. IMO he's one of the more exciting young scholars on historical representational painting techniques. I don't think he has written anything himself yet, but he does do workshops and has a big informational update for his website coming.
*the definition of painting is now so broad it is meaningless. Here, I mean "attempting some degree of visual accuracy" which can be achieved in endless creative ways. Compare (easily on instagram):
- Colleen Barry - Peder Mørk Mønsted - Cecelia Beaux - Jas Knight - Felicia Forte - Eric Johnson - Ksenya Istomena - Sergei Danchev - Glenn Dean - Blair Atherholt - Jose Lopez Vegara - Hongnian Zhang - Hans Baluschek
These artists all have their own voice and stylistic choices. They also all represent things they see with some sincere accuracy. Look up Hockney's ipad paintings he got lauded for. People treat them like they're some misunderstood genius, but really, they're just bad paintings.
I'm sure he's a sweet old man and I'd drink tea with him. But if it weren't for his ilk I might have found proper instruction 10 years earlier in life. Modernists and postmodernists robbed generations of proper art instruction. Imagine if all the music teachers burned all the sheet music and refused (or forgot how) to teach the diatonic scale. "Hit the keys in a new way! Don't let yourself be bound by conformist ideals!"
P.S. I'm not fun at parties
That is what really bothered me in art lessons in high school. When discussing any famous work it was always about concepts, ideas, composition,... and execution was very much secondary. But for your own work all that is completely ignored if your coloring is just slightly uneven or lines are too rough. If you could hand in a photorealistic drawing of anything, no matter how boring, that would give you much higher marks than a rough drawing of something worthwhile.
The true artists in this one are the coders, no doubt (corrolar to the inteligence debate).
On the other hand, you mention an important point with layout but you underestimate the progress these days. Surely there are companies who are working on automated design beyond CAD (computer aided design), eg. for specialized antenna.
> we'll all start to feel a bit disappointed that the average human's imagination is pretty limited and novel/original ideas remain somewhat rare as the patterns and tropes in all the generated art emerge
Well, one might argue that Richter's most highly priced piece looks a little like prehistoric art of the pleistocene. It's a little vain to mention it, because I can much better relate to the more basic form, of course. A more frequently sore point would be the pop music industry between professionals and the amateurish.
Anyway, this may be thinking too big. For the time being, the bunch of techniques is better understood as a toolbox, because it will be a long time before it trumps demo-scene productions, for instance. Here it is the technique that counts more often than not. The rest is an acquired taste.
https://old.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
Quite funny to see OpenAI panicking and falling on their own sword, as they were supposed to be 'Open' in the first place and are now being disrupted by open source.
I'd argue they're imperfect, but they don't look like arses. Big gap between the two, too.
https://projects.propublica.org/nonprofits/organizations/810...
100x seems like a pretty generous cap to me.
"Open" is not well-defined.
I'm not saying I agree with them, I'm a software person and open means to me what it means to you. I'd prefer if it was truly open by our definition. But I don't think we can bash the choice of the word so easily.
Well, I can't speak for anyone else, but it makes it wrong for me (and judging by the number of upvotes I got from the original comment, a lot more than just me).
It's not just the code left to an exercise to the reader...it's the training sets. You don't get to trade on the suggestion of "open" while keeping everything closed. They aren't idiots, they knew exactly what they were implying when they picked the name.
What if the training sets are proprietary (as in, they don't have licence to share)? Should they keep the research to themselves just because of that? I don't think that's better than not sharing the training set. It also doesn't mean the research is invalidated - find your own pictures and it's going to work. Same as - find your own plants and it's going to work.
TBH, I just don't see the training set as part of the research... In my case, I'd feed it tons of electronic circuits to try to teach it generate some. Why should I care about some random other pictures? I care about the research and I have my own training sets.
There's no CSS here but you can clearly see the MediaWiki template: https://web.archive.org/web/20090422103636/http://www.wikile...
I agree that Dall-E isn't obsolete. I'd also add MidJourney to that list. All three are great models in their own right with their own pronounced strengths and weaknesses. But when it comes to enabling novel workflows Stable Diffusion seems lightyears ahead of the others.
That's not what that means.
In case it is not obvious, every time a user generates an additional section of an image using the outpainting feature, it costs a credit.
And while OpenAI released this feature now, it's probably just a matter of days until even better features built on Stable Diffusion will be released, given how much community energy is focussed on it right now.
Not sure I'd want them in 16:9, hd 4:3 like the other HD releases of TNG and TOS would do me. I understand they shot on video so an official true HD remaster is likely to never happen.
1. https://www.extremetech.com/extreme/324466-tutorial-how-to-u...
But it did use a lot of early CGI that would need to be redone.
Do the black bars actually bother you that much? You know there are cropped 16:9 widescreen versions of some of these shows (which I personally detest, but I work in the business of moving images).
Genuinely interested in why this bothers people.
I think we still have a ways to go before the results would work well without looking like those nightmarish deepdream videos where things are constantly shifting, so society will have plenty of time to discuss the merits.
It’s not perfect but if the filmmakers thought of it it can be ok.
6 months from now, how we interact with these models will probably look entirely different.
http://elateachers.weebly.com/uploads/2/7/0/1/27012625/virtu...
Doesn't mean I'm not excited though. This kind-of feels like I'm watching the camera or printing press being invented. Everyone is comparing it to fine art, but I think ultimately it's going in a different and bigger direction.
1. How long would it have taken me to get good enough at painting to exceed what I generated in under an hour? How many people have the motivation to spend that time?
2. How much would I have had to pay an art student to make a painting better than what I generated in under an hour?
Ten million sub-par Monet knock-offs didn't exist, but could exist very shortly at minimal cost. Even if it never gets any better this is already potentially disruptive, and the models are getting better every month.
https://www.vice.com/en/article/bvmvqm/an-ai-generated-artwo...
This picture is also unusually coherent for Midjourney; if you just ask for a 16:9 image the sides tend to evolve into totally different pictures.
I do wonder how expensive the outpainting is. I'm assuming that each additional step in the timelapse is a full generation, in which case ~15 generations is about $1 total.
Imagen isn't special anymore.
Organizational scar tissue making them more risk averse about the PR risks of letting the genpop use AI generation tools, and create something offensive. With the safe assumption that Google will get blamed, not the user.
Fear of government regulation on AI if they don't self-regulate.
No need to actually release it, since this isn't the core business but just research. (While openai needs to actually create the business.) Corollary: more to lose -- a scandal around offensive content will not hurt openai's non-existent other businessess. It might make some advertisers pull their ads from Google.
The opportunity cost of building a self serve platform is too high. (Can't pull in people writing those kind of apps from projects with more commercial importance. Can't make the ML researchers do that.)
They misjudged how much demand there would be, and thought that building a platform would not be useful for a few years. And if it now turns out to actually be a great business it'll now take them a year to productionize and build a platform.
Their compute requirements are so high that selling access is not viable, the costs are prohibitive for real users.
It's not that different from e.g. self driving cars. Pretty obviously they had better tech from early on, but were not willing to take the risks that Tesla was.
I imagine there's a market for a stock-photo style service that has a very large number of good images for very diverse topics, but DALL-E etc are a bit low level for that, there's a lot of product development that needs to happen on top of that. A stock photo service doesn't seem to be the sort of thing that Google would get into.
Maybe it's an art-helper plugin to image editors? The Stable Diffusion based plugin sounds promising for this, but will _artists_ want to use it? Surely the point of art is (except for art making A Point) that the artist produced it with their talent? If you're not trying to make Art then maybe you need the stock photo service.
Perhaps I'm just not very imaginative, but I can't think of any use-cases that aren't either extremely niche, or are better served by a higher level style service that happens to use a model like this under the hood.
But my guess is that the market for “programmatically created good enough images quickly” is larger than the market for “inspired, perfected hand drawn digital images”.
Regardless, this is a great way for people to fight the lack of detail in Dall-E which I think is one of it's largest flaw.
Obviously, it’s going to be an incredible boon for content creation. I suppose that in the future it’ll make creating videos an order of magnitude easier, which will allow a single person or a small team to make a high quality movie where all the assets are generated, so that’ll really give us an eye into a lot of people’s imaginations, for better or worse.
To leave a thought provoking example, what’s going to happen when every adolescent has the ability to make a convincing deepfake?
It’ll put nation states in a similar position than they already have with crypto, where they wonder if they should ban, or regulate… doing nothing wont be an option.
Seems like it is currently not working on their site.
Go from Hello to any topic, and see how it responds. Trained on everything ever written.
Seems like it would/could generate realistic discourse. Is that possible, to work with words instead of images?
I've used DALL-E/DALL-E-2, but have yet to try Stable Diffusion. Can someone give me some insight?
Img2img mode where you provide SD a sketch and a textual description of what you want to achieve abd it figures out result is a killer!
https://www.reddit.com/r/StableDiffusion/comments/wzlmty/its...
Krita plugin: https://www.reddit.com/r/StableDiffusion/comments/x209sb/pre...
I uploaded a digital painting, selected "Edit mode", added a generation frame and prompted "complete the painting in frame" ...but it just added a completed unrelated photo related to painting in that frame.
You know companies sold software that you paid once, and then ran as much as you wanted on your pc?
Putting "Girl with a Pearl Earring by Johannes Vermeer" in the kitchen in 2022 does not look good!