Stable Diffusion is a big deal
simonwillison.net
simonwillison.net
But two things I've noticed:
First, artists will still have a massive advantage over non-artists with this tool. A photographer who intimately knows the different lenses and cameras and industry terms will get to a representation of their idea much faster than someone without that experience. Without that depth of knowledge, someone might have to rely instead on random luck to create what's in their head. Art curators might be well-positioned here since having a wide breadth of knowledge and point of reference is their advantage.
Second, we need the ability to persist a design. If I create a character using SD, I need to be able to persist that character across different scenarios, poses, emotions, lighting, etc. Based on what I know about the methods SD/Midjourney/Dall-E are using, I'm not sure how easy this will be to implement, or if it's even possible at all. There will always be subtle differences and that's where being an artist who can use SD for inspiration instead of merely creation will retain their advantage over a non-artist.
That said, holy crap. This tech is insane.
See yesterday's Stable Diffusion article: https://news.ycombinator.com/item?id=32643564
I think that thing will neccesarily contained representations of dimension, behavior (physics/bones), and "style". Without the "ai thing", if only using an image/text, the character would have to be impossibly represented in the model, so it could guess all of these things predictably. For example, what that character looks like from a side profile, or behind. What if it's an alien, and its arms should always bend backwards? Could a text representation ever be made to completely describe this, with good reproducibility? Probably not. But, I assume some non-human representation would have a better chance.
As is, if something known is required, I think the behavior of these models can be considered "destructive" to the input image, more often than not. For this reason, I think artists are safe, for the time being. :)
[1]: Multi-View 3D Face Reconstruction with Deep Recurrent Neural Networks [2]: Deep Neural Network Augmentation: Generating Faces for Affect Analysis
You could use Stable Diffusion et al to create new characters based on a prompt, then farm the concept out to artists to produce individual works. Kind of like hiring a super expensive agency to design your new logo or brand identity, then using a stable of in-house designers to translate the concept into UI, ads, etc.
Any signals humans give (painting x is better than y) is another signal to encode. Take billions of such ratings and improve the AI’s taste to superhuman levels.
In short, anything that humans would add to artistically improve the outcome is just another signal to be encoded. It’s weird to write it, but artistic creativity is deciding what new pixels go where, which is a search problem (in a large search space) which AIs are apparently doing great at.
We have a bias: we’re humans, we must be important somehow! But it comes down to a bigger neural net eventually outperforming the one in our heads.
A neural net that can communicate with intent to human minds as well as a biological human is nothing short of strong AI. We’ll get there, but not with generative models.
They did this for Backgammon and found styles of play humanity hasn’t discovered in millennia. Now humans can use those when playing each other and it makes an old boring game feel fresh and exciting.
In video games upon release everyone is bad and its an even playing field. But as people practice and gain experience they improve their usage, refine their approaches. You eventually see metas develop and best practices.
This will all be exactly the same. Just give it some time and watch professionals create professional outputs compared to paying for some base level work on Fiverr or whatever.
How's the story there?
I think Copilot is going to live off hype for a while then tank and be looked back on as a failed experiment. Whereas I think that this kind of AI will eventually get to a point where it's extremely useful and could change up certain industries (game assets, marketing materials etc).
I mostly do maintenance of legacy codebases (also known as codebases, lol) where a lot of the work is figuring out where the changes need to be made and actually making the changes is frequently just a few lines here and there.
When I do have to figure out how to use some API, it's often not an open source one, so Copilot would not have it in its corpus.
I think these kinds of conditions are really common since software tends to last for maintenance longer than it is in initial greenfield development.
So I'm confused what kind of work benefits from Copilot. Just pumping out greenfield development of new websites/webapps that don't use much legacy or closed source code or services, just using existing popular open source libraries in commonplace ways?
The other thing I wonder about is code quality. When I look up API docs and stackoverflow examples, I get to read them all, maybe test some examples out in a CLI/REPL, and then decide carefully exactly what to do myself, what special cases to worry about or not, what errors to handle, etc.
Maybe what I end up writing is even the same as Copilot would have written. But in the process, I learn about finer details of the library and make detailed decisions about how to deal with rare edge cases. Might even end up writing a comment calling out a common pitfall I realized might exist.
My question is -- in order to save so much time with Copilot, are you still able to do all this extra thinking and deciding and learning (in cases that warrant it)? Or would doing that just end up consuming most of the time Copilot "saved"?
In other words do you end up producing code much more rapidly, but at the expense of code that looks more like a junior than a senior wrote it, because it is most concerned with working at all, and does not have time to worry about finer details? At the expense of not being as deeply familiar with the foibles of the API you're working on?
Honest questions as I haven't tried Copilot, and these are the thoughts that make me imagine it won't be of value. A lot of what I know I learned from doing the parts of the work that Copilot would be automating. Sure, Copilot would save me time when initially writing it. But would I then have less deep knowledge available when there's a fire in production because I never explored the fine details of my dependencies as much?
One of the major annoyances of working as a team with legacy code is when someone forgets to, or deliberately avoids, conforming their code to the style and techniques of the surrounding code. Nothing grinds my gears like working in a 500 line C++ file, where_every_function_uses_underbars, has consistent 4 space indentions, avoids exceptions, and passes by reference, but right in the middle is that functionThatMiltonWrote that uses camel-case, has 8 space indentions, throws exceptions, and passes-by-pointer.
They have a 60 day free trial. Try it out, it's one of the most interesting changes in developer tools in a while. I feel like I'm living in the future sometimes when using it.
Whoa! Really? Seems like modern software releases would be excited to achieve 98% with the incredible amount of bug fixes/patches released very quickly after the massive beta test known as release day.
Code 0.1% wrong, sends you to the Sun instead of the Moon, debits your account instead of crediting...
Presumably if AI-generated code passes every test case, but would fail on edge cases that some human programmer(s) did not anticipate in their suite of tests, the humans potentially might have made similar coding mistakes as the AI if they had had to personally write the code.
Code that is 98% correct is actually much worse than no code at all. That’s the kind of code that will introduce subtle, systemic faults, and years later result in catastrophic failure when the company realizes they’ve been calculating millions of payments without sales tax, or a clever hacker discovers they can get escalated permissions by passing a well-crafted request etc.
That wouldn't be a meme if it was something that doesn't happen
Currently, working with creatives (including programmers) is a very iterative process for non-creatives.
"I want this."
"No, I meant this."
"Can we try making that line longer?"
"Eh, I'm not feeling it. Why don't we try brighter colors?"
"Ugh. That looks obnoxious. Can we tone down the red?"
etc.
If the non-creative can use this system to reduce that iteration, it will make life easier for them (maybe not for the creative, if they are banking on the hours doing the iterations).
Well, the bottle is open, and the genie is flying around, so we’ll have to see what happens…
https://labs.openai.com/e/bzJBnS0dkgvNeZ681ZTA49nK
Sadly, "highly rendered ai vomit" is filtered.
Camera settings is just a short hand to describe the field of view and depth of focus (at the very least). If you make that implicit you'd still need to give the network the steradians, focal length, circle of confusion, etc. etc. etc. that you want your image to use.
You'd need to understand everything in Hecht's Optics to tweak all the parameters of an AI generated image.
There's nothing special about art. People will still do it as a hobby, like they do with a lot of other dead trades and professions. In many ways I think the personal aspect of the creation and sharing with people you know is the special thing, rather than the creation or performance of a commercial success (not that I've had any experience with the latter, mind you). And I'm not against the mass consumption of art, just that it needn't be produced by people if it has the same entertainment value then that's great.
Repairing expensive shoes is not an automated process. It’s more like fixing a roof leak, landscaping, or changing a flat tire. Jobs for those things still exist and aren’t going anywhere.
Someone is still going to pay for a person to perform or create art for them. Some professional driving jobs will continue to exist long after most are automated. When I said "nobody", that's what's known as hyperbole.
https://melbournechimneysweeping.com.au/ is one near me.
Not wanting to pile on the GP but their point is moot.
To be nuanced, maybe they might have said, "cobblers are less in demand now that many people have moved from owning fewer pairs of shoes they make last through repair to owning more pairs of shoes that they tend to get rid of when they are worn out due to changes in construction materials used in production," but if people have to write like that to make points, nobody will ever make a point.
But to know that these jobs still exist, and might be more common than you realize, now that’s interesting!
That wasn't the larger point. And if that was your issue, why didn't you ask about it so I could have explained it to you better?
The point was that art as a profession and a commercial venture is not special or different from any other industry with respect to automation.
Cobblers are in just as much demand in most of the world as they always have been. They only fell out of demand in car-dependent areas, which is a small minority of the world population (but a vast majority of the HN commenting population since most of the USA outside of a few cities is car dependent)
I don't know if it has anything to do with construction but doubt it. If you actually walk everywhere shoes don't last very long these days, especially shoes under $100.
There are a lot of other urban services that exist in almost the entire populated world, but that most Americans think quaint because they are not relevant to a car dependent highway world.
All that said, this really is a nitpick and the original point stands very well. Some of us just don't like it when car-dependent people forget that they are a small minority worldwide and instead treat urban walkable people as the insignificant minority! Or rather, HN being a forum for all things interesting, we find it interesting to make it a teachable moment. What could be more interesting than finding out that something that has always seemed obvious to you is actually backwards?
By construction, I mean the material and design of shoes people tend to wear. I can't say I've ever met someone who takes sneakers or running shoes to a cobbler and these shoes are more common nowadays.
These kinds of shoes are most of the market because most people don't walk much in the USA. If you walk a lot, you might still not change anything and keep buying the disposable sneakers, throwing away $160 a year.
But if you walk a lot AND are disposed to think critically about the situation, you find that if you pay a bit more for shoes you can make them last many years as long as you resole them periodically. And as I recall, a good $50 sole on a good $150 shoe costs half as much and lasts three times as long as a disposable $80 sneaker.
Not only do you save money (not really a ton) but it actually is more convenient, since even counting resolings, you get more miles between having to go repair or replace your shoe. And you don't have to wear in your leather uppers again. It is truly a luxurious feeling when you come back from the cobbler and have shoes that are worn in and fit your foot just perfectly like a glove... yet the soles are brand new and strong and comfortable and ready for another thousand miles.
Why do you think there's a stereotype of leather boots being popular in NYC? I'm sure the resoleability and longevity in the face of large amounts of daily walking have a lot to do with it.
I think that's why there's a lot of people online who think cobblers aren't a thing anymore. They're from car dependent areas. If you drive everywhere, shoe soles don't wear out much faster than shoe uppers anyway, so it doesn't make sense to care. But if you move to a walkable city, you'll suddenly find it quite economical, since the soles wear out far faster, and the cost of a sole replacement is less than a new pair of shoes, so you might replace the sole a couple times before discarding the shoe.
SD is limited to only digital art first of all, so let's not get too over-excited and hyped around this tool just yet.
You might make painting that looks identical from far away but when you come close it will be completely different.
"Just" need a decent simulation of how that works, first!
People who need illustration or graphics with no particular style can meet their needs with this tool but that is far from art. This replaces commercial illustration, not artists.
It might replace the "nice pic, do you have it at {size} so I can use it as a desktop image" comments on various social media sites.
I don't see it replacing an actual photograph as a wall hanging (or a painting) because they are subtly wrong in some ways. The reflection in the lake doesn't match the landscape... that one cloud has its lighting at slightly different angle than the rest.
Well, they are now. Give it a year and we'll see.
It's not that you can't paint that picture... but you'll never be able to capture that scene in camera. If someone was presenting that as a photograph, it would feel wrong to me because of an understanding of the meteorological criteria for the scene.
I believe that they will generate images that are impressive. I watch https://www.youtube.com/channel/UCbfYPyITQ-7l4upoX8nvctg and have been impressed with the pace of technology.
Yet, I am doubtful that I'll have a generated image at 11x17 that holds up to the same scrutiny that I apply to my own photographs.
I am absolutely certain that it will be able to generate images that are completely appropriate for images that you don't look at for more than a minute at a time or are used as complimentary material for other content.
All that said, I am not concerned that I will get more or less sales of my photographs with AI generated art competing. The people who are going to pay for a photograph are going to pay for a photograph. Those who aren't - weren't going to in the first place.
40 years ago was 1980. There's a vanishingly small number of fields that have had continuous job security from 1980-2020 (even ignoring Covid).
You might as well say that careers are over for everyone, and have been for a while.
A tiny percentage of lawyers do really well. The rest struggle under crushing debt and insane working conditions.
So you've got one example in the entire economy.
It used to be that you could find work as a political cartoonist and draw a picture of some satirizable politician each day - maybe Dr. Oz with a funny-looking vegetable platter, or Biden falling off a bike labeled "build back better," or something - and that would be a career. Now it seems you can just type "Dr. Oz holding a vegetable platter, political cartoon" into one of these AIs and the bulk of the work is automated. Sure, you could spend the rest of the day refining it, but nobody's really looking for perfection or transcendent skill in their political cartoons.
You can still find work making political videos of Dr. Oz and vegetables (e.g. https://twitter.com/JohnFetterman/status/1564432981841907713). Today's image generation AIs cannot do videos like that. Tomorrow's will, before we know it, and again nobody's looking for more than a baseline level of quality there either.
And even the local newspaper that might have been employing a political cartoonist is being swallowed by high-capital-ownership companies that can replace a lot of the writing with AIs. The expectations of quality are a bit higher - though the AIs have mostly gotten the hang of sports writing - but again people don't expect The Smalltown Gazette to have the standards of The New Yorker.
Sure, the highest-quality drawings, feature films, and longform journalism will get a lot better. And that's great! But most people don't work on such things. What are they going to do? If nothing else, how will they remain an audience with money to spend on the highest-quality works?
(I am not advocating for stopping this process, to be clear. Smashing the AIs isn't a coherent proposal, and smashing the physical machines of automation didn't actually work when the Luddites tried. I'm advocating for admitting that this process is happening and drying up employment, admitting that keeping as much of humanity as possible under a good standard of living is important for humanity, and figuring out what to do about it.)
I've been hearing this for awhile, yet ironically right now we have the biggest shortage of blue-collar workers we've had in a long time.
It seems someone indeed did this (the p&d, not the index): https://www.coingecko.com/en/coins/artonlineGood; I hope the future of software engineering AI generative tooling is also more useful for software engineers.
There are already websites which sells tailored "professional" prompts for dall-e, GPT3, etc
Creativity demonstrated by Alpha zero chess engine blows Magnus Carlsen’s mind (from his recent interview with Lex Fridman), I wonder if at some point in the future, we’ll finally throw in the towel and get out of the denial phase.
The fact is, we know no other intelligence except ours. We are the only known being that appreciate art or books, from our subjective conscience. Can AI create its own art, understood and appreciated by itself ? Has intelligence a meaning without « humans » ? It’s part of the same thing. So the AI is modeled after ours. So it’s not its own thing and cannot understand what it creates and if what is created is relevant. AI is just an amazingly powerful tool.
Ditto for me regarding being surrounded by creative people. This isn't going to lose them any business any time soon, not even the younger ones five years from now.
For threatening artists it should be much better. Maybe it's not the model, maybe it's the interface human-computer that fails, but in the end it's what we have.
Maybe someone with a lot of time in their hands can iterate enough times to come up with something nice and depicting what they wanted. Not for me.
It was an absolutely amazing accomplishment. I legitimately thought I could hold out long enough that my next car would be fully self driving. Truckers were an endangered species.
But here we are 20 years later, and we’re still almost there.
We’ve made amazing progress and I love the self driving features I do have on my car, but how many jobs have been replaced by self driving cars?
SQL was designed so that business people could describe queries in English. How did that work out?
Now, of course, we have ORM's, which will get you data, sure, but often in egregiously inefficient ways if not used correctly. If you want to to get it right, you still need to pop the hood and adjust things, and you have to know what you are doing.
I see this as similar.
Sure. Only now the gap has been shrunk by a factor of infinity, because the time to representation was infinite for someone without that experience before.
For example let's say you're a well known designer with a distinct style, and you train your own instance on your lifetime's body of work. Now you can generate whatever you want in 'your style' (just like you can now ask for a painting in Dali's style).
Now you've turned your style and design process into a factory - for every new client you can create whatever they want, in your style, with multiple examples, in a button press. Perhaps you can even sell that 'training library' to other designers?
isn't that Textual Inversion (https://textual-inversion.github.io/ ) ?
It's more or less implemented in some forks (e.g. https://github.com/lstein/stable-diffusion#personalizing-tex... or https://github.com/hlky/sd-enable-textual-inversion (discussed previously ( https://news.ycombinator.com/item?id=32643564 ) )
that's assuming the people artists sell their expertise to are able to understand the difference or what they are looking at
https://www.vice.com/en/article/bvmvqm/an-ai-generated-artwo...
I just find it a bit ironic that programmers are irate about Github Copilot using their copyrighted material to train. However, if it's an ML model training off of copyrighted artists material, clearly its a transformative work. I just find the opposing sentiments for these scenarios a bit funny.
If you use my copyrighted (public) code to train your brain, and then produce new functions that do new things, there is no problem.
Replace "brain" with "AI" and you get my position on this (and a lot of others' positions). It seems to be a lot easier to transform art than code.
That's the amazing part: the training dataset contains 5B images. Yet it distilled all of them in this mere 4Gb of data and can produce an infinite amount of content with those. It really feel like it learned how to draw, in the same sense a human learn and does not simply reproduce exact copy of what it already saw.
https://i0.wp.com/www.technollama.co.uk/wp-content/uploads/2...
Not really a copy. I'm sure you could cajole it into something very similar to the actual piece, but at that point it's more the model + the extensive prompt instead of the model itself.
For example, if you enter in something simple and high profile, what it returns looks pretty close to an existing work...
Try, for example, typing in "Banksy", "Brad Pitt" or "Starbucks".
So if you type in "Photo of a Coffee", "Impressionist painting of a fruit bowl" or "Blue canvas with one red line in the middle", how do you guarantee that the image you get back isn't actually a copy of someone's work?
Copyright exists to allow an artist time to reap the benefits of his labor. Without this time, they say, no rational person would invest his labor into making art. Dubious… but, whatever. The point is, if you remove the ‘labor’ from the mix, there’s no need for copyright.
If I can produce spectacular images with zero individual labor, there will be little reason for me to copy someone else’s work.
Can I interest you in some recently discovered DaVinci paintings? 10% off when you buy 2 or more.
That said, this is very very different from just copying at a conceptual level. This is going to end up being a an interesting legal question going forward. I'm curious to see how it turns out.
I think it’s either
A) almost like a dunning Kruger effect where people are less familiar with a domain and therefore think this has superseded it and cheer it on, whereas they’re more familiar with their own craft and can see where it falls short.
B) They see it as technology conquering a new area , but get nervous when it starts infringing on their own.
I think it’s a bit of both in all likelihood
With these artwork models, they can emulate general styles, occasionally known characters or bits of text will show up in the output, but it is (as far as I've seen) never a 1 to 1 faithful reproduction of the training material with no changes made.
This is an important difference from a legal perspective, not just a moral one. Whether or not a use of copyrighted work is transformative is a big part of whether that use is fair or not.
It actually, by default I believe, now checks for exact collisions with any existing GitHub code (the training data) and removes them. These are somewhat rare in any case, is my understanding.
Not to detract from your other points. Just a common misconception I see a lot.
Similarly, Copilot can be seen to reproduce 'snippets of code' but would probably struggle with a full application.
The images generated from these AI models are not transformative. In fact they are not even generating art, they are only generating pictures.
What is a problem, in my opinion, is the tendency of large corporations and small circles on top of these to monopolize access to these models, and if some of the functionality gets available to the public, it's going through a very paternalistic, corporate, puritan censorship pipeline.
If you train artificial intelligence on our hard-won data, the resulting artifact should be available to us. StableDiffusion executed very well here.
Problem is, there may be no stability.ai for general-purpose multimodal AIs that are coming this decade, and this technology is a dystopia fuel when it's owned by a select few.
AI should be democratized.
My brain has been trained on even more copyrighted material. Every book I read, every tv show I watch, the toys I played with as a child. It's hard to imagine that I could come up with anything that is not inspired by copyrighted work.
Sending me ads doesn't constitute a transaction, besides I usually don't choose to look at them anyway.
If a human has a right to view and learn from a copyrighted image on the internet, why shouldn't an AI?
because by buying the book you are not buying the right to distribute derivative works of that book.
I prefer straightforward movie piracy than the "hidden behind a curtain of supposedly artistic value" one.
1: you cannot produce thousands of detailed pictures in a day, this program can. The argument gets pretty clear if you transpose it to other objects, ie why is it fair to ride a bicycle in the sidewalk but not drive a car?
2: copyright laws. You may not see a picture and imitate it. How do you know this AI didn't just imitate one of the million pictures it saw? And if you distribute it and the author wants to sue, who should it sue? The author of the AI or the person who prompted it?
thought is not a crime yet.
thought police is not a thing, yet...
Car manufacturers in the future may offer to take on the liability themselves for autopilot mistakes, but that's not yet the deal offered.
At the end of the day, there's no legal magic or loopholes. Somebody is ultimately the operator of the vehicle, even if their hands aren't on the steering wheel.
I don't think we should distinguish between meat and silicon neural networks. People should be able to use whichever neural network they see fit for a particular task. If a person has a right to observe and incorporate an image, so should a ML model operated by that person.
Also, you may see a picture and imitate it. You may not copy or redistribute it directly.
> How do you know this AI didn't just imitate one of the million pictures it saw
I think the onus is on the copyright holder to prove a violation to a specific work.
> who should it sue?
I think the operator of the network is liable. If you distribute an image that is found to be in violation of a copyright, then you are liable no matter how you came about it.
That is legitimately a problem, people get sued for it all the time.
That's in fact a problem: plagiarism is considered cheating/intellectual dishonesty and can ruin a career, copyright infringement is punished by civil law (and fined), counterfeit is a crime, etc. etc.
You must retain yourself from copying too much, but only the original author (and eventually a judge) can decide if that happened.
If stable diffusion is used for counterfeiting or the result infringe the copyright, do you think the fact that the model was trained on unlicensed copyrighted material is irrelevant?
Nitpicking, but only unless it's not counterfeit material.
Claiming painting is from "famous painter X" while it's instead a copy it's an infringement.
How long before "famous painter X's new painting found in the attic of old lady"?
anyway, copying money is a crime, almost anywhere in the World, the simple act of making a copy is enough.
Money usually contain artworks
I would also not generate pictures of child pornography
There's a reason why SD apply censorship filters to generated images
Generating a digital image of money isn't the same thing as counterfeiting currency.
> I would also not generate pictures of child pornography
Legality aside, why not? Who is harmed?
> There's a reason why SD apply censorship filters to generated images
I'm not sure there is. I don't think any one group of people is uniquely equipped to limit what images another group of people can generate with ML.
of course there is, you're free as long as you respect the rules.
I'd expect that level of legislative response, and also would bet on lawsuits over any authorized data in their training corpus.
What your brain has learned cannot be transferred with an USB stick in seconds
Not even your offspring will receive any of it
If I want to learn everything you know, I have to learn what you learned , assuming I will be able to
Kinda of a big difference, don't you think?
These kinds of comments are embarrassingly low effort, just because we threw rocks at each other when we all were chimps, doesn't mean that guns haven't been a game changer and made homicide easier even for people that would never be able to hit anybody by throwing rocks.
Your comment implies that DALL-E 2 is morally okay, because they don't distribute the model ("copy everything it knows to a USB stick") but only sell access to the algorithm to generate images, while Stable Diffusions open source model is a problem because it can be copied.
Most people would take the exact opposite stance I guess.
but I am replying to
"My brain has been trained on even more copyrighted material. Every book I read, every tv show I watch, the toys I played with as a child. It's hard to imagine that I could come up with anything that is not inspired by copyrighted work"
Difference being your brain has not been trained by someone (for profit), you have trained it using YEARS OF YOUR LIFE TO ACQUIRE KNOWLEDGE AND EXPERIENCE
which is morally acceptable (does not imply that the use you do of it is legally acceptable), given that you paid a very high price, sacrificing your own time for the objective.
And that your knowledge is only yours, you can't transfer it to anyone, it doesn't even show up in your DNA.
> Your comment implies that DALL-E 2 is morally okay, because they don't distribute the model
Implication doesn't mean what you think it means.
My comment doesn't imply anything of the sort, you are
> Most people would take the exact opposite stance I guess.
The problem here isn't that the model was trained on copyrighted works, the problem is copyright itself and a cultural focus on collecting rent. We probably need UBI or equivalent to deal with the outcomes here in a healthy way.
Am I missing something here?
It's also quite obvious just by looking at the generated images that it's clearly transformative. The images generated are unique and you can't trace the original copyrighted image from what's generated.
You really don't need a judge to see that Fair Use covers Stable Diffusion.
Federal lawsuits are not cheap and the default is that you pay your own costs, you have to win the argument that you should get court costs & attorney's fees.
Or for written works, start with a sentance from a copyrighted work, or part of licensed code. Will it start reproducing that work word for word (like code pilot can do with the GPL license)? Getting these to generate copies of GPL'd, company owned, or other code with restrictions can lead to complex issues for the person/company using that code. Or likewise if a story contains significant elements of copyrighted works; worse if the works have trademarked elements.
Explain how you are not a copyright violating machine.
That might not always be true. I've gotten some results back that had the Getty watermark on them and others with the artist's signature. Unless the AI is adding that to images that never had one before (which might be a trademark issue), then you might be able to determine the provenance of the image components.
Is there something equivalent to the yellow dots printers add to their output that would survive the AI transformation?
If you take a step back, you can see that there are different ways to frame what is happening. One frame is: “Defendant built an algorithm that memorized features of Plaintiff’s IP. Defendant’s algorithm recombines parts of those features in order to produce works in the same domain that compete with Plaintiff’s work, all without Plaintiff’s consent.”
Bear in mind that copyright holders are among the most litigious out there. If generative art becomes as big a deal as some people expect, they will have every incentive to use their huge litigation budgets to claim a piece of the action.
Because the law develops very slowly, the legal process has not yet had the occasion to really evaluate what transformative use means in this novel context. I’m personally interested in seeing where things go, but it’s going to be a while before we know where the law is headed.
The fun part is, this is how human artists learn too.
Will the same image be legal if a human made it but not if it was created by Stable Diffusion? How will someone even know, short of a legal discovery process?
The model is open source and free for people to make their own modifications. I don't see how any watermark can survive in those conditions.
All this is super, super interesting to me. I don't think anybody really knows how it will play out, there's a lot of good arguments floating around.
I don't think computer-generated works will be easily distinguishable from human unless they're desired to be or shipped with metadata. It's already hard enough to distinguish human artists from other human artists without having names attached up front.
Yes, and tracing counts as art fraud.
Collage is a bit different, because you are mixing many clones of many other objects such that you create a new object; additionally the way you assemble the clones may transform them (a photo of the mona lisa has different surface texture than a painted version, even more different if it is clipped from newsprint), but while the borders of this are not clear, it is clear when people are far enough over the border. Think of hip-hop, sampling, and remixing music, and some of the legal battles which have come out of that.
I just spit out my coffee, turns out I'm not a bad artist, I'm actually a fraudster!
Yeah I am sure a lot of lawyers are going to have a lot of fun arguing every way imaginable.
Deep learning is very powerful and impressive in its applications to date. However, it’s so saturated with hype (and humans are so prone to anthropomorphizing things) that it’s often viewed as something much more profound than it actually is. Neural networks, despite their name, don’t model the brain. And they lack a whole array of “intelligence” features that humans possess and use constantly.
All of this is to say that there are very significant differences between computer algorithms and human cognition, and I tend to think the legal system will be unpersuaded by arguments that ignore those differences.
Also, this is to say nothing of the public policy interests that shape the law. Regardless of what’s “under the hood,” the law can simply treat human and machine output differently. I’m not a copyright lawyer, of course, so I can’t speak to the norms or technicalities of copyright law itself.
There are a lot of things we don’t know, but it is not magic. There is no discussion that artificial neurones don’t have much in common with the real ones, which are very non-linear and much more connected. But in the end it’s all electrochemistry.
> Neural networks, despite their name, don’t model the brain.
But that’s not directly related to my point. My point that even in the case of a ML model, you cannot get an exact reproduction any more than you can get from a human’s memory. In one case it’s scrambled somewhere in someone’s brain, in the other on a hard drive but the difference is not really relevant. Subjecting an AI’s production to the copyrights of all the things it’s been exposed to is very similar to subjecting a painter’s production to the copyrights of all the painting they have seen.
That’s an element of a belief system you may choose to subscribe to, not a fact. It can’t be a fact because there’s no way to prove or disprove it.
> “My point that even in the case of a ML model, you cannot get an exact reproduction any more than you can get from a human’s memory.”
Not quite. ML models can and do memorize and regurgitate near-exact features of the inputs. It’s not the goal, but it happens.
In the (actual) neurons, is there a representation of real numbers? Where are the numbers in the brain stored?
I feel like people who assert this neither understand what neural networks are and how brains work.
None of these details are relevant to the bigger picture similarities of non-hardcoded learning from training data. None of these details change the ethics of what's being discussed here.
> I feel like people who assert this neither understand what neural networks are and how brains work.
That's just your bias showing.
Don’t they, though? What else do you think is behind the curtain if not mathematics?
We don't actually know exactly how human artists learn, and human artists are capable of innovation, nobody knew pointillism or Bauhaus before they were invented.
A little know fact is that for humans it takes a long long time to learn, while they learn, they develop a style, if they don't they are not "real" artists, but merely executors, artists evolve, sometimes dramatically, in unexpected ways [1] [2].
So for us humans learning is an experience, not just recombining parts of features of other things.
We are also highly influenced by feelings, unfortunately, so sometimes we do things a certain way because we felt that way, not because we wanted to paint that thing that way, or because we are not good enough to do exactly what we wanted to do.
Is Mona Lisa happy? Who can tell?
Was Leonardo happy when he painted it?
What was Leonardo thinking when he painted it?
What was happening in his life?
Is that the best smile Leonardo could paint or it's an enigma he put there for future generations?
These questions are more important for an artist than the mere features of the painting.
The philosophical question is: is art discovered or invented?
If it's discovered, then SD can generate art, if it's invented, than SD it's not even generative work, because to invent something from something else, you need inventiveness.
[1] Picasso 1896 https://mymodernmet.com/wp/wp-content/uploads/2018/01/pablo-...
[2] Picasso 1946 https://www.photo.rmn.fr/CorexDoc/RMN/Media/TR1/MS4GY/16-515...
- It appears people can train AIs from scratch or at least fine-tune them at home.
- Even if your art isn’t in “the training set”, that does not prevent the AI from learning its style. (Someone can decode it to CLIP embeddings. It could have a really good text model trained on vivid art museum descriptions of your art.)
- The ability of an image model to generate your art means it could also be trained in reverse to recognize it, producing a caption model, which would give vision to the blind. And surely you’d feel bad about that.
You can also require cloud providers to enforce a ban on training (and deploying) such models, it's doable. Good luck training it in your basement, it will probably take you a decade.
If this is banned, it will become a lot like piracy - yes, it's available, no, most people (at least in the West) don't do it, practically no businesses do it.
Either use a CC0 set like Wikimedia/Flickr and throw in some dead artists like Brueghel, or train on data from a country we don’t respect the IP of. Lots of Taobao product photos out there. It’s enough.
As of about a week ago this tech runs on consumer GPUs. The weights have been downloaded 100s of thousands of times, and fine-tuning / modifying is possible.
Training from scratch is about $500k still, but it will only get cheaper and easier.
But I would not be surprised if this was trainable on a commercial GPU at home within that time. But I think another important trend that we are seeing is that you don't need to train these models from scratch.
Open-source "foundation models" means that you can usually get away with the much easier task of fine-tuning, as to not throw away / re-learn everything that these large models have already fit.
Edit: I initially said 2-5 years, but on more reflection this does seem optimistic (for training from scratch).
I don't know enough about diffusion models but if LLMs (of current size) have to use only public domain, they will be undertrained and we will see significant degradation in performance. Not to mention that Codex will be effectively dead.
Why would you, though? If art is for the sake of art, then all art is valuable regardless of origin. If art is for the sake of providing human employment, AI being better in no way stops performative make-work from existing. If art is for the sake of copyright trolls to troll harder, then fuck art, feed it to the AI!
Your "logic" for making AI art illegal is basically "don't like it". Your personal and subjective opinion is that it's not art by definition.. This is like refusing to eat artificially grown meat because you have some strange idea about what food "should" be. Even if the meat was made MORE delicious you would still claim it wasn't food and turn it away. There's no logical consistency to your position, it's purely reactionary.
What is not logically consistent is to claim that a black box utilizing statistical relationships between pixels in a giant dataset is an "artist" and that its products create "value".
The compiler is not a programmer, AI can never be an artist.
Step 1: Tweak settings and type text
Step 2: Look at the result
Step 3: If you like the result, go to step 4, otherwise go back to step 1.
Step 4: Save and share the result
Feels like art to me. Ultimately it's still a human using a tool to create art. For me AI art is just the name of the art style.
All this boils down to a simple fact that egos like to think of themselves (and of artistic interaction) much more than there actually is.
I remember a story when a literature teacher insisted on a definite symbolism of some minor detail in a novel. People contacted the author about it and he said no, there is nothing behind it. It was just a filler without any second thought. Makes you think how much symbolism is far-fetched in classics, where you cannot simply email an author.
If you tried learning, let's say, the chiaroscuro technique from Caravaggio you'd be analyzing the way the painter simulated volumetric space by using white and dark tones in place of natural lighting and shadows. You wouldn't even think of splitting the whole painting into puzzle size pieces while checking how many how those look similar when put close one another.
Given somewhat decent painting skills, you'd be able to steadily apply this technique for the rest of your life just by looking at a very small sample of Caravaggio's corpus.
On the other hand if you tried removing even just a single work from the original Stable Diffusion data set you used to generate your painting, it would be absolutely impossible to recreate a similar enough picture even by starting from the same prompt and seed values.
Given how smart some of the people working on this are, I'm starting to believe they're intentionally playing dumb to make sure nobody is going ask them to prove this during a copyright infringement case.
Both my parents (though retired now) were commercial artists. I was trying to be an artist at one point in my life before moving in Engineering and Science. All my parents friends are artists so I grew up around artists.
Ask any artist here who is using illustrator, Photoshop, Krita etc. How often do they google image search for textures, or reference images that gets incorporated into their artwork? The final artwork is their own but it may incorporate many elements from others artwork.
>If you tried learning, let's say, the chiaroscuro technique from Caravaggio.. You wouldn't even think of splitting the whole painting into puzzle size pieces while checking how many how those look similar when put close one another.
Ever seen hyperealistic pointillism?
Who are you to be the arbitrator of how an artists creates their work? Have you ever gone to a modern art gallery and seen all the different methods people use to create artwork?
Art is boundless and unique to each who creates it.
If a Artist uses a tool to create art, everyone agrees that is art. It could be a paint brush, clay, software on a computer etc etc. If an artist uses AI as a tool to create art then suddenly it's not art.
What potentially human-creatable images have I just taken ownership of?
Let's extend: what if I claim IP ownership of every image which StableDiffusion could produce?
For example, it has previously been litigated that nobody has IP ownership of an image taken with a camera by an animal.
Note: I am absolutely not a legal expert.
I don't think SD will replace most top artists for now. It's hard for me to believe SD is going to come up with images like those from top concept artists. But I can imagine SD replacing lots of situations, like maybe stock photography, when you can just ask the AI to draw "people in front of whiteboard discussing sales chart"
What's a similar thing that has come before this? I can't think of any, this is very novel. You'd want to wait for some rulings before you jump to conclusions.
As far as I understand it, this is still considered copyright infringement in most IP law systems. (If the samples aren't cleared)
So when your image generator keeps spitting out Gettyimages watermarks, while you are building a service that is in direct competition with Gettyimages for stock images, there is an argument to be made that Fair Use really doesn't apply here. As what you are doing is essentially stealing Gettyimages' work, AI laundering it and selling it back to their previous customers.
With StableDiffusion a Fair Use defense might have an easier time, as the results are released to the public. But it's still not exactly clear cut. If you type in "Mona Lisa", you'll still get something that looks like a copy of the Mona Lisa, not like an original work.
Everyone should be able to freely use any public works, and governments should ensure that people whose art is being used to train these models are able to continue to do so.
I first read this as a prediction that AIs will be employed to generate future IP legislation, and now I don't know if I'll be able to go back to sleep.
1. Artists that use SD etc., will save a lot of time and have a big edge. Example: Generating references instead of searching for them, Generate random inspiration(s) for abstract ideas, Variations of a particular idea to compare and contrast.
2. Non artists who have powerful imagination but lack the artistic skill will stand to benefit as well. For example, I look forward to generating visual companions for my poetry, essays or fiction. I have the choice of commissioning an artist if my piece is successful.
3. Artists that are purists and masters of their craft with a very unique style will continue to thrive in their niche. These people have to lobby and come up with a license that explicitly prohibits feeding their images for learning. But a skilled artist can replicate the style and feed it. There is an ethical line here that may require creative tweaks to existing laws to protect people that fall in this category.
Super exciting tech and I can't wait for this to move over to 3D models. I know 2D image to 3D model is already pretty close to being real(nvidia ominiverse, nerf etc) and SD can be the starting point of that pipeline. Prompt -> 3D world that can be decomposed to meshes, textures and USD will change the game development landscape quite a bit.
"An elf fighting an orc in a forest"
Half the images had christmas elves so it's obviously bad at context, half looked like miniatures (models like games workshop), so it's obviously got a load of its data from pictures of miniature but doesn't realize and thinks that's what it should draw, and the "fighting" always looked pathetic.
"A human berserker with an axe fighting 3 orcs in the desert"
That had problems with actually putting a human in. Also no axe. Orcs designs were obviously stolen from LotR films orc design, the actual breath of orc designs in fantasy didn't seem to be represented at all. Often the weapons are for some inexplicable reason only part drawn.
"A space elevator leading up from new york"
Was a little more promising, but looks pretty amateurish, just a picture of a space elevator transposed over a picture of new york with terrible lighting, though it was choosing half decent perspectives.
"A 12th century castle on a hill in a hellscape surrounded by demon hordes"
Pretty much ignored everything after "castle", just a bunch of 12th century looking castles.
"A castle on a hill in a hellscape surrounded by demon hordes"
None of the 9 images generated has any sort of demons, it just coloured some castle a bit red.
I think artists are safe for another few years.
https://news.ycombinator.com/item?id=32634074#32661992
This isn't supposed to be used the same as Dall-E mini. You're supposed to be very detailed in your prompts.
Also the online demos tend to not let you tweak parameters. You get more options if you run the model yourself
You seem to have missed the main point of the article this thread is based-on though. With img2img, you can input a kid-level MS paint input and control composition in detail. With inpainting you fix individual parts, with outpainting the surroundings/background. With text inversion, you can reuse successful parts in any context. People are even making the first videos and animations.
All of this is available today, not years from now. The workflows are still clumsy, but it's been a week. Many people are working on the UI to streamline all of this. You're severely underestimating progress.
I think as technologists we want to think that code can "solve" some of the problems in the art world... but I think we still have a really, really long way to go. I tried to get style transfer adopted at work (worked at a creative technology firm in NY) but frankly I think deep learning methods for art generation tend to be really unpredictable, which make them pretty hard to use for professional applications. Imagine deploying production code that only worked 85% of the time... would be a nightmare. I felt, and feel similarly about deep learning approaches to art. They're just so finnicky and unpredictable, for example, add a single extra pixel to that example in this article and the output would look completely different.
Either way, cynicism aside, stable diffusion is awesome :).
Don't think the metaphor works. Code that only works 85% of the time is obviously broken but Art is subjective so an 85% solution to a creative problem could be more than enough for most consumers.
I can find a good prompt within 30 minutes to 1 hour.
My GPU can generate 100 images in 5 minutes.
Out of those 100 images, 10 is very close to what I exactly meant at professional concept artist level.
So, in this case Stable Diffusion only working 10% of the time is fine.
Future is already here, I’m already incorporating stable diffusion generated images to my professional work.
Are you using 512x512 images or larger ones?
Best workflow is to keep images close to 512x512, record the seed and then upscale.
If you are not using this library already, give it a shot.
Also, I'm using Nvidia Studio drivers though I'm not sure if that would make a difference.
Maybe that's because we never really thought about it. In hindsight, it's only logical. For an artistic rendering, correctness doesn't matter much nor does understanding the model. For flying a plane or driving a car or just transforming code from one language to another, it very much does.
(The “cow tools” Far Side comic is a dumb example.)
Since these AIs can’t count, that extra weird thing is probably going to be people with too many fingers. I actually have a personal collection of weird Midjourney images because if you ask for wide aspect ratios it starts generating multiples of the same thing and fusing them together…
Wouldn't this create more jobs for artists? It's just a different tool
Getting higher level languages and various libraries/frameworks didn't make coders obsolete, because there was still tons of work to do at a "higher level" that required algorithmic/engineering-style thinking. Yes, previously difficult tasks got a lot easier, but that just meant that even more ambitious projects could be tackled, and that was useful.
In contrast, it may be the case with art that the things a relative newbie can spit out with the help of image generation models are "good enough" for most use cases. Yes, a real artist would be able to squeeze out even more, but many companies may be willing to go with the option that's enormously cheaper and still seems to produce okay results.
What we're going to find out is whether novice users of image generator AIs frequently get stuck with inadequate results. IMHO the answer is likely to be "no", especially after we've iterated on the tools for a few more years.
It's not clear that you'll see something similar here, that easier-to-produce high complexity art will greatly increase demand for said art.
It's like some of the stuff that GPT-3 generates, it sounds extrememly plausible and realistic but then you read it properly and realise that parts just don't make any sense.
However I don't think these things are show-stoppers, I think we'll see a lot of artists getting good at providing the "feedstock" and guidance to these systems to deal with those quirks and generating some really interesing work.
Symbolic AI dominated the field for 50 years: thanks partially to an accident of history (the highly influential work of Minsky, Newell, etc) and partially due to lack of the data needed to try anything else.
Now we’ve seen essentially 50 years of research on data-driven methods compressed into 5. In retrospect it makes sense that applications to which symbolic reasoning is especially ill-suited (including artistic rendering) would be the lowest-hanging fruit.
Francis Bacon has a quote that is apt about how he paints the things that cannot be described with words.
So Donald Buck who is a duck is too derivative, but Ronald Cluck who is a bear might be fine.
Depending on how close it is it may not even be legal. To give a music analogy: at least in my country, if you take a melody, change all the notes' durations, and transpose it you'd still infringe
I can write a book about a boy wizard's adventures wizard school and that's legal, but if I call them Harry Potter it isn't.
I can create Harry Potter fanart and distribute it online pretty freely - but slap it on a mug and sell it, and that's illegal.
I can record an audio description of a painting that's as detailed as I like and it's legal to distribute - but take a photograph of the same painting and it's a derivative work, no matter how artistic my choice of camera settings.
We don't really have any prior examples that are precisely like these huge ML models trained on copyrighted data - and depending on which imprecise analogy you choose, you can come to a different conclusion.
If your fan art is infringing, it was infringing whether or not it was on a mug or on dropbox.
Photograph one is not true. For commentary it can be by audio, or printed on a mug or whatever, commentary is transformative. Go out and take a photo of the world outside, if you live in a city you've captured thousands of copyrighted materials in your image. Maybe it captures someones painting, maybe it doesn't. Whether its fair or not is if it's transformative, the format doesn't matter.
I'm not seeing an ethical difference from any of this. Or did I miss the point?
Perhaps I should have said Darth Vader, then - the point is you can copyright a character independent of the copyright on a book's text, and the trademark on the series name, and the fact that broad concepts like "black-clad masked evil overlord" are uncopyrightable. And that copyright can persist even if you transform a book character into an engraved coffee mug.
> I'm not seeing an ethical difference from any of this. Or did I miss the point?
The difference is:
If a person says "Stable diffusion is to its copyrighted training data as an audio description is to a painting" or "Stable diffusion is to its copyrighted training data as the concept of boy wizards is to harry potter" they would probably say it's ethically fine.
If a person says "Stable diffusion is to its copyrighted training data as a photograph of a painting is to the painting" or "Stable diffusion is to its copyrighted training data as video lecture is to a single image in its slides" they might well say it's not ethical.
So since automated driving will result in financial suffering to those working in trucking and Uber and whatnot, automating driving is not ethical?
Think of this instead: a doctor collects a big volume of symptoms and analyses and creates a statistical way to cure people more easily. They publish many examples of their work without licensing anyone (legally and morally) to use it freely. Now some algorithm collects their data and many others data and transforms it into a better method. A doctor suffers from going out of business. Is that ethical? On one hand, the algorithm invented something new and easier to access. On the other, it basically stole parts of their and similar researches on a previously unthinkable scale. We humans copy ideas all the time and this is somewhat normal, but this enormous at-scale capability was never a thing.
Personally I don’t care for optometrists, uber drivers or designers. Nature will find a way. But when we talk about fundamental social contracts like property or accumulated knowledge protection, I think it is unethical to break them, regardless of technicalities. If it’s such a great advancement benefiting everyone, why can’t AI creators just ask permission for 2.3B of datapoints they used?
I’m not optimistic that their impact will be positive though.
BUT it will likely be disrupting the upper middle class more than the lower middle classes, as it's mostly disrupting to creative work. Upper middle class people have a louder voice.
With the risk of looking silly, I declare "meh" once more, just as I "meh"-eh when GPT-3 came out.
The so-called "AI" is not fundamentally different from the AI of the 80s, it's just that now we have much better hardware. The main problem of the past AI winters still remains - the existing algorithms focus on statistical methods, which can be rather inexact. Imagine a nuclear plant controlled by an AI, or an airplane flown by a neural network. They completely lack reasoning capability and therefore you can't trust they will be able to adapt to unpredictable situations.
Either way, I think the current "AI" is fundamentally limited by the statistical approach to problem solving. Without any reasoning capabilities, no amount of incremental improvements will change the fact that neural networks are simply making guesses based on existing data sets. Nothing magical or revolutionary, it's the same thing we've known for many decades.
Not lifetime. Right now. Just one week of StableDiffusion in the wild has seen the development of many clients, GUIs, optimizations, plugins for commercial software, editing workflows, the list is endless.
The speed of it all is dazzling. It will not take long before this is assembled into a smooth experience that runs everywhere, with obvious end state being your phone.
And that's still just artistic image generation. The open sourcing of it means people can make any vertical, precisely optimized for a particular (commercial) domain.
Almost all shortcomings seem to be crushed in no time. Bad at drawing eyes? Here's a new encoder that fixes it.
The creator of StableDiffusion indicated to soon tackle music generation and if I remember correctly, even poetry. And there's 3D and video.
The ultimate end state: if you can imagine it, it can generate it. Not only are we much closer to that point than society realizes, we're also moving at an exponential speed towards it.
https://www.paepper.com/blog/posts/how-and-why-stable-diffus...
This tech is a huge deal. Huge. Ubiquitous unique, beautiful art. Entire movies made by machines. Too many to even watch them all. Actor models that aren't real people owned by the thousands by people and corporation s.
But maybe the feedback will make marginal defects/artefacts worse. (Which could be manually corrected before re-entering this loop).
Possibly but not necessarily. The human curation prior to posting as well as additional textual context associated with the image could be valuable training signal, even if there is some feedback.
It's a fantastically great tool, and a very exciting space, but reducing the function of creatives to people who draw pretty pictures is staggeringly ignorant.
A better comparison is like saying AI text generation bots will replace authors, or that AI drug testing will replace drug development - which clearly isn't the case. The core issue is that people with no idea what the creative fields offer are throwing their hat in with what they think is going to happen. People with experience are saying the opposite because they know better.
At best, this will remove those websites where you can pay $5 for "a designer" to make you "an image" - is this a loss though? Such things have never been a threat to creative fields.
With imagine generation you just look at output. There's no difference between seeing the output of a human-created image or an AI-created image, people can't tell.
The way I see it, if you'd consider the art world a pyramid, the bottom is about to fall out. A lot of commercial artwork serves no deeper meaning but pretty decoration.
The emphasis will move to ideas instead of just execution. Artists will soon find out about the avalanche of people that have creative ideas yet can't draw or paint. They'll be unlocked.
If people were saying "oh hey this is going to give the lazystock on iStockPhoto a run for their money", I wouldn't debate that point, it's true. However that's not the industry, and it's certainly not where the money is - neither in # of customers nor total spend. Those people who you might think are customers simply put: aren't, they get by with images stolen from google images and bundled clipart, or frankly: nothing at all.
Now this isn't to say that txt2img isn't useful or exciting. I can say that it is the largest and most significant expansion of creative tech since the advent of DTP. This will absolutely accelerate and open the door to not just higher standards, but new ways of rapidly ideating concepts. I've already seen fantastic examples of txt-to-image-to-mesh-to-live animation. All automated through AI.
This is also why I speak against the other kinds of naysayers: the ones that think this tech is unimportant. These types are being incredibly short sighted and acting like we're looking at this tech's endpoint, rather than its infancy.
tl,dr: No creatives are not being put out of the job. Yes this tech is incredibly important.
Surely you have a point that this does not cover the entire world of art, but I think it would be helpful if you constructively explain which parts are less or not affected, instead of calling people ignorant.
A better approach for people is to ask questions, rather than trying to write controversial falsehoods.
Right now there is at least one high ranking submissions on HN where a creative details how this won’t end their career, but you don’t need to read it - social media is filled with creatives literally rejoicing - no one is sweating this.
So to that: I say that ignorance to this is definitely a choice.
It would be great if this would turn into a community-driven chat-based text2image offering - which is certainly challenging as the required GPU-powered server instances aren't exactly cheap. Maybe this could grow into an open network where people provide the GPU power of their personal PCs? Or we find a way to cover hosting expenses through a credit system or some kind of sponsoring?
Feel welcome to join the community [1]. You can easily test drive the current Bot implementation on this server.
[0] https://github.com/manuelkiessling/stable-diffusion-discord-... [1] https://discord.gg/nsfeutx35z
% find . -type f
./migrations/Version20220723092633.php
./composer.lock
./.env.local
./bin/connect-to-db.sh
./bin/_init.sh
./bin/deploy.sh
./bin/visualize.sh
./bin/console
./config/services.yaml
./config/bundles.php
./config/packages/messenger.yaml
./config/packages/framework.yaml
./config/packages/doctrine.yaml
./config/packages/translation.yaml
./config/packages/cache.yaml
./config/packages/routing.yaml
./config/packages/doctrine_migrations.yaml
./config/preload.php
./config/routes.yaml
./config/routes/framework.yaml
./public/index.php
./symfony.lock
./.env
./composer.json
./.env.local.dist
./src/SymfonyMessageHandler/VisualizeSymfonyMessageHandler.php
./src/Command/BotRegister.php
./src/Command/BotRun.php
./src/SymfonyMessage/VisualizeSymfonyMessage.php
./src/Kernel.php ./migrations/Version20220723092633.php
./migrations/.gitignore
./composer.lock
./.env.local
./bin/connect-to-db.sh
./bin/_init.sh
./bin/deploy.sh
./bin/visualize.sh
./bin/console
./config/services.yaml
./config/bundles.php
./config/packages/messenger.yaml
./config/packages/framework.yaml
./config/packages/doctrine.yaml
./config/packages/translation.yaml
./config/packages/cache.yaml
./config/packages/routing.yaml
./config/packages/doctrine_migrations.yaml
./config/preload.php
./config/routes.yaml
./config/routes/framework.yaml
./public/index.php
./symfony.lock
./.gitignore
./.env
./translations/.gitignore
./composer.json
./.env.local.dist
./src/Repository/.gitignore
./src/Entity/.gitignore
./src/SymfonyMessageHandler/VisualizeSymfonyMessageHandler.php
./src/Controller/.gitignore
./src/Command/BotRegister.php
./src/Command/BotRun.php
./src/SymfonyMessage/VisualizeSymfonyMessage.php
./src/Kernel.php
That's a total of 36 files in 7 folders.So why do you come here claiming "like a thousand files spread over a hundred folders"? Have you considered the guidelines [0] of this community?
[0] Be kind. Don't be snarky. Have curious conversation; don't cross-examine. Please don't fulminate. Please don't sneer, including at the rest of the community. https://news.ycombinator.com/newsguidelines.html
(Also, very happy for somebody to steal this idea!)
It is for authors that got their works stolen for "the greater good" without even being notified.
A friend of mine found his works in the Stable Diffusion dataset, the work was not meant for public use, he's never been notified and, most of all, he would have never agreed if they cared to ask.
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
That said, it did get a little better at drawing like your friend from that. But so does every other person who looks at his art. The model however might be better at it than most people. It's a dilemma but it's hard to say any individual image is "stolen"
We kinda already went through this once when GitHub scanned all public codebases to build the model for Copilot. When they did that, Copilot got the ability to program a little more like me (sorry everyone). But when I publicly put my code out there, I was also giving people the opportunity to learn from my code and reproduce similar things to it as well.
I spoke with the author, the author disagrees with you.
The reason he consider that stealing is because they took that image from the online shop where the disclaimer specifically states "2022 © Magnetic Press LLC by Arenathemes. All Rights Reserved" without even trying to contact the copyright holder.
https://store.magnetic-press.com/products/viewpoint-by-lrnz
> But when I publicly put my code out there, I was also giving people the opportunity to learn from my code and reproduce similar things to it as well.
Not really, IMO
I can enter a bookshop and look around, open a book, read some pages and maybe learn something.
It doesn't give me the right to republish a similar book putting together what I saw, unless the license specifically grants me that right or I obtained a license that does.
Code Pilot is a commercial product, free licenses don't automatically grant the right to commercial use of that code.
> took that image from the online shop
They didn't take the image. They looked at it and learned a few datapoints from it. It's not like they compressed the image and added it into the model, able to be reversed by anyone who downloads it. In fact it's likely there's other completely distinct images that could have contributed the exact same tweaks to the model, like a hash.
>I can enter a bookshop and look around, open a book, read some pages and maybe learn something.
>It doesn't give me the right to republish a similar book putting together what I saw, unless the license specifically grants me that right or I obtained a license that does.
No, but the analogy would be if you learned something from the blurb you read - like a word, or a factoid, or an expression - and it influenced some writing you did later. You didn't steal from that book's author.
All rights reserved means "all rights"
Nobody authorized them to do what they did.
> No, but the analogy would be if you learned something from the blurb you read - like a word, or a factoid, or an expression - and it influenced some writing you did later. You didn't steal from that book's author
that's what clean room implementation is for
https://en.wikipedia.org/wiki/Clean_room_design
Wine contributors were not allowed to see Windows leaked source code to avoid copyright infringement claims.
both WINE and ReactOS have refused to use the leaks; ReactOS doesn't even allow people who have worked legitimately at MS in the past to be developers, simply because even the smell of contamination would expose the projects to enormous legal risks
So, to be fair, Stable Diffusion authors should take inspiration from all the source material, replicate it the best way they can using their own abilities and then train the model on what they produced.
They didn't do it for two reasons:
- it would have required centuries
- the result would have been much less compelling
stable diffusion is so interesting due to the fact that the source material was of high quality.
So, we can conclude that it's the merit of the authors of the source material, not simply the model itself.
Shouldn't they be rewarded or at least consulted before using their works?
What if some of them is training their own model on works they have the right to use and now Stable Diffusion took all their hard work away?
If stable diffusion was trained on my drawings, it would produce a steaming pile of shit.
And nobody would be talking about it right now...
“Copyright” and fair use are human defined not absolutes, so the real debate is whether AI models should be allowed to learn from copyright works like a human does. This use case is also close to people who use clips or samples snippets of works in other bits of art which they then profit off of. It is way more efficient at this process than previous methods (having an artist study and practice a style then reproduce it), so that is where the debate is.
There are a gazillion artist who take clear inspiration from other artists, and there is no copywrite violation in doing so. The AI viewing his work isn't any different.
Actually he does.
> There are a gazillion artist who take clear inspiration from other artists
Inspiration is not the same thing Stable Diffusion does.
We can't even define inspiration in a proper manner, but for sure we can say that if someone wants to draw comics in the way Tezuka made them, they have to study, exercise, rinse and repeat for at least a few years.
No human can scrape billions of images and take inspiration from all of them, not even in 10 life times.
Also, no human will make something similar to something else seen for the first time in 10 seconds.
Not even if it's "La Linea" from Cavandoli.
Unless your friend thinks the AI is copying snippets of his work into generated images, I'm not sure where they are getting these ideas from.
My friend is just upset.
Legally he has every right in the World, he's the author for Christ Sake!
Will he try anything?
Of course not.
Are you okay with this?
Well, then you should reconsider your values.
> People do "like copies" of works all the time,
If those copies are authorized, I don't see the problem.
Try to recreate a Star Wars image and sell it on the Internet.
See what happens.
> Unless your friend thinks the AI is copying snippets of his work into generated images
Would you bet your life on the fact that it doesn't?
You know why we Europeans came up with the GDPR?
Yes, exactly, because processing of the data, automatic or not, must be authorized by the holder of the rights on that data.
We are not talking about artistic expression here, I'm not sure where you are getting this idea from, this is not inspiration or art or human expression, this is simply data processing leading to algorithmic replica of other people's works.
Without people's work, no model could replicate it.
Impressive, still completely dependent on some kind of source material, that they scraped, they haven't produced it by themselves.
There will be a day when people like you will realize that appropriation is not right.
Books are different, most of the works models are trained on are in the public domain, Shakespeare is public Domain, nobody will ever contest that, but if you think a living author has no right to have a say before someone process their work, you're the one with crazy ideas.
Yes, because that's not how it works.
If your friend is still upset, maybe they should consider the artists they "stole" their learning material from.
I think you don't know how laws and Author rights work
The simple fact his work became part of something else he did not authorize is the problem here.
And yes, it could spit out something that is very close to the original, so close that fair use could not stand.
Fair use is not a right!
> If your friend is still upset, maybe they should consider the artists they "stole" their learning material from.
He does, don't imply differently, ad hominem are a stupid argument for very stupid people.
That's why he spent 30 years of his life learning and in the end he became good enough to meet the artists he "stole" from to thank them of what they did.
You seem to lack the ability to understand the difference between being a good person and being a senseless automata...
That's the one argument, but it's not the most important important.
The most compelling issue here is that SD used copyrighted data scraped from the Internet without even informing the authors, who were not unknown to SD authors, because they tagged them in the model.
If your friend published images on the internet, he clearly intended them to be viewed by (meat) neural networks. Why are silicon ones operated by humans any different?
Edit: To elaborate, if (10 years ago) someone saw your friends works and started producing derivatives and publishing them to DeviantArt, would your friend have any good reason to be upset?
And how useful has this policy been for the public?
Eg with other AI problems out there I can see a potential application to medicine, self driving cars etc, but I just don’t see what the bigger goal of this is going to be.
I have seen some very impressive pictures but nothing that seems to "emote"
The rest is up to the skill of the pilot. You can do a lot with color palettes for instance.
MidJourney's production model is good about dirtying things up, you could use it as a post processor. I think its default style of turning everything into fine art-cyberpunk-movie matte painting gets old though.
I've gotten some good atmospheric photorealism results out of DALLE, but currently trying to get outpainting to work to extend them and it's tricky.
Human being never characterize the world with linear rules.
This captures my thoughts very well. That's why these images get old very quickly - you can basically imagine the same thing in your mind. There's no real whimsy, surprise or creativity there. For now at least.
I encourage you to check out the subreddit where people share some prompts that you can reuse.
There's usually only a part that describes what is in the image, while half the prompt will be used to describe how does the image look.
Thinking of this one for example: https://www.reddit.com/r/StableDiffusion/comments/wp26lp/i_s... I have reused it successfully.
It's not a search engine. It takes a while to get a feel for what prompts work and how to phrase things. I suggest browsing https://lexica.art/ and using some prompts from there as a template.
If you zoom in on a DALL-E generated image the illusion falls appart and you can see it's a vortex of pixel that fuzzily creates a shape.
So far my experience with Stable Diffusion is that, no matter how basic the prompt, the ouput will be sharper and more coherent.
Linus didn't invent the core concepts of Unix. He copied the (arguably) good parts from an (arguably) closed ecosystem. His big innovation was leveraging the internet to create a new kind of community not really seen before it. The Linux bazaar gave smart developers excluded from the Bell Labs / BSD cathedral a place to be productive in that style on their own terms and others used it to disrupt the OS business. A lot of shit gets into Linux that the community turns into useful things.
I look at the inputs/outputs of Stable Diffusion and am reminded of the fractal craze of the 80's which revealed the underlying simplicity of incredibly complex looking things. Lots of interesting art and technology came out of that but I don't think anything like the Linux community did.
Stable Diffusion is open source, and was being compared unfavorably to a technically superior but closed source system.
This is definitely really important, but these systems generate output that is frequently off or incoherent in really important ways.
I think this is going to lead to some fascinating art, but stories of the death of art are wildly premature.
Blue hamsters filling donuts with nails
mostly edible donuts and quasi-donuts; some blue, but no hamsters and no nails. purple hamsters in free fall, eating gold nuggets
no free fall, some purple objects and backgrounds, somewhat metallic hamster hair, a five-limbed black and white hamster, credible gold nuggets without rodents, eating hamsters without food, multiple almost identical instances of eyeless white and yellow hamsters in the same pose. Martian hamsters wrestling
generic reddish hamsters, not wrestling at all. a hamster is the CEO of a financially struggling startup, in San Francisco
remarkably standard hamsters, in somewhat office-like environments (wooden tables and harsh light), with curiously unbalanced camera angles that might be random or inspired by the source material. the most beautiful hamster in the world, parading on a toy car and wearing sunglasses
No complete pair of sunglasses, but an interesting hybrid between round black lenses and round black hamster eyes; unusually deformed hamsters; an incoherent furry thing with a plastic part floating in the air; a deformed toy car with its back nicely replaced by a hamster. Vampire hamsters feeding on unsuspecting tourists in a dirty alley
No vampirism and no tourists, but nice dirty alleys.I insisted on hamsters because they are usually photographed in very few poses and activities, leading to completely ignored indications.
The system works much better in img2img mode. Slapping together a rough shape of what you want in GIMP/Photopea and then applying a prompt to blend everything together make the entire process a lot more reliable.
Ooooo. Ouch. That's... Kind of a death knell for a good tool in my experience, and suggests the tool in question is a glorified search engine in a sense. With a really confusing query syntax composed of words, and graphical starting states.
If I have to become <an expert on SD's training data> to get anything done... Why not just learn to draw/hire a creative?
If you want something unique and abstract, you're going to need to go through a lot of trial and error to get what you want. That's still a lot easier than teaching yourself how to create such art.
"Graphical starting states" in this instance isn't as hard, a very rough MS Paint picture of the general shapes you want things to appear in is enough. Alternatively you can grab rough cutouts from stock art, position them right, and the algorithm should figure out how to turn it into a single, flowing picture.
Take a look at these examples (https://huggingface.co/spaces/huggingface/diffuse-the-rest/d...), they're far from perfect but the autocompletion is done quite well. It should be noted that the demo application doesn't expose a lot of the flexibility the underlying model provides (like blend strength and such) but I don't know a free online alternative that does.
Trollface: Disambiguation of lines in the source image is poorly executed. The model appears confused as to whether those lines are indicative of depth, or lighting artifacts. The shape and perspective are poorly chosen, and in all the resulting images the lighting arrangement is quite inconsistent.
The ears are completely unspecified, so too the nose. This is somewhat of a deliberate omission in a trollface, and adding them in without careful thought as to how it changes the piece is... Well, not the best move. The eyes are terribly arranged in all submissions.
The plate of meat, fries, and beans. You can barely see the beans in the first sample, they are hidden underneath the fries, enough that an inattentive eye may miss them entirely. No specifi ation was given as to the state of the meat, or kind, so I suppose the being cut is a nice bonus. Interesting in a sense since one may get the impression the model may have confused the grammatical deep structure such that "fries and beans" was taken as a compound predicate.
The second with the meat surrounded by the beans is an interesting contrast, but without more samples, I have questions about why all the curated samples include rare beef instead of say, sausage.
The Colloseum: I too could use Photoshop, and select a particular palate. The more interesting aspect here seems to be the color pallete processing, and I'll admit that I wasn't able to find source works of the artist being initated to compare against. Still looking for those.
The Unicorn/Butterfly: These still disturb me in the sense that once again, we're replacing actual artistic technique, with the ability to tweak prompts or assemble graphical starting states/prompt combos. Is it making some hellish form of combined Natural Language/graphical programming pipeline? Yes.
However, none of this would have any value without being trained on works done by previous artists who likely were not asked whether or not they wanted their works included in the dataset.
As the guy who blew up a Philosophy of Art class by positing that a well executed forgery was as much a work of Art as the imitated piece, I still see here more problems than solutions. Yes, a new art form may have emerged. However, with it comes serious questions around data curation practices. As for the efficacy of the model/runtime characteristics/how this bodes for the environment... I'm increasingly concerned the more I apply ny "what if everyone started doing this?" supposition.
In short, see a hell of a lot of hype, but precious lite coming to terms with what will ultimately be the hard questions.
Here's a few of "martian hamsters wrestling" I generated:
Prompt (modified from a recent post on the /r/StableDiffusion subreddit):
A picture of 2 hamsters wrestling on the surface of mars, intricate, elegant, highly detailed, digital painting, artstation, concept art, matte, sharp focus, illustration, art by greg rutkowski and alphonse mucha
However, shortly, there will be millions of extremely detailed descriptions of new images… that is, a person puts in a detailed prompt, and if you can capture how pleased the user is with the result, you could then add that new picture — along with the user’s prompt — to your training data. Eventually, you will have millions of well-described images, which will make the system even more amazing.
BTW generating images using Anaconda on Windows using my RTX 2080 Super was very easy (followed the instructions here: https://rentry.org/SDInstallGuide ). The only thing was, because it has just (!) 8GB of ram, I had to use --H 256 --W 256 to limit the output image to its minimum size, which barely fit into its memory.
https://www.reddit.com/r/StableDiffusion/comments/wys3w5/app...
...but it seems to just be a copy of this scene from Aladdin:
https://www.youtube.com/watch?v=inzkJ34VMfk&t=20s
Am I missing something?
I think this applies to all AI generated images though.
What I'm going to do is take all those new image-generation AIs, and use them to train an image-generation-AIs-generation AI, this way I won't be bound by copyright anymore.
I wonder if it will be possible to train a neural network to do our programming tasks for us?
Apparently the computational power required to make them larger grows exponentially (or geometrically?) so until we find new algorithms it might be a long time before the same technique can generate whole films and games and any coherent way.
Of course maybe that breakthrough will be announced tomorrow.
In art humans always look for novelty and authenticity and not for mechanical reproduction, which quickly becomes generic. When every guy or girl on the planet could create a DeviantArt account and start drawing, did that have an impact on professional art? Not really.
We've increased the total artistic output and reduced cost several magnitudes over through tech, and if anything it's increased the demand for human novelty rather than reduce it. But as soon as the 'AI' label gets slapped on it people start to have weird Terminator fantasies.
I could see it becoming a marker of low-brow taste where vaguely related illustrations become something akin to decorating your home with velvet paintings of Elvis.
Is it because it represents that AI can do things that humans thought would never be replaced by AI before?
To restate my previous sentiment, I am quite attached to image generation with GPT-2, which is now legacy. Even though GPT-2 is not as advanced in many ways as more modern alternatives, I am artistically attached to its' shortcomings/artifacts. I really hope older image generation methods like gpt-2 remain available as a service to the general public.
https://cdn.discordapp.com/attachments/999426920376717513/10...
What I got was "NSFW content detected, please try a different prompt".
I must say it didn't exactly exceed my expectations.
But bloggers who are surprisingly often on the side of the current mainstream apparently think different --- for now.
"I mean, what stops us a life-imitates-art system where we have speech2img like in the Westworld ( the narative creating scenes ) ? I guess, I hope someone reads this and will pick up this. Maybe coupled with a VR set?"
Now I really hope on this :) Imagine giving this power to kids ( they cannot write but talk! )
"elon musk giving donald trump a massage using pizza sauce instead of oil, in a majestic room filled with flowers and golden toilets"
...and the result was a crappy AI-generated picture of not-quite donald trump holding a terrible rendering of a pizza, and some hands sticking out of random places. There were some red flowers in the picture at least.
I realize this isn't a "serious" use case, but clearly the tech isn't doing what I'm hoping it does.
I tried "coffee beans with cartoon mouths" and it's just a picture of some coffee beans. I don't get it.
"A cup on a plate".
And then replace cup for other objects. It generates nonsense pretty quickly. It seems most able to generate stuff close to something which already exists in a complete form. Sort of a pastiche machine.
That's the absense of reason sticking out: the "ai" mimics, but doesnt understand, so its creations are frankenstein-like bodies, almost human and for this reason more revulsing. So this "ai" represents an insane painter.
I wish people would stop exaggerating or making statements without some insight behind them.
I've seen hundreds, if not thousands of photorealistic result that range from acceptable to remarkably good.
One example it took seconds to find:
https://lexica.art/prompt/3fbc30ee-ca0f-42f6-8ea5-6890d469ab...
> [...] It takes a while to get a feel for what prompts work and how to phrase things. I suggest browsing https://lexica.art/ and using some prompts from there as a template.
"8k, highly detailed, ultra super very detailed, epic lighting, octane render"
When is somebody going to point out that this isn't Stable Diffusion’s fault, but capitalism’s? Why are we willing to stifle innovation for this instead of first making sure that everybody can live and create?
How often have you wanted to hire a conceptual artist? I have never done so in my life. While I have visited many art museums and appreciated it, this has not affected my commercial endeavors.
I think if I was a creative writer or professional artist, the current generation of AI tools could be useful for inspiration. But the bar for quality is just so, so high in the creative industries, I am not concerned that AI will replace them.
It's September 2022 - where is my self driving car?
Nearly impossible because logic doesn’t work in the real world like it does to make pictures.
https://en.wikipedia.org/wiki/Moravec%27s_paradox
https://metarationality.com/rationalism
I can think of two good businesses you could build with SD right now. Neither of them are text to image models or necessarily "art", both involve using it to visualize other kinds of sensing by transforming them to image embeddings.
I don't think AI will replace professional artists, but instead supplement them. Creation of Photoshop, Wacom tablets, 3d modelling and animation tools didn't put professional artists out of their jobs, it gave them new tools to create even better stuff in less time, and I think this is what's happening here too.
All that these algorithms are going to do is slightly bump up the productivity of some people. But that's amazing. Productivity is basically the free money tool. If people are more productive, they get richer for doing less, and the single best thing we can do to improve people's lives is make them richer. Poor people? Make them richer and they'll be happier and healthier. Rich people? Make them richer and tax the hell out of them, make the world better for everyone. Productivity is the best tool we have for making some people richer without making others poorer.
Agree that Stable Diffusion seems a bit over hyped. It is very cool! I am definitely going to play with it. I’m not sure really how much it’s going to change the world - there will be some neat apps and it may put a certain class of concept artist out of business (note I am not saying it will eliminate them - there will always always always be human artists creating wonderful art. But as a business model / way of earning a living, that might change).
I am going to disagree with you on GitHub Copilot. It has radically changed the way I do software engineering on my personal projects and it has increased the efficiency of the engineers at my company by at least 10%, maybe more. You should check it out. It eliminates an entire class of puzzle solving, and the least efficient kind (“how do I drop a column by condition in a pandas dataframe again?”). Simple answer via Google, but maybe 3-5 minutes of reading. 10 seconds via Copilot.
Even on your own terms, you're engaging in hyperbole. I've looked though thousands of these images and the harshest I can come up with is "assuming the result isn't obviously broken (i.e. some generated anatomy) then the results are sometimes a little bit uncanny or 'off'")
I could be wrong but this seems like an area where our progress will look like an S-curve. Getting that last 10% could be the real achievement.