4.2 Gigabytes, Or: How to Draw Anything
andys.page
andys.page
Basically, drawing sketches, editing (rudimentaly) in image editing software, img2img, edit, img2img, and a few more rounds, and you can get to something really, really cool.
This Photoshop plugin demo blew my mind yesterday: https://www.reddit.com/r/StableDiffusion/comments/wyduk1/sho...
Here's the before[0], and here's the after[1].
And an example of a step in-between: The base was [2] which I changed into [3]. That was I think the last step before the final generation.
It's not great by any means, but it's miles ahead of what I could hope to achieve myself. The biggest problem was to stop stable diffusion from turning my flying island into a pretty standard mountain. It kept trying to connect it to the ground. Especially in further iterations.
[0]: https://imgur.com/a/wKraDWn
[1]: https://imgur.com/a/Vn0RS9O
/coat
Posted video example here https://twitter.com/P_Galbraith/status/1564051042890702848
What actually happened is that it put mediocre graphic artists out of business and highlighted the difference between one that was mediocre and one that was good.
I feel like this will happen again here with digital artists. The mediocre ones will be indistinguishable from AI, but the good ones will still stand out.
I also wouldn't be shocked if a big portion of the market for this technology ends up being the artists themselves. I personally know a painter whose creative process has been overhauled by DALL-E. Brainstorming the next project is easier than ever, and unlike the DALL-E images inspiring them, the resulting paintings actually have the "human-touch" necessary to illicit the deeper emotional response that a good painting should bring about. Adding depth to a model doesn't necessarily add depth to the output.
But like I said, I don't think anyone really knows what's going to happen. We'll see, I guess.
YouTubers are a perfect example of that, it used to take entire television broadcast studios to do what they do, and now it can be done solo or with just a few people. And YouTubers are exactly the audience for this new tech - the smaller ones who want to put out branded merch can not currently do so at a level of quality. But you hire one artist to go out and make 5-100 pictures of your branding, tell them to take their time and make those few images to perfection, and that can now be molded to anything they want to create.
Indie game devs must be thrilled. Aspiring indie directors looking to make green-screen backgrounds are thrilled. VRchat users looking for 3d models are thrilled.
Though much of YouTube can be summed up as what was the three cheapest TV shows to produce - standup comedy, talk shows, and howto. The amount of YouTube sitcoms is a much lower number.
The path of a digital artist is long and arduous. For a time on this path, the artist may be considered mediocre, or to put it better, they are an apprentice.
Just as in other physical trades, an apprentice who is mediocre at their craft can still practice aspects of that craft well enough to be useful and earn some money. It is also through practice that the apprentice improves their skills. In this way, the apprentice is financially supported and even incentivized to improve at their trade, until one day they become truly good at it.
So what things like DALL-E and Github Co-Pilot and your clip art package do is displace the apprentice. With no path of mediocrity for the apprentice to walk, to earn a stipend for training, how then can they receive the financial support necessary to train until they're a master? They would need to already be independently wealthy or receive financial assistance.
In order to train more master artists and programmers, we would need to provide them with financial support while they train without us receiving anything useful in return.
This is the perfect example -- with a UBI the apprentice no longer needs to get paid to learn. They can live off of the UBI while learning, until they are good enough to charge for their services.
Just to illustrate what the problem is using an extreme example: Oh good, we made it so anyone can turn the whole of the earth's crust into paperclips with a push of a button in a fully automated way that doesn't require any human labor and the energy to do it is completely sustainable. Hmm... Maybe that wasn't such a good idea.
I'm not so sure it would actually happen though. We already give support to a lot of poor people through various programs like EBT and Medicaid. This just converts that help to cash, which gives people more freedom on what they want to spend on.
That is, I don't think UBI adds a new problem beyond "how do we make sure that humanity properly accounts for externalities" and "how do we make sure that AI does what we want it to do".
If you want a living, earn it. If you want wealth, earn it. Might not happen with your favorite school of craft, but the vast majority of people don't/can't make money doing something they are passionate about.
So far every experiment in UBI has shown that almost everyone getting it does something useful with the money and doesn't just sit on it.
And frankly, I have no problem with paying someone to sit on their ass drawing lines, if it means they aren't starving and homeless.
Why don't you? I am sure that you can support at least one such person with your income
Please don't expect everyone else to have the same generosity
When "earning it" takes much more than it used to due to technological shifts or otherwise, the only ones who can afford to walk the path toward mastery are the very well-off. This of course violates the modern western liberal ethos of equality for all, particularly in regards to educational pursuits.
We end up with a McDonald's worker class, their menial profession determined from birth, and their noble masters.
Maybe c'est la vie and there's nothing we should or even can do about it. But it's unpleasant, to say the least, knowing there's an entire class who's destined from birth to perform cheap menial labor their whole lives, without the slightest hope of doing anything else. After all, slavery is necessary for civilization, always has been.
I can accrue money doing what I'm doing - they can't.
Even if that only kill half the positions, we're still looking to a situation where humans overall don't have anything attractive to the market, if you can't earn a living wht would you do?
Quite a few of us already do that for people who don't even so much as draw lines on paper. (cough cough landlords cough cough)
No I don't. I receive a temporary lease to something of value that is fundamentally necessary for meaningful existence in modern society. Landlords are pure middlemen - and while there's a place for middlemen in society to provide initial capital, at some point that value dwindles down to zero as that initial investment is repaid, and then dwindles past zero as the landlord continues to parasitically rent-seek despite contributing nothing that the tenants themselves could accomplish for far cheaper.
Your retort to my retort would be valid (or at least actually equivalent to your cheeseburger analogy) if - in exchange for my rent checks every month - I received ownership stake in the property and/or the company that owns it. Such an arrangement has more in common with a housing cooperative than with a typical landlord/tenant relationship.
A lovely principle. We can start by taxing all inherited wealth and using it to compensate people who do essential but underpaid jobs, like teachers.
1) We need people to do low level jobs. So if UBI exists, wages will need to rise until people are willing to do them. This will happen along with price raises until an equilibrium is found where poor people need to work in order to survive. No need for narratives about landlords raising rent, though it is possible. The poor people aren't in an overall worse position here though, because although they're still earning just enough to live, a portion of that minimum is now guaranteed. However:
2) By raising your domestic (or local) wages/prices, you've just given yourself an absolute disadvantage against every other economic entity in the world. Anything that is outsourceable is now more appealing to outsource than before. This removes jobs and puts downwards pressure on wages.
If everyone just "lives off UBI while learning" society won't function because the jobs they do are important.
How many appliances do we build to last for a few years and then break? How many economic resources could we save by building fewer products to last longer? If the economic engine were tilted towards quality rather than churn, we could be much more efficient about our use of resources.
Durable products would be nice.
Curse autocomplete/swipe keyboards
Energy crisis in Europe is entirely man-made because, as I said, we are shitty.
Is the idea that it suddenly tips?
https://fred.stlouisfed.org/series/LNS11300001
For women, since about 2000: https://fred.stlouisfed.org/series/LNS11300002
https://fred.stlouisfed.org/series/LNS11300060
The shape is the same, with a peak at ~1999.
There is a large difference between the labor force participation rate and the unemployment rate because the LFPR can’t be gamed. A stark way to think about it is to consider that, if every person currently looking for work were to fail to such an extent that they just gave up, the unemployment rate would drop to zero.
Using your preferred metric, we have about the same LFPR we had in 1980. It has gradually varied since then between about 81% and 85%.
> The shape is the same, with a peak at ~1999.
The trend is up since 2015, apart from the Covid dip.
Automation and AI have not had a dramatic effect yet on LFPR. Maybe there's a tipping point to come.
The entire time, automation people like Yang have been claiming unemployment was already happening, and yet it was constantly going down. Unemployment in the US is now lower than it’s been in decades.
> We need people to do low level jobs
Is solved through automation and immigration (only citizens get UBI). Of course this is a major downside, because you end up with a slave class unless you make sure those immigrant workers are well protected.
Your point 2 has already happened. But the wealth still remains here in the US. So if that wealth were redistributed to the poor it would actually make things better.
Orders of magnitude more than we have now
> and immigration (only citizens get UBI). Of course this is a major downside, because you end up with a slave class unless you make sure those immigrant workers are well protected.
... Uh... yeah I would prefer not to have a slave class
What do you think we have now, where lower class people need to work to survive? If slavery is having the option between work or death, then I don't see how our current economic setup is not producing slaves. Sure, we don't treat the lower class as property (at least not normally), but they are definitely forced to work under any reasonable definition.
People are not property and they are most certainly not forced to work. Wanting things and needing money to buy the things you want is not “being forced”.
It's arguably the most regressive fiscal system one can imagine. A class of people who must earn an entire cost of living paycheck to net zero with their non-working, unskilled citizen equivalents.
Andrew Yang's premise is that those low level jobs are increasingly being automated away anyway - meaning that no, we don't need people to do them.
Even without that premise, this argument presupposes that people receiving UBI will do so in exclusion to working. That doesn't really logically or practically follow; it's just as possible that people will work anyway because extra spending money is extra spending money. They'll work because they want to work, not because they're being actively coerced to work.
> This will happen along with price raises until an equilibrium is found where poor people need to work in order to survive.
With Yang's proposal, the "price raises" part is probably true, yes. However, that has nothing to do with UBI; instead, it has to do with VAT. VAT advocates oft insist that it's somehow "not a sales tax" and therefore "totally not regressive like a sales tax", but at the end of the day consumers are paying more than they otherwise would for goods - and since consumer spending is disproportionately higher (relative to income/wealth) for the working class than the ownership class (or low/middle v. high, if that's the terminology you prefer), that's going to have the same regressive tax effects.
However, a VAT ain't the only way...
> No need for narratives about landlords raising rent, though it is possible.
Not if the UBI is instead funded by taxing the unimproved value of land - a.k.a. a land value tax, or LVT. We Georgists tend to call that a "citizen's dividend", but it's just a special case of UBI: a basic income intended to compensate citizens for occupying less than their equal share of land value within a given jurisdiction. There are a lot of implications of this (I could go on and on about the economic efficiency and ethical justifications), but relevant to this conversation is that the lack of deadweight loss means replacing other taxes with LVT would if anything reduce the consumer-facing cost of goods by reducing the effective tax burden of those producing said goods.
> Anything that is outsourceable is now more appealing to outsource than before.
That has already happened, without UBI. UBI is if anything necessary because of outsourcing - again, because we don't need local people doing those particular low level jobs, because they're now being done overseas.
UBI also might even help correct outsourcing; it's a lot easier to start a business if you know that if it fails (like most businesses do) you won't be homeless and starving as a result, and that's exactly the sort of safety net that UBI enables.
Well Andrew Yang is wrong. That's not what automation does. Automation reduces the amount of skill required to do jobs, reducing both the amount, but also the value. You still need people, and often more people because it becomes economical to employ poor people at a higher scale.
> Not if the UBI is instead funded by taxing the unimproved value of land - a.k.a. a land value tax, or LVT.
A land value tax is a great idea, but irrelevant to what I was saying. We need people to do low wage jobs. If they get some wages for free, we need to pay them more to do the jobs. If we pay them more, then we need to raise prices on the goods in order to not go bankrupt. The natural level of wages/prices is the one where people need to work in order to survive. The tax system and funding of the UBI is a separate problem.
> That has already happened, without UBI. UBI is if anything necessary because of outsourcing - again, because we don't need local people doing those particular low level jobs, because they're now being done overseas.
Economic Comparative and Absolute Advantages are not binary events. Doing things that make domestic businesses less competitive across the board in a globalized international economy is suicidal.
> UBI also might even help correct outsourcing; it's a lot easier to start a business if you know that if it fails (like most businesses do) you won't be homeless and starving as a result, and that's exactly the sort of safety net that UBI enables.
It's just a naive thing to focus on this founder idea.
Which translates to one worker being able to produce the same output as what required multiple workers previously. And sure, you could hire three entry-level workers at $15/hour for the price of one specialist at $45/hour, but chances are high that the same automation that enables those workers to do the specialist's job at all also enables said specialist to do considerably more than merely triple one's output.
Even ignoring the above, automation doesn't cause demand to materialize out of thin air; if you're a widget manufacturer and your sales team is able to sell 10,000 widgets a day, then multiplying the daily output of each widget factory worker from 10/day to 100/day will necessitate one of four things:
1. Figuring out how to multiply customer demand at the current widget price
2. Slashing widget prices
3. Slashing factory headcount
4. Slashing factory wages
1, 3, and 4 all minimize COGS and thus maximize profit margins. Unfortunately, 3 and 4 are both much easier than 1 (since 1 typically entails considerable effort to execute), so those are the options most companies pick. Both represent a severe loss of worker income - and thus, both necessitate UBI to compensate.
> A land value tax is a great idea, but irrelevant to what I was saying.
Assessing where the tax burden lies - and the impacts on that tax burden on spending ability, and the impacts of that on demanded wages - is pretty darn relevant to what you're saying. If you're paying an extra 10% (or whatever) on everything you buy, then you're going to adjust your wage expectations accordingly.
> We need people to do low wage jobs. If they get some wages for free, we need to pay them more to do the jobs.
Good. We should be paying workers a lot more than they're currently getting. The American (and for that matter, global) working class has been chronically shafted under capitalism for centuries now; God forbid we get shafted a little bit less.
> If we pay them more, then we need to raise prices on the goods in order to not go bankrupt.
Or the management could take a pay cut. I have very little sympathy for the "but what about our profits?" argument when C-level execs of even small businesses are skimming enough money on the output of our labor to be able to afford multi-million dollar homes and fancy cars.
> It's just a naive thing to focus on this founder idea.
Doesn't seem any more naïve than the idea that workers will somehow manage to "pull themselves up by their bootstraps" in a socioeconomic system deliberately designed to ensure we're never able to accumulate enough capital to do so (even at all, let alone without significantly impacting our physical and mental health in the process). Entrepreneurship currently skews hard toward those who already have money. That's a problem which in and of itself needs solved in order for a society to actually have any semblance of that "equality of opportunity" to which "laissez-faire" capitalists pay lip service; that maximizing the ability for working class people to start their own businesses (be it as individuals or cooperatively with others) happens to also at least partially alleviate outsourcing-induced job loss is a nice side benefit.
> Is energy really scarce? Currently the entire world consumes about 165.000 TWh (Tera Watt per hour) which is a lot, but it just a tiny fraction of the energy we receive daily from the sun, which is about 174000 * 0.7 * 3600 TWh = 430.000.000 TWh. On top of this there is all the energy stored inside our planet and atmosphere, which has been subjected for billions of years to the sun’s energy transfer.
By this logic energy has never been scarce. Energy is one the most important scarce resources. Much of the world is currently undergoing an energy crisis.
> Is food really scarce? No, as it can be obtained by a mix of energy and chemical elements.
We have the sun, and there is soil, ergo food is not scarce.
> Are chemical elements really scarce? Let’s pick gold, an element which is notoriously considered rare. The total gold mined in all human history is 200.000 metric tons. But if we look at the abundance of elements on earth, even though the mass fraction of gold is just 0.16 part per million, knowing earth’s mass we can estimate the total gold on earth to be about kg * * 0.16 = metric tons. If we only consider earth’s crust, that’s about 100 times less, which is still a huge number.
Minerals that require more energy to retrieve than value they provide.
You state
> Now I’m going to state something which may either hit you as a profound insight or as an obviousness. Basic resources are not scarce per se, what’s limited is the ability to transform them and make them usable. The fact that we need a human to perform the job is what creates scarcity.
It's a really poorly reasoned thesis and your arrogance to call it a profound insight is just bad. If you have one human and you have a water pump that requires two humans' labor to retrieve one human's water, it is nonsense to say you have a labor shortage. You have a water shortage.That one resource can be used to acquire another does not meant there's only one resource on the board. Everything you've said about labor could be restated as useable energy.
Energy, food, materials, labor, land, time. It's all scarce.
A UBI also opens up the possibility of removing the minimum wage which not only allows for more people to obtain jobs, but also raises the competitiveness with other countries, potentially (it depends on whether the minimum wage is actually effective in raising wages above the market rate).
How do you incentivize people to work? Pay them more.
If mass unemployment was a substantial problem, this may well be an acceptable tradeoff, but in the current economy it is not.
If you get just enough to survive from the government, and employers try to reduce wages because you'll net out the same, you probably won't accept the job. It's not worth your time. Even as a dirt poor person, your time is valuable. An employer needs to pay you more so that it's worth your effort again. You have much higher freedom to shop around too.
If the government does not pay you enough to survive (non basic income), you're still in a precarious position, but you are still in a BETTER position than you were without the funds. This will RAISE wages, which will in turn RAISE prices, until the equilibrium is found where the UBI is distinctly not sufficient and you need to work to survive. You will still be poor. You will make more wages, but have about the same level of real wealth. You will be a bit safer due the guaranteed portion of income.
Until globalization kicks in and makes your specific local circumstances much worse.
We aren’t even close to automating basic needs. We can certainly automate the manufacturing of some complex individual items though. Seems like a pretty fundamentally flawed assumption of UBI.
UBI is the solution, not the problem - and on the topic of Native Americans, funding it can and should start by taxing the unimproved value of the very lands we conquered from them.
Cool, now that we all have $250/mo[1] for the rest of our lives, our problem is solved.
1. This is to a rough approximation what we can maybe sustain in the US now.
What's the market for mediocre art today? I long ago worked on tech for magazines, which would sometimes use adequate commissioned art to jazz things up. But that was before the rise of vast stock art collections that were instantly accessible. Looking at some popular web-based magazines, it seems like the still commission the occasional original illustration, but that it's mainly stock photos or photo-composite illustrations.
I guess the same happened with those creating and animating 3D models for the likes of Toy Story.
There was never a market for mediocrity... but people will happily pay for exposure (to play in a bar, rent a space as a gallery and so on). The problem is that even for good art it's hard, and it has always been. The rise of accessible stock art doesn't help, and AI will not.
Still one point is important: if you want to create something new, and not reassess (derive) the same thing, I guess we (human) have still a place. At least for now.
As an aside, I'm not sure I'd use the term "apprentice" -- I'd maybe say "junior" as in "junior developer" or "junior designer". They're learning how to make good work, but they're still a professional in the field.
Presumably we need fewer of the new more efficient jobs to displace the work done by the old jobs. But that lowers the price of ”mediocre” graphics, and therefore increases demand; maybe we actually end up with more jobs at the more accessible level.
These market dynamics are quite hard to predict but either way, bad time to be an entry level professional artist, great time to be just about anyone else.
Look at music production: historically the barriers to music creation were the dexterity and practice needed to master an instrument, or multiple instruments.
But when you put tools in the hands of more people without the filter of needing the time, training and skill to coax the sounds in your head from the instrument you have… and suddenly you get stuff nobody was making before.
I mean. Do you think that there are people out there with good taste in software and imaginative ideas, who can’t pass the skill barrier currently required to write code?
Why would art be any different?
Software seems like mostly thought-stuff to me, there’s no mechanical skill-based piece. But even so, when AI can generate a full app with the ease of iteration displayed in the OP, then sure, I think you’ll see some people with app ideas generating those apps themselves, instead of having to hire contract developers. Right now just completing stubs of functions doesn’t seem useful enough to allow someone that can’t code to make an app.
Over the history of humanity the printing press, the photograph, the computers etc. destroyed some profession only to make something else flourish.
Yes, for a day or a week. The AI wont stop scanning and learning.
Scenes with more than one person in them, or multiple people and objects interacting in complex ways, are the most obvious cases.
I’m sure the technology will keep improving but I think it’s possible to be overly optimistic as well.
20 years from now they'll click a button and you'll get a fully randomly generated pixar movie thats as good as the originals.
Im software engineer and amateur illustrator. I have always welcomed technology. Always felt good about it. Copilot? No worries, please automate my job, if humanity dont code anymore, I dont think its gonna kill us the slightest on the contrary. Art? Mark my words, this wont go well with people souls. This is an obvious evil mistake. Im still confident some wises will stop this heresy before civilization collapses (I like to dramatize like that but still this is bad imo)
I'm also the type that welcome new technology into my workflow, always one of those early adopter, but I have a hard time this time
...or maybe I'm just overthinking (Life finds a way!, right?)
Please rephrase I dont get it. Do I not sound sincere? Rest assured I absolutely am.
StableDiffusion was trained by an academic group in Germany and they do believe it’s legally compliant though.
It's like software development.
Back in the day you could make good money with websites. Knew HTML? You got the job!
But the goalposts are constantly moving. You had to add CSS and later JS to your skills to keep up.
Today, you won't get paid for writing HTML anymore.
Same for other industries.
You have to move with the industry, gather new skills, etc.
Also, it will be interesting to watch whether and how people will be assessing art created by AI. Will there be something like Connoisseurship for AI art?
Here are the samples https://imgur.com/a/XMnMi
Then I found that the end of the era of handmade digital art is coming. Only transistors are limited and future digital artists will differ in memory size and teraflops like bitcoin miners.
the all-time classic in the genre:
https://slashdot.org/story/01/10/23/1816257/apple-releases-i...
The creative process, like many kinds of fields applicable to business is about problem-solving. AI text-to-image generation doesn't replace that function, however it does form an excellent tool, especially when it comes to rapid conceptualisation. This will allow more people to be creative problem solvers without needing to possess technical skills in image creation. Much in the same way that graphics apps allowed people to make image without needing to learn studio or art skills. Or DTP tools allowed more people to publish without the tedium and high set up costs. I will still be hiring illustrators and designers, and this may be one of their tools, and it would be their responsibility to be experts in it, but it doesn't replace them - it makes them better illustrators and better artists. The right way to think about this is not that it shrinks the field, rather it opens it and accelerates it - for that it's a very welcome addition. No creative is scared of this - they're looking forward to the next-generation approach, and it's clear that 2D images are not the end point. Soon we'll have 3D (already in progress), soon we'll have music, soon we'll have this for animation and programming.
People fiddling with the technology have noticed some obvious short-comings, such as getting consistent results - for example it's currently not possible to develop a series of story boards where the character is obviously the same. Instead some level of reseeding the image with the desired character or outright recomposing the graphics later is needed. These aren't things that can't be fixed however, what we're seeing now is definitely an exciting new tool in its infancy.
When most people can self-serve for pennies, demand will drop. Random indie game needing art assets? Done. In days, not months.
Travel agents are actually a great weird example of what might happen... we didn't really get rid of 99% - more like 80-90%. They just stuck around as good salesmen who use Expedia etc better than the average person and (in theory) destress the process. Just like artists of the future almost certainly will be quite good with these AI tools and use them primarily - but maybe have a bit of artistic talent themselves.
Back to the subject at hand, and as other people have already said, the generated images seem to lack emotion and "feel". They're good to maybe put as wall-art in a rented out AirBnb, but that's pretty much it. Still a cool thing from a technical perspective, though.
For comparison, the human brain is estimated to have a (equivalent to a computer) capacity of about 2.5 petabytes [0].
I think I read in the past that the human mind holds memories in a picture like way, I wander if that's why these image based models are so incredible when compared to the text models.
Maybe we are in a new "Moore's Law" like period where the complexity and size of these models is going to double something like every 18 months. It's going to be fascinating what's possible in a few short years time, I fully expect to be continually surprised.
I'm looking forward to seeing a video model trained on 10 second clips, someone somewhere is working on it.
0: https://www.scientificamerican.com/article/what-is-the-memor...
I’ll hold off declaring it dead till it is well and truly dead. And even then we could expect cost improvements as the great wheel of investment into the next node would no longer need to turn and the last node would become a final standard.
As to physical limits, there are plenty of weird quantum particle effects to explore so that seems overstated. We are still just flipping on and off electromagnetic charge. Haven’t even gotten to the quarks yet!
The classical Moore's law formulation has been dead for 15 years already. What we have now is whataboutism about why it still holds.
https://www.researchgate.net/figure/Density-of-logic-transis...
If you are noticing that this seems to fundamentally limit model performance on certain tasks to aggregate human capability, you are noticing correctly.
To give you some idea of what these benchmarks look like, here’s the prompt list from DrawBench which Google created as part of training their Imagen model.
https://docs.google.com/spreadsheets/u/0/d/1y7nAbmR4FREi6npB...
Model size scaled 100x faster than compute over a decade. We are paying for this difference by using more energy and hardware, but it's already too expensive to train except for a few labs, and deployment is restricted.
Can't even load GPT-3 on most computers. Stable Diffusion is an exception, they did a good job and were lucky the model can be so small.
First 'how good' is an ill-defined metric -- that it seems in this case is a measurement of how much surprise and wonder they generate in the audience.
Second, it might just be that the real wonder of the models is the their compression -- that is there is a space of mappings of line drawings and simple descriptions into art and this technique was able to lossy compress that space down to 4.2G. If you only compress it down to 42G, you'll be looking at the JPEG that's 90% compressed instead of 99%. Yeah it will be better but incrementally, not necessarily "Wow!" better.
Honestly it's not obvious that it will be better at all.
> For comparison, the human brain is estimated to have a (equivalent to a computer) capacity of about 2.5 petabytes [0].
That's a terrible and basically non-sensical comparison to make.
It was already a genre that highly incorporated computer assisted methods. There is a lot of doom and glooming going around, but honestly the modern process of creating 'concept art' was already extremely commodified and efficient. These weren't exactly your idealized vision of some artisan craftsman laboring weeks over a picture, they churned this stuff out in a few hours (if that)
These models let anyone achieve similar results in minutes. Without any prior learning. It is not even lowering the bar, it is literally dropping the bar to the ground.
Besides, stable diffusion is able to generate not only painterly scenery, but also photographic images that are almost indistinguishable from actual pictures (certainly helped by the fact images have a heavy digital look in our era).
From https://simonwillison.net/2022/Jun/23/dall-e/
Have you tried DALL-E or Stable Diffusion yet? I bet you could generate a black and white image that met your standards for being impressive, if you spent a few minutes on it.
You can try Stable Diffusion free here: https://beta.dreamstudio.ai/
So yeah ANY style, I’m pretty sure of it.
AI really doesn't handle styles with restrictions like that well. I tried the stable diffusion website with variations on "black and white silhouette stencil image of a cat". It kept wanting to give the cat colored eyes, or it used shading, or the cat didn't have a coherent anatomy, along with the typical AI art "duds" that aren't really anything at all.
To be fair, I did get a couple of passable results when I replaced 'cat' with 'dog'. They were simple, but didn't have any obvious errors.
To be fair in the other direction, replacing 'cat' with 'abacus' gave me an (admittedly pretty) grid of numbers and some chainmail, and 'helicopter' suggested a novel design where two helicopter bodies would be stacked vertically, connected by a vertical shaft through the rotor, and which turned into a palm tree trunk above the top unit. Once you get out of the sample data, it starts to fall apart.
I feel like other people here are willing to forgive more errors than I am. They see an incoherent splotch in an image and assume more development can iron out all the problems, and I see a unavoidable artifact of the fact that these systems don't have a real understanding of what they're making.
Edit:
I gave this another shot to see if I could get a more complex stencil. This was my very first try again, so truly not cherry picked. Prompt was: “Stencil image of a tiger face. Clip art. Vector art.” This looks like an infinite stencil making machine to me
The second one is representative of the upper end of what I was getting. It's almost passable, but doesn't hold past about five seconds. The left and right half don't look like they belong to the same animal. The blank space in the middle of the face is huge and detracts from any sense of structure, and the whole mouth area is just odd.
It's seriously impressive for AI, but it's not end-of-artists type stuff. I can google "tiger black and white stencil" and get a bunch of tiger faces, and every one of them is noticeably better. People imagine there are plenty of art jobs where discrepancies like this don't matter, but there really aren't.
Again those were my first try and I know nothing about stencil beyond what 2 seconds on Google images could give me. Certainly better than I could produce if you gave me Adobe Illustrator and a weekend. And the image is mine to use as I please, unlike what I could rip off Google Images.
Also, I thought the cat was cute, but there’s really no accounting for taste. Here’s a silly and swirly cat that might be more your thing? This was a cherry picked one of 10 since you have high standards;)
I could run pretty much any of these through adobe illustrators auto trace and end up with an amazing vector image.
I could also leave it generating these for an hour and I'd have over 1000 results to choose from.
Today I can get quick, effortless renders from Blender with a zillion available assets on the internet on my laptop. I can drop that directly into something like Clip Studio and paint right over it.
In the 80s you needed an extraordinarily expensive workstation like the Quantel Paintbox to even do primitive Photoshop type stuff. If you wanted a 3D render you needed a whole farm of servers.
So quantity is indeed a quality.
The neat new applications that have taken over this site for the last couple days sometimes require CLI steps to install because they are in active development and it can be easier to experiment with something local. I'm sure they'll either be moved online or wrapped in nice installers over the next couple weeks.
4 out of 5 people globally would be able to submit a stable diffusion prompt and view a result. Most would have no idea what the hell was going on or even why it was interesting.
This is the funniest part to me, because so many people already think this is how digital art worked to begin with.
Using a CLI-based tool is inaccessible for most people... but building a GUI around this would be very easy. I'm too lazy to google it, but I would bet someone already has a GUI, or is working on one.
12GB of VRAM may not be accessible on most computers, but there's nothing innovative about offloading that task to an EC2 instance. It just requires an opportunistic developer to tie the pieces together.
I would be monumentally surprised if Figma/Canva/InVision/Adobe are not already working on this.
CLI-based tools are perfectly accessible to most people.
They just can't be arsed to learn them, unless they need to. And most of the time, they don't, because good-enough alternatives exist.
If a CLI-based tool is the only way that an average person can get their work done, that's what they'll use.
There's a WebUI with a docker container if you're on Linux w/ GPU; https://github.com/AbdBarho/stable-diffusion-webui and https://github.com/AbdBarho/stable-diffusion-webui-docker.
If you don't have a GPU, there's a Colab UI (Google hosted GPU). https://github.com/pinilpypinilpy/sd-webui-colab-simplified
Think about it, give the user a few basic, MS-Paint level pencil tools, colors, shape makers. Ask for a description, the application can even push you in the right direction for putting together good, detailed prompts, gives you a list of art-styles, artists, filtering methods, etc all with reference images so you don't need to memorize names. You can zoom into sections of the image to work on independently (like the birds in the article), then blend it into the greater image. Drag and drop image files onto the project and iterate on them.
Implementing the glue to simplify the "tough" parts of this process is honestly pretty trivial.
Probably not many in general, but the RTX 3060 has 12GB or ram and it is around $350. And I saw a RTX 2060 12GB for $250 the other day. That's a pretty reasonable entry fee IMO.
Itll comfortably run on 6gb now. gtx1600 series cards need to run in full precision mode to produce output. The HLKY fork has improved the Gradio GUI and integrated realesrgan and gfpgan for those with beefier cards.
Someone else also figured out how to load and run it all on a CPU, so pretty much anyone can in theory run the model now.
There is an elaborate Colab notebook linked in the HLKY repo that seems to get more point and click user friendly every time i look at it. I think it even launches the gradio webui so you can use the Colab instance with a webui remotely.
Yet. This is a huge leap forward, to get more basic prompts generating things will be a much smaller leap IMO.
This is more like the literacy/printing press transition.
Used to be, people had to learn to memorize a lifetime of stories and lore. Now nobody learns to make a memory palace or form a mnemonic couplet. Why would you bother? You just write things down.
Today, people learn to draw. In a generation, why would you bother?
There will still be specialist jobs for people generating images, but instead of learning to make them up, the specialists will be very good at picking them, suggesting them, consuming them.
Humans will be the managers and the editors, not the creators.
The same thing will happen to other arts. First (and very easily) to music. Eventually, perhaps, to writing and whole movies...
The only thing stopping that is that the models can't maintain a reality between frames. They can't make an arc. It's all dreamlike.
If we find a way to nail object persistence it will be a singularity-level event. The moment you can say "make another version of this movie, but I want Edgar to be more sarcastic and Lisa should break up with him in the second act" we will close the feedback loop.
It's a lot bigger than "lost jobs".
I agree. It is more than just "lost jobs", like artist impressionists, court room sketch artists, etc. it is a complete dystopia and it doesn't help artists at all, but displaces them. At least the value of actual paintings will be more valuable that the abundance of this highly generated digital rubbish.
So given that the technologists have so-called 'democratized' and cheapened digital art, I really can't wait until we get an open-source version of Copilot AI that would create full programs, apps, full stack websites with no-code so that we would be seeing very cheap Co-pilot AI shops in the south east of the world generating software that effectively eliminates the need for a senior full stack engineer.
Easy cheap business solution for the majority of engineering managers on a tight budget who know they need to offshore tech jobs without the need for any skill as it is offloaded to cheap Copilot prompters.
So we will have no problems with that and be happy with that dystopia. Wouldn't we?
But I think society will find a way. Who knows, maybe we'll all work less and enjoy life more? One can hope.
Being an artist isn't really being able to draw well, it is able to do a lot more than that in harmony, and so I believe these tools will just get incorporated and some new artists will appear and older artists will adapt.
My only worry with this, and it's not something that I see being pointed out too much. Is that due to these models being able to produce art from previous art they've seen we might find it difficult to come up with new novel styles. But then again, this might precisely be a new kind of avenue for human artist expression.
In the realm or “real” art I’m actually very excited since I believe there are hundreds of very imaginative and patient people who just can’t paint well but will be able to create new art with tools like this. It can also synthesize new and alien things.
A race to the bottom and the cheapening of 'art' in general for the sake of replacing artists is a shame to see and nothing to celebrate. I was against both the gatekeeping of GPT-3 and DALL-E by Open 'faux' AI. But now it seems that every-time an open-source alternative or version was released into the wild, it seems that the uses become even more dystopian; especially with DeepFakes, fake news propaganda / articles and catfishing with generated hyperrealistic faces.
> And it’s not NFTs. Remember that last year this would have sounded mostly like sci-fi unless you were following cutting edge research.
Stable Diffusion is the reason why JPEG NFTs will always be worthless. Both of them will fuel JPEG NFT prices to the floor value of zero. But as NFT proponents cheered in believing that they will help artists, here we are seeing DALL-E 2 and Stable Diffusion fans screaming that it will help artists. No it will not.
> In the realm or “real” art I’m actually very excited since I believe there are hundreds of very imaginative and patient people who just can’t paint well but will be able to create new art with tools like this. It can also synthesize new and alien things.
This isn't the 'democratization of digital art', it is the complete devaluation and displacement of digital artists and it now makes 'real art paintings' much valuable.
A dystopian creation.
So are you still against gatekeeping? Are you in favor of releasing AI advances to the wild?
Even with the release of GPT-3, there seems to be very little good in such a system despite it being generally underwhelming at generating convincing sentences. However with DALL-E 2, it has gotten much better for worse on digital images, to the point where even gatekeeping it would spur on an open source competitor superseding DALL-E 2 anyway.
But it was actually after the release of Stable Diffusion that done it for me when most here hyping just want to aid the race to the bottom and at the same time are screaming that it will help artists when (like NFTs) it won’t.
So looking at both DALL-E and Stable Diffusion, it is yet another contribution that advances the dystopian AI industry which will just be used for fake news, surveillance and catfishing. Worse part is that they haven't built any detectors for this.
Given that there are more powerful models that have already been developed, should they be gatekeeped or released?
> I am still against OpenAI’s gatekeeping and gave AI itself a chance to be more used for good and significantly less dystopian.
> Worse part is that they haven't built any detectors for this.
So it is neither. If a given AI project has no detectors or a straight indicator of knowing that it is generated by an AI, then the whole project should be effectively scrapped and cancelled, postponed, etc until it has one. It is that simple. And No. DALL-E 2's tiny watermark doesn't count.
'AI researchers' know the dystopian scam that they are creating and they know that they need detectors and analyzers for them to significantly reduce the risk of malicious use. So it doesn't matter if there are others that are more powerful as the conditions are still the same.
I think you're directionally correct, but overstating the case in a few ways.
One, as a not particularly visual person, even this example involves some skills of composition and perspective. If you asked me to do something practical, like creating an illustration to go at the top of a blog post, I would not do nearly as well as somebody with art skills, and I would take a lot longer.
Two, this is the beginning. In the same way that digital artists took tools I could use and got really good at them, I expect the same will happen here. What will a good artist be able to do with a solid workflow and a few years of picking up tricks? Given the opaqueness and quirkiness of models, I expect a person who puts in the time, especially one with a strong command of art styles, composition, and the practical uses of visuals, will be able to run rings around me.
Three, people are quite accepting of AI images right now, but they're novel and exciting and decontextualized from how we normally use images. That's a playing field that advantages the novice. But what happens once these images are no longer fun and novel, but boring and overdone? As we learn to discern novice-grade work from what real artists can do with AI assistance, I think our bar as image consumers will rise.
Usually it gets downloaded into someone’s mind, triggering some kind of cascade of baffling imagery.
Feels kind of odd that this model data actually… works sort of like that?
But only so much can be encoded through history. I'd love to see a sci-fi movie combined with the butterfly effect. A somewhat advanced civilization in the past and another (maybe present) civilization where people try to find out stuff about the other one, maybe they're successful at the beginning at depicting how they were but they start to think they know everything and the whole perspective of the civilization changes.
https://twitter.com/EMostaque/status/1564655464406650881
With all the model optimization, distillation papers already out, a few hundred Mb doesn't look impossible with similar quality outputs.
function closeModal() {
let modal = document.getElementById("myModal");
modal.style.display = "none";
window.location.hash = '';
}
// Left and right arrow keys scrub through the gallery.
// Everything else closes the modal.
document.onkeydown = function (e) {
if (e.keyCode === 37) {
// Left arrow
goBackPrevImage();
} else if (e.keyCode === 39) {
// Right arrow
advanceNextImage();
} else {
closeModal();
}
};Don't get me wrong. I'm sure there is a skill and what I'm seeing in demos is the happy path where it all happens to work well. But damn it's impressive.
Companies could be hired to develop a particular keyword over a period of weeks or months to allow for more specific prompts.
Love that idea. The prompt economy.
That's actually built in to some of these systems - Google Imagen for example generates a 64x64 image and then up-scales it to 1024x1024: https://www.assemblyai.com/blog/how-imagen-actually-works/
StableDiffusionPipeline.from_pretrained("CompVis/stable-diffusion-v1-4",
revision="fp16", torch_dtype=torch.float16, ...)In some domains this will almost certainly transform the way people create art, and it will feel like it's happening overnight.
Digital artists will surely adapt and use these technologies to their advantage, but transitioning will take time.
Personally, this also feels like merely a first glimpse into a world where humans use AI-based "cultural technologies" (to borrow Alison Gopnik's term) to "fill in the blanks" on ideas across many domains.
What we are seeing now is likely just the first pass at what this technology can do.
Great article about one of the little known features of stable-diffusion. The img2img.py awesomeness that turns your 4 year olds finger paints into Picasso or Monet. It’s just mind boggling!
1989: Coding for the first time (Apple IIe)
1994: Getting on the Internet
1999: Using Google when it started to get really good
2005: Buying GOOG Stock
2014: Buying Tesla Stock
2022: Building a local Pipeline with GANs, SD, Nvidia 1660
Things are about to go nuts. Tim Urban explains it well. https://waitbutwhy.com/2015/01/artificial-intelligence-revol...
The technology might be there but the capital structures...are not?
If this isn’t crazy then what is crazy?
What is the initial drawing being done with?
What is this stable diffusion process? (Amazing is one answer.)
A very brief description of the software used.
Color me astounded. I actually liked (was delighted by) the results in step-6. When I realized it was done programmatically somehow, I had to pick myself up and get back in my chair.
Thanks for the explain.
Try for yourself a bit. Here’s my current favorite interface to Stable Diffusion. Give it a little sketch, then go wild with the prompt description. Try different style descriptions like oil painting or comic book.
> these models were trained on image-text pairs from a broad internet scrape
… yep.
I've the same issue with this as with Github Copilot.
I will admit, it is technically impressive, and something I would love to use, as someone who cannot draw worth a darn. And it is that I cannot draw that I do not feel morally comfortable with this: I am using a — complicated, admittedly — tool to just derive art from the unwilling talents of others. (Admittedly, my skill in prompting & editing might matter, but that's true of "normal" derivative works, too!)
People have shown fairly convincing examples of this in the more general sense: e.g., they've had well-known stock image (e.g., iStockPhoto) watermarks get produced in the output from the AI models (when not prompted). An artist with "experience" would not reproduce a watermark. Or in this article[1], where an AI was requests to mimic another artists style, and the output was (attempting to) reproduce the artist's signature.
(IANAL.) If you make a film that directly incorporates aspects from Star Wars (what I believe to be the more accurate version of what these models do), then yes, I would expect that you will be handed a C&D. "Glowing space swords" aren't copyrighted, but if you include something indistinguishable from a lightsaber & call it a lightsaber? I bet Disney would have something to say about that.
[1]: https://kotaku.com/ai-art-dall-e-midjourney-stable-diffusion...
1. Find a photo of a person as reference 2. Create portrait 3. See how well the portrait compares to the reference and the stylized art I was drawing inspiration from.
The work I was doing was original in colloquial sense, but also I see zero reason why what the AI's process is fundamentally inferior to mine.
But broadly the courts have upheld the rights of companies to use copyrighted works as inputs to commercial algorithmic derivative works like neural networks.
Now you might argue this doesn't apply here. A key aspect of the decision rested on the fact that the original copyright holders (book authors & publishers) were not directly harmed by Google's indexing of them, since it probably drove more sales of those books. In this case it's not so clear. Is somebody using a diffusion model doing so instead of buying a piece of commercial art? If they're generating a new piece of art, I'd say probably not. But if they're generating something specifically similar to an existing specific piece of art, perhaps, but if it's deliberately different, it's still a tough argument. If the ML model is being used to deliberately replicate a specific artist's style, then I think you can make that case pretty strongly. But if you're building something that's an aggregate of a bunch of styles (almost always the case unless you specifically prompt it otherwise) then I don't think the courts would find that any damage has been done, and thus nobody taking this to court would have standing.
I think it's likely we will see this end up in the courts somehow. But being able to prove actual harm is critical to the US court system. And it's difficult to see how the courts would rule against the kinds of broad general use that is most common for this kind of generative art.
> Now you might argue this doesn't apply here.
Indeed, I would. In particular,
> and the revelations do not provide a significant market substitute for the protected aspects of the originals
I'm not sure if that holds here. In Google's case, the product (a search engine) was completely different from the input (a book). Here … we're replacing art with art, or code with code, admittedly different art. And … uh, maybe? different code. I'm also less certain due to the extreme views on what constitutes de minimis copying the courts have taken.
> I think it's likely we will see this end up in the courts somehow.
I agree.
> But being able to prove actual harm is critical to the US court system. And it's difficult to see how the courts would rule against the kinds of broad general use that is most common for this kind of generative art.
This is a good argument, too, though I'd like to see it tried in court, I think.
> If the ML model is being used to deliberately replicate a specific artist's style, then I think you can make that case pretty strongly.
I'll link the same example I linked in a comment, [1]. Seek to "On the left is a piece by award-winning Hollywood artist Michael Kutsche, while on the right is a piece of AI art that’s claimed to have copied his iconic style, including a blurred, incomplete signature"
[1]: https://kotaku.com/ai-art-dall-e-midjourney-stable-diffusion...
The benefits of this are clear, but the problem is that artistic expression and being able to receive small-scale rewards and genuine encouragement—at least in one's family or social circle—for even minor talent seem to be very healthy and fulfilling things for people to do. Taking that away came at an ongoing cost that none of the beneficiaries of that change had to pay. A kind of social negative externality.
Relatedly, consider the sections of Graeber's Bullshit Jobs where he treats of the sort of work people tend to find fulfilling or are otherwise proud to do, or are very interested in doing (sometimes to the point that supply of eager workers badly exceeds demand and pay is through the floor, as in e.g. most roles related to writing or publishing). What kind of work is it? Mainly plainly pro-social work (to take one of Graeber's examples, the disaster-relief side of what the US military does, which is by no coincidence often heavily featured in their recruiting advertising; or teaching, for another one) or: creative, artistic work.
Graeber notes an apparent trend whereby these jobs tend to pay pretty poorly either due to the aforementioned over-supply of interested workers, or because there's some societal expectation that you ought to just be glad to have a job that's obviously-good and accept the sacrifice of poor pay, and that you must be bad at it or otherwise unsuited if you want to make actual money doing it (teaching's a major case of the latter—I've seen that "if you care about being paid well you must be a bad teacher" POV, and the related "if we raised teacher pay it'd result in worse teachers", advanced on this very site, more than once—it's super-common).
People really, really want plainly-good and/or creative jobs, but those don't pay worth a damn unless you're at the tip-top, either of talent level, or of some organization. This seems like another blow to the creative category of desirable jobs, at least.
My point is: I wonder and worry about the effect this latest wave of AI art (in a broad sense—music and writing, too) generation is going to have on already-endangered basic human needs to feel useful and wanted, and to act creatively and be appreciated for it by those they're close to. There's already a gulf between the among-the-best-in-the-world art we actually enjoy and, should our friends present their creations, how we "enjoy" those these days, with the latter being much closer to how a parent enjoys their child's art, and everyone involved knows it. Used to be, hobby-level artistic talent and effort was useful and valuable to others in one's life. Now, that stuff's just for yourself, and others indulge you, at best.
Why, with this tech, you can't even get by doing very-custom art, such that the customization, rather than the already-devalued-to-almost-nothing skill itself, is what delivers the value. Now the customization is practically free, too, and most anyone can do it.
Getting real last-nail-in-the-coffin vibes from all this. I'm sure it'll enable some cool things, but I can't help but think we're exchanging some novelty and a certain kind of improved productivity, for the loss of the last shreds of a fundamental part of our humanity. I wonder if we'd do this (among other things) if we could charge the various players a fair value for harm to social and psychological well-being that happens as a side-effect of their "disruption"—alas, that pool's a free-for-all to piss in all one likes, in the name of profit (see also: advertising)
I've been blowing minds amongst friends and family for over a week now with my Midjourney creations. Just saying.
I feel that this apocalyptic perspective is just us older folks trying to grasp how we would do things we like under different conditions. Kids will find a way to make a career using these models, and maybe now some not-so-skilled but extremely smart person can finally show us their creations in a meaningful way.
And for me personally, I find extreme beauty in these technologies. A fundamental part of our humanity, artistic expression, and we can manipulate it like this? Another unique human trait, language, how enormously fortunate I am to see another part of that puzzle being worked out and reduced to math.
I'd love to follow this process to generate several that I have in mind to put up on my walls, but I keep running into the resolution limitation. You need a pretty high resolution to get a crisp image at a poster size. Is there a trick or a setting to get the models to output images suitable for posters?
With img2txt you could give it an audio file, call it "S" and tell "music in S style, but with flute".
So even though to a casual glance, these images look amazing, if you look closely, you see all the flaws that come with them being generated. Weird artifacts, a lack of symmetry that humans usually add to their creations.
These flaws would not exist (to such a degree) when a professional artist paints the same scene.
But it also dillutes their own ideas. I know of painters painting AI generated stuff and that is likely a new genre but a lot of genuine artwork will lose interest on the market. It is what it is and don’t I love or hate it…
Say you know you want to do a portrait of a woman in armor, you can generate a dozen of those in around a minute on a 3090, look at the generated armors, the faces (usually all sorts of screwed up), and the composition. Its just a way to kickstart the creation process.
People so quickly assume that access to these tools will make everyone an artist, but the raw output is so lacking in a voice and intentionality. If you supply the voice and intentionality through your iterative process and a hybrid visual/text language playing the generator like a violin… you're playing the generator like a violin.
Your artistic skills have been translated to a wholly new set of vocabularies, and it's your eye that is tested most. Can you see/imagine better than the next guy?
Imagine pairing this with bespoke automated clothing output - take a photo of an outfit, verbally describe the changes you want to it ("a little more debonair, dark lapels, 1920's styling"), click a button, preview it as it would look on you in several recent pictures you took, and a week later your tailored suit arrives. Now for sneakers. Hats. Bags. Watches.
The 20th century was about mass production to ensure everyone could have things: food, clothing, transport, entertainment. The 21st century may turn to expression: allowing each person to express themselves however they want in their goods and services. Or just following along to buy whatever your favorite tastemakers recommend. However involved you want to be!
The world is about to get a lot weirder and more interesting.
I would agree that this is like 20th century mass production. To be clear, I don't necessarily think that mass production is a good thing either. In fact, it has been probably the most detrimental thing to our environment that humans have ever done.
Mobile, Indy, and Retro are all very popular, just look at what people are playing on Twitch.
And maybe the number of places paying for art will grow by a factor of 8.
The flaws in AI generated art were 100x as obvious in systems like this only a few years ago. In 5 years I doubt anyone will be able to tell the difference between AI and human art.
When the photograph displaced most portrait painters, we invented a new type of artist - the photographer. I hope we’ll see the same thing here - artists who specialise in using stable diffusion (and friends) to make new art in a new way. This blog post is like one of the world’s first photographers saying “hey look how the photo changes when I move the subject relative to a light source!”. I can’t wait to see what results we get with deep expertise (and better algorithms).
How long before we have filmmakers using AI to cast, direct and shoot their films?
The mulchers, as Bruce Sterling calls them, have a fresh meat problem. They've consumed all the words, and all the pictures, and we already know that feeding them their own mulch gives worse results.
We're not at the scale limit for data but we know where it is. It's not clear that refinements to the mulching process will create mulch good enough to tell apart from creation. It might. But it might not.
I'm not so sure about this. While some scenes are still obviously using CGI, I think a lot of CGI in movies now passes unnoticed, even entirely digital characters.
We certainly notice when those characters do things humans can't do, of course, and when budget or schedule or both result in things being pushed out too early, but how would we know when digital characters look natural on-screen? We wouldn't!
You generate the image in 30 minutes (maybe less if you get the process down to a science), then wait around for a few weeks to keep up the illusion that you're actually doing the drawing by hand, and send it off to your satisfied client. You could be charging hundreds of dollars for your "artistic services," and have dozens of clients going on simultaneously.
That makes the process a little less simple, but still easier than doing the real work
Manual labor still exists, but it’s a vastly smaller percentage of overall jobs. Productivity and automation seem like the same thing on the surface. However, the argument for a long tail of creators in an ever more wealth society breaks down when AI can start writing niche romance novels not just barely coherent news articles etc.
In theory we might have ever more new types of jobs, but automation isn’t just getting better it’s also getting faster.
What about women? Yeah, yeah, I know it's uncouth to worry about gendered language. On to the real topic ...
> reducing [a job] to a kitschy "artisan" label
Automation doesn't (only) reduce jobs, it (also) allows people to get meta. They are now free to think about how the job gets done, rather than constrained by time to only do the job the way they did it yesterday. Or, they can do other jobs that they prefer.
I used to live in an apartment without a washing machine, and where the closest laundromat was a 20-minute walk away. I got used to a bucket-and-plunger method for washing my clothes. It was enjoyable in that "I am the salt of the earth" way. When I found a used sink-attachable tiny washing machine on Craigslist, I had a smile on my face the rest of the day.
It's just the speed and the scale that have been following logistic distribution. We're still before the midway point as a global society, but we can certainly see the big speedup as more and more work becomes automated.
If people's work can be automated away as a whole and people somehow become poorer rather than richer, then the benefits of automation are sucked away from them. And here comes the flaws you mention - or the flaw, I think, singular. They happen exactly in this one point.
Many aspects of computer setting of works for printing also show a significant regression in potential quality (prose, musical engraving, &c.), though the right software and proper human tweaking can balance that out—but people seldom actually do as much tweaking as would be done automatically by the experts of old. And as for text presentation on screen… well, that’s just lousy compared to what an expert setter would do. But it does adapt to different media with no or minimal effort required, doing in one second what used to take days, and that’s a rather big deal.
“Unfortunately, the venture was so successful that Magrathea soon became the richest planet of all time and the rest of the Galaxy was reduced to abject poverty. The Magratheans went into hibernation, awaiting an economic recovery that could afford their services once more.“
If he was right, it can be "patched" and made to work with a sufficient level of redistribution to avoid such crises, or left to fail catastrophically without.
The fundamental question is: Why have owners?
It sounds like you might be so squarely in one of those camps, that you see it as a fundamental flaw that people may disagree with you. To me, that sounds even worse!
It doesn't mean that we shouldn't make things more efficient, but economies exist to serve people, not the other way around. That sometimes means bailing some people out, sometimes it means phasing something in, and sometimes it means changing the definition of value.
For example, painting used to be how you'd get a portrait. But photos did that much better. Painting shifted: it became much more abstract, and because photos were cheap and easy, they didn't have the same cachet as a portrait. Not many people hang big photo portraits on their walls they way they might have done with paintings.
I suspect authenticity, or some other thing which the machine cannot replicate, may be valued more in future.
If the highly ethical are avoiding involvement the advances will be left to the greedy and oppressive.
Further, when you automate one portion of the work you still need the human brain to strategize and orchestrate at a higher level. Job opportunities have only increased as a result of this, not decreased.
However, much more positions in machine learning and data science.
Software _is_ eating the world. The only viable survival strategy is learning to code. I don't believe that not everyone can learn to code. I teach people to code, and I have yet to meet one that couldn't learn, assuming some general intelligence (about the same that you need to learn a foreign language).
I haven't gotten around to faffing about with Python and Tensorflow or whatever yet.
As software eats the world, the value of people who can talk to computers will increase. Even if it takes fewer programmers to make a website, there will be more jobs for programmers to automate concept art production pipelines. You may not be writing javascript or python in fifteen years (I bet you will) but there will be code to tell the automation services what to do.
The interesting question is the general population becoming more tech savvy? Will this change in work encourage more students to learn how to code (whatever that looks like in twenty years)? Or will the demand for coders rise without a corresponding increase in supply?
Technology has always enabled the creation of jobs faster than it displaces workers though. Sure, horse manure shovelers lamented the automobile, but people who became mechanics and petrol pump attendants didn't. The same will likely be true for artists - this will suck for them, but the proliferation of easily generated art assets will likely enable the creation of entirely new jobs we haven't considered yet.
Yes, and isn't it wonderful?
Ideally machines will do all of our work as soon as possible.
The ratio of people the printing press helped to those it hurt approaches 1.
As Andy points out in this article, the model itself is a 4.2GB file. That's way, way too small to work as a "database" of examples it can stitch together.
I think of it instead as an enormous mass of loosley assembled impressions of concepts - everything from a low-level primitive like a triangle to a Star Destroyer. You can use text prompts to combine those primitives - so you could get it to generate something just from dots and lines and shapes, or you could mix in extremely detailed concepts like the Seattle skyline - or anything in between.
https://huggingface.co/spaces/huggingface/diffuse-the-rest
To understand how stable diffusion works, see:
https://www.paepper.com/blog/posts/how-and-why-stable-diffus...