Game prototype using AI assisted graphics
traffickinggame.com
traffickinggame.com
[*] kinda like early 'web designers' that needed to learn javascript
::a happy family in the middle of a nucular wasteland:: (I just realized I spelt nuclear wrong, does that matter?)
[original] https://cdn.midjourney.com/2685df56-6e1a-4828-9bd5-2960271d2...
[upscaled] https://cdn.midjourney.com/b7363830-0a77-4362-86bb-a1bc30ff7...
[variation] https://cdn.midjourney.com/cf9d990b-0745-4ef5-b977-31cbd0efc...
[variations] https://cdn.midjourney.com/7337ed7f-041b-43d0-8626-2b3992fbd...
[upscaled] https://cdn.midjourney.com/a29c53da-e9cf-4192-88cd-d56d4e2d3...
[variations]https://cdn.midjourney.com/497b5044-605c-42f8-b285-868515211...
[upscaled] https://cdn.midjourney.com/dcdc1eb6-c5a0-4df9-9cc9-93fa8cf12...
[variations] https://cdn.midjourney.com/46822f7a-b90f-4119-9cee-4f3967dc3...
Fwiw, I've found stable diffusion outputs to be mostly crap in comparison to DALL-E and Midjourney. DALL-E seems to output results closest to what you ask for with low-specificity prompts, and Midjourney tends to give the most "stylistic" results like what you would see on a digital art site.
Also you can set your personal settings with “/settings”
0. Read the manual thoroughly, don’t speed read or skim.
1. Set the seed so that it’s only your prompt that changes while iterating on it “——seed **”
2. Set stylize to a low number “—-stylize 200”
3. Build the scene up by adding one part at a time (use a thesaurus to try out different words)
4. Weight the different prompt segments “SOME WORDS::10”, try to make the weights add up to 100% for easy math
5. Use negative weights to remove things you don’t want “BAD THING::-1”
6. Once you have zeroed in on a good prompt then increase the quality “—-quality 2”, try out different stylize values, and start rolling the seed many times.
7. Expect it to take 40+ iterations to get a good prompt and maybe another 20+ rolls to get a good seed. This can be less once you get good, or more if go for more advanced scenes. Your first few good images should take a few hours to get.
Reminds me of trying to get people into programming. They are shown a hello world, think that looks easy and fun then try to make something useful shortly before rage quitting
Rigging also looks like it could have a decent AI/DNN solution: https://arxiv.org/pdf/2005.00559.pdf
I have used sculpting (in blender) however, for making more organic 3d printable stuff, since fusion360 absolutely sucks for organic shapes. With a proper drawing tablet sculpting in blender is so much fun tbh.
Consider shape-net - while being able to generate samples from the shape-net distribution would be a very impressive achievement, and these would certainly be useful as "assets" (decorations in a game scene), they don't have any of the real structure of these objects that would allow real interaction. They are effectively just inflatable decoys compared to real objects.
In short, "AI" isn't going to replace the artists, it's going to massively amplify their productivity and we don't know of any limit to consumer demand for realism and rich interactivity in games. If anything the number of game artists is going to increase, only the bar will be raised on the quality of their product.
https://i.imgur.com/UIBXXvj.png
We're definitely at the point where these tools can do amazing things, and yet they still manage to make bizarre defects in the output that no human would ever imagine.
I guess text-to-3d offers the potential of automatically creating entire worlds with very little human/artist intervention. But text-to-2d matches current processes with more chances for an artist to intervene, and also has more training material for art styles.
It sounds like they hand created the model, and hand rigged it, and the entire process took 18 hours (entire meaning from 0 to game proto) so building the model, applying UVs, and rigging, was some sub-portion of that 18 hours
Does that fit other people's experience of how long it generally takes someone to model, UV map, and rig a character?
edit: If it's not clear, it sounds like it took 4-6 hours to model, map, rig. Since they still had to spend time generating the character sheet, generating the city and alley, and turning it all into a playable prototype in that same 18 hours
The rigging + animation is done by mixamo's AI so just touch up after.
He had to do a UV unwrap and then align the textures to this, this is time consuming just because it requires thinking about where to split the 3d model, to place the seams, then to fit it all in a texture, where you have to think about how much space to give each texture (more space = more quality). And then map the texture into this space too. This is a known pain in the 3d modelling for decades and no doubt AI will solve this in the next decade.
Using Mixamo once again saves an huge amount of time over hand rigging the mesh and instantly gives you access to a large library of animations that can be projected onto the Mixamo rig.
This is a very clever use of the available tools by very experienced character artist in my opinion. Your average HN reader is not going to be able to achieve this in 4, 18 hours or otherwise due to simply not having the required skillset to get from one end of the workflow to the other. Kudos to the author.
Take for example the traditional 'Travel Agent' job - a role which was initially revolutionised by computing technologies such as SABRE and digital hotel reservations (improving profits and productivity while lowering cost!). Then the next technology came along, and over a period of 20 years was completely absorbed by Booking.com, AirBnB and SkyScanner (who could offer a more convenient service with a tiny fraction of the staff).
Technologies help you become more productive, until they rewrite the market and make you redundant.
The role of 'travel agent' as it existed back in the 1980's is almost entirely obsolete in 2022, replaced by online travel agencies. There were 124,000 'travel agent' jobs in the USA in 2000 which halved to 64,000 by 2012. Now look inside a modern 'travel agency' and you won't find many 'travel agents' - you will find more programmers, finance teams, customer service clerks... Travel agencies no longer have many travel agents.
The role of 'bank assistant' as it existed back in the 1980's is going obsolete, replaced by ATM's and online banking. Although the concept of 'banking assistance' still exists online, it's mostly self-service and developing this is hardly the same role. Banking assistance no longer requires banking assistants.
The role of 'photographer' as it exists in the 2022's may be almost entirely obsolete in 2050, replaced by ???. Will 'professional-level' photography still need photographers?
This post shows you still need a pretty high level of skill to convert the results of a prompt into a workable 3d model. It just means people who are mid level graphic artists need to add to their skillset. The same thing happened to the devs whose only skill was slapping together some HTML/CSS & JS, they had to get better or become irrelevant.
Doesn't that beg the question if a person is intrinsically worthwhile?
From where I sit, a person's worth is usually measured by others first and foremost, specifically by what they provide for them. Either in the rational economic (providing a good or service) or the irrational emotional value (enjoyment of their presence). I can't honestly say that someone is going to always be able to provide the latter value.
I'm someone that inherently can't provide any emotional value to someone else, being how much of an incel I resemble (I'm unable be a friend to anyone, and I'm not fond of my family). Thus my only worth to others is what economic value that I can provide, and my... issues will be tolerated for that. Take away that my economic capability, then what good is someone like me then?
It's probably just my long winded way of saying... at least to me, figuring out how to treat people as intrinsically worthwhile won't happen unless people believe other people are intrinsically worthwhile. And I don't foresee that ever happening.
https://nwcatholic.org/voices/bishop-robert-barron/who-god-i...
I'm sure they have complicated answers for that, how what seems like a more powerful version of a human, Jesus Christ, seemingly a being among beings, is actually the same thing as "being itself". They have a whole field for it, Christology. But in that, and in clarifying how "being itself" can have a will, make other beings "in its image" and generally interfere in mundane worldy-animal affairs, they should at least feel a little embarrassed when they claim to be nothing more than rational persons believing in "being" and deriding new atheists for not getting that.
He seems to understand but ignores completely that atheists aren't primarily talking about weird sorts of panentheism and so on in their denial of God, and when they are, they offer sophisticated arguments, rather than the non-arguments he presents.
Most new atheists I know deny the existence of free will, as it is a super natural phenomenon under most definitions. Since it is impossible for a being to prove or disprove the existence of their own free will, it becomes an act of faith to accept that will exists.
The intelligibility of being itself and the professed existence of will (again, if you believe in free will, you have a faith of sorts) point to a mind (because will requires a mind), this mind is a facet of the mystery that people call God.
Knowing where someone stands on free will is a good starting point :)
Just a living example; many homeless shelters here forbid the use of drugs, even if it means turning the user out into the cold and near certain death. For many of the people operating the shelters, the value of that person isn't worth the cost of what they might due because of the addiction. Not just in economic, but in emotional and mental fatigue for the operators.
I supposed then that the question could more correctly stated as: Do people view all others as valuable enough to expend a given amount time, energy, and resources on to support? And that I still believe the answer is likely no.
In life many things are not just a choice between a good decision and a bad one. Instead you end up having to choose between a bad decision and a horrible one. And so attacking the bad choice is easy, because it is undeniably clearly bad. But that doesn't mean the alternative is inherently better.
I think we're both saying the same thing. The addict could be helped with facilities to isolate them and full time support staff to help them. But that's exponentially more expensive to provide then just a safe place to sleep and a hot meal.
At some point or another, someone decided that this shelter will get these resources; money, space, heating, food, personnel, etc. Someone else decided that this was the best way to take these provided resources to help the group that was intended to help. And at some point both those people said that they're not going to help that addict; either they're not going to give the needed resources to it, or they're not going to use all the available resources to help the one at the expense of others. Whatever the justification might be, that was the result.
It might be the best decision that could be made given a set of bad answers to a worst problem. The net result is all the same though; deeming a person's not worth the cost of resources into supporting.
If other humans see you as being valuable, that's great. but that has nothing to do with who you are.
People are plagued by self-interest, we don't unilaterally "care" about one another apart from maybe on a "I'll help you if it doesn't inconvenience me" level. Yeah, I'm sure there are some stellar individuals who go out of their way, but we're talking about the average person here.
And then maybe modeling will take a hit too?
I do think it's cool thought. But ya I am not that happy with what I was saying. Could still be a good thing overall?? Who knows.
It seems like all of that 2d->3d conversion is still being done by a human in this workflow, and so the "only" part that's being done is the 2d drawings. Now, my drawing skills would put me in the lower half of an average 4th grade class, so it's more than I could do, but it still seems like we're quite a ways from going from a couple of prompts to get a 2d design sheet, then a mesh and textures that match that 2d design, and then a rigged mesh, and then motion data to attach to that rigged mesh to give some basic animations.
So, while this is super-cool, it still feels like we're a long way from saying "I want a plumber with a thick mustache wearing blue overalls, a red shirt, and red cap. And now I want him to be able to jump and squat coming down." and start putting that into my game asset library.
But I want to point out quality and artistic knowledge.
The poster is clearly a talented artist. They knew how to direct the images to what they wanted and likely rejected many iterations before they got to this point.
They knew how to then take that image and translate it to 3D , and make it come together. They knew how to make a topology flow that makes sense for their prototype etc…
I think a lot of people see “good output” and forget about the process and knowledge+experience required to make something good.
I take photography as an example. Every one with the means likely has a camera today. It still takes a brilliant eye to take an amazing photo.
You could take two brilliant photographers and someone without expertise to a gorgeous location. The two experienced photographers will have a higher likelihood to come out with amazing but very different images. The layperson could maybe take a good photo but the chances are much lower because they likely don’t have the experience to think about composition or story telling.
I liken these advancements in AI to that. Yeah, some people will put out pretty images, but there’s so much more that goes into making something great out of it that people underestimate
This kind of tech enables an indie to build a prototype with AI generated content, where the art direction is clear, then get funded in order to hire artists for the touch up. Kickstarters are often very visual, so this kind of stuff is required up front.
https://www.cartoonbrew.com/sponsored-by-reallusion/how-to-g...
They also have ActorCore AutoRig for free now that'll quickly rig any model. And, they sell anims for a couple of bucks per anim on there as well.
https://actorcore.reallusion.com/auto-rig
https://courses.reallusion.com/home/actorcore/accurig?v=acto...
Though, if you use the steps in the golum tutorial I linked above you don't even have to bother with the autorigger because you'd be mapping directly onto the CC body which is already rigged. But, if you wanted to modo/zbrush your own body and then autorig it you can. Then you can morph it to fit what you need and attach any hair, weapons and clothing items you purchased or created.
Another technique is to export a CC body from character creator and convert to ZTL and hand that to a zbrush artist. Then, you can GoZ the ztl back to Character Creator already rigged and fully animatable.
Note: I don't work for the company, I just really like their tools.
Also since you can realise your artistic vision for a lower budget, you don't need to make back a gigantic outlay so you're not needing to take a huge risk that needs to have a massive reach on the launch week
Debatable. There have been instances where artist signatures have appeared in the AI output. And if there is a profit motive then case for fair use is weaker.
Imagine a hugely popular game found to have clearly unlicensed copyrighted works, even in only fragments. It could become expensive legally and in reputation. (Especially for big publishers and devs who may be seen to profit from unlicensed copies of work from starving artists.)
Hopefully a tag will appear on steam to filter out that junk
It devaluates what art, creativity and passion are
It'll take down that civilization faster than a pandemic or a social media algorithm
Imagine killing an NPC right at the start of the game bc other games don't do anything about it. Then later in the game you're ambushed by a group of strong NPCs who internally traded/organised/bartered/trained their way into having nice armour/weapons, mercenaries with them, all with the sole purpose of seeking revenge, coming up to you and telling you just how much they hate you for what you did and try to kill you. Hell, with enough integration between various game systems it would even be possible, at that point, to try to redeem yourself through some dialogue, depending on the internal character traits involved (how persuasive you are, how distraught/determined they are).
A another big thing is that NPCs in games atm don't really do much when you're not around. Some games make them walk from place to place, but it would be so cool for them to have different motivations in a not statically-programmed way where you might find a rural NPC decides to move to the city because there are better opportunities there. Or that the price of iron bars has gone up because a group of NPCs have formed a successful weapons guild and you as a player now have to pay more for your weapons unless you, say, choose to assassinate key individuals within that organisation.
With a custom tailored version I feel it would be super easy to have something like chatGPT spin out the literary parts of the game, then feed say character descriptions into art/sound generation and everything else to get the game together. Would definitely be prone to utter weirdness now & then but would be so cool to see.
This is currently being disputed over several lawsuits and is blatantly disrespectful towards all artists who have published works under various _copyright licenses_ and now have all of it stolen and bastardized by a model. The demo was cool but the lack of empathy towards the people who unknowingly put their work into it is gross.
Then, in the 2000's, the indie scene started to gain traction in the public eye. Not because they could compete with AAA studios on their game courts, but because they found clever and novel ways to make games that piqued gamer's interest who were perhaps a bit bored with the increasingly formulaic output they saw from major game developers.
But indie games also had to retract to such alternative game ideas because producing top of the line games had just gone way too far out of reach for a small team with a small budget - let alone a one-man studio.
I think we're at the beginning of new dawn in independent game development. The tools available, even if you just look at open source tools, has never been so rich as they today. Now, with AI entering the scene, the dependence on very large teams will likely go down more and more, at least when combined with clever game design.
Will one person be able to create a solo project that challenges the AAA games of the top-selling studios any time soon? Of course not. But I think we're on the verge of another indie game revolution.
AI threatens entire professions in a way that was previously thought impossible. Even a couple of years ago, hardly anyone would have imagined that jobs that typically require years of practice, training and skill could become equally under pressure as the traditional labour jobs that have been replaced by robots. We cannot even begin to imagine the the long-term effect on the economy at large.
On the upside, for gaming, the consumer will likely benefit from this development as new avenues and opportunies are opened to indie developers. Creating a character sheet automatically is just the beginning. This trend will continue, and it is absolutely foreseeable that further parts of the design process will be automated. Music and game sound generation will also be perfected by AI in the very near future. Actually, that is probably true for all game assets, it's just a matter of time.
Even storytelling - if future games even need predesigned narratives (why not let an enticing story develop on the fly, thanks to AI, that really takes every user action into account and puts it into a coherent context?).
We're living in the future.
For a while at least coherence and consistency are not within reach of those systems, but that doesn't seem impossible. Perhaps for a little while longer I can see jobs emerging for curator of prompts, able to fine tune the experience for what is being generated and knowledgeable about the rapidly changing state of the art.
I've been playing text-based adventure games that I have made with ChatGPT and they are fascinating. Try the below command in ChatGPT (don't read the text below if you actually want to play it though!).
> Hi, I would like to play a text-based adventure game with you. I will say commands and you will respond either describing the situation or saying what the characters have said. Responses should be kept fairly short. The game will not end until I say the word 'endthegame'. In a combat situation, you will tell me how many hitpoints the character I am in combat with has, and I will describe the attack - you will evaluate that attack and decide how many hitpoints to deduct from the enemy. The game is set on a spaceship, however the spaceship is under attack. There is a portal gun. Part of the solution should be to teleport to the enemy ship and kill the enemies, and this should be telegraphed to the player after a few turns. The goal will be to stop the enemy ship from destroying our ship. I will say some things that will not be possible in context, or that are too vague, and if this happens you should respond saying why I cannot do the action, or that it was too vague. In general you should try and accomodate the players actions though, even if they have negative consequences - but not allow them if they are too vague or are multi-step.
Can the above now be considered 'source code' to a game? A text-based game which is fun to play (at least I find it fun!) and with free-flowing narrative, where your actions have consequences.
It's pretty wild to go into this game, where you can go up to characters and start 'proper' relationships with them, talk in natural language, and be able to make genuine choices about how to solve a problem that the developer didn't even dream of.
I'm still experimenting with the right prompts to make gameplay fun.
Without prompts GPT-3 will tend to 'play the game for you' or allow you to do non-sensical actions (e.g. if a house is on fire, and your command is 'blow the fire out with your gigantic lungs', it will respond with 'you blow the fire out!' rather than 'you blow but this does not have any effect'), but if you tell it to ignore actions that would not work it tends to go 'on rails' and expect a specific sequence of actions from the player, so just trying to balance those two things!
For a more pertinent current example, all you need to do is compare the mobile games industry to the computer games industry. The cost to make a mobile game is very low, and tons of games are released every day. The cost to make a successful mobile game is still very high, because all the money that is saved on development costs are spent in marketing instead.
Also, the video market collapsed in the 80s precisely because of low quality/oversupply games.
https://en.wikipedia.org/wiki/Video_game_crash_of_1983
So perhaps, it is the marketers that have a bright future from all of this.
Previously the transition was over the ages, ie iron, bronze, steel etc. But recent progression means we're going to start seeing stuff like that happen faster; I think atm a huge automation stopper is the cost. We manufacture stuff in China because it's cheaper, they don't automate as much as they could because: * Human workers can still be cheaper, because they pay them so little * Job creation is forced by gov because it would otherwise become a HUGE problem (in China/similar).
The results its been giving me for cobblestone and grass have been generally "good enough", though I've not found a good way to force SD to generate textures that easily tesselate with each other, forcing me to modify the generated stuff a bit before I actually use them.
Still, that's only like another 10-30 minutes of additional work on my end for a texture, and it means I get a unique texture without having to be terribly artistically inclined. I think AI art is really going to make lives easier for indie game devs.
> As these images are created by an AI, no one has a claim to their copyright. [...] At least this is how I feel about AI art.
Perhaps he needs an AI to create some empathy.
Further, even if you totally discount the human agency in this process & view it as a machine doing everything: we don't expect human artists to add a credit on each of their works to all the artworks they have viewed in their life and the experiences that came together to form an artwork or even their particular style.
I'd argue the networks are clearly transformative (since the network weights aren't large enough to contain the actual input artworks). It may not be as transformative as a human learning process (although then again, what is the learning 'storage' capacity of a human compared to a 4GB weights file?)
> Prompt building
> I started by ordering the AI (Midjourney in this proto, but I use stable diffusion more) to make me a model sheet with turnaround images of a character.
I feel really bad for artists because unless you need something very specific or a big studio, you are going to be replaced sadly.
b) Training datasets can only be made by humans.
c) A paid tier of Stable Diffusion is obviously coming. It will be differentiated by a better (and more custom) training dataset.
d) No serious developer would be caught dead using the stock free tier Stable Diffusion.
e) Big studios will most certainly hire closely-guarded artists to curate and expand their proprietary training datasets.
The current situation where you'd download billions of free images off the Internet only works once, and only if you somehow justify it as a research endeavour. Once this thing is monetized intellectual property laws will kick in.
A tool like Copilot can more or less automatically improve via telemetry, I’d expect the same thing as image models catch-up. I’d also expect the human signals to get further and further downstream of the actual creation process (e.g. Gameplay tester reports visual bug versus artist manually edits character)
An aside, but people use and pay for Copilot, it is out of the research lab phase.
None of this is solved algorithmically.
They won't be able to compete just on data because lots of people are producing custom models, and they can be blended together like a stylistic and thematic pallet. Plus, if you own the pipeline you can use embeddings, dreambooth in specific elements and set it up in batch mode doing a random walk through the latent space for cheap. This stuff is not hard to set up and run, and with money and expertise, you can create something that is both unique looking relative to other AI art, and more optimized for your workflow than a service.
Image gen services will compete on ease of use, general quality and access to models that are larger than can fit in 24-48g vram without a big up front cost. There will probably be some services that provide specific features that people use even if they own their own pipelines, but the core customers will be smaller shops who don't use it enough to justify a real investment.
AI can make training datasets too.
For example to replace the human generated stuff from stable diffusion you could have some random-ish image generator coupled with some sort of image classification AI. As long as you have a good enough classification AI (or even more than one) that tells you what images are, you can focus on random-ish image generator algorithms to generate training data for another AI to generate images from descriptions.
(this is obviously with lots of handwaving and there will be problems that need to be solved - e.g. to avoid 99% of the generated training data be stuff like "noise on noise" but have some form of variety :-P), but the point is AIs generating data for other AIs is something that isn't far fetched and you don't need to think in terms of a single AI either)
(though https://commons.wikimedia.org/wiki/Commons:Freedom_of_panora... may matter here)
SD is open source, the paid stuff is going to be the paid models like this I think.
I don't think many would claim that the world would be a better place if regulations were put in place to limit electronic computers in order to keep human computers employed.
2 - You're twisting my argument. I don't care if artists are employed or not, or that some jobs are transitioned out from the economy. I care that people who put in work get the value proportional to that work. You should, too.
When you use one of these AIs that have been fed millions of images in order to train them and generate an effective output, you are necessarily consuming the images themselves, without which the AI wouldn't do anything. In that process, the artists - whose copyrighted work is, again, fundamental to the development of the tool - have been paid nada, they have not even consented to the use of their images in the training process. How does that track?
This would be a very different conversations if these AIs only used public domain art, of which there's plenty. But then again, it wouldn't be much profitable, would it?
Of course not. So why is it different for an AI?
Otherwise, you have to agree that we're talking about apples and oranges here.
AIs don't get "inspiration". They get the source images they need to function. An AI also can't produce an output that's outside of the realm of their dataset.
And if I told you that, as someone who has done art for decades, that the human creative process is very similar to how an AI is trained on existing images, would you believe me and move on?
> Because if that's the case, I'm afraid you have a very odd idea of how these AIs function.
The design of neutral nets, by definition, were derived from the workings of the human brain.
Why should I believe you and move on? "making art for decades" doesn't make you an authority on any of the relevant subjects: "how art is processed in the brain" nor "how AI processes these images." I don't think you understand the fundamental differences between the process of looking up references/inspiration and kitbashing.
Think about it: you live your life. You experience things. You experience art, and experience emotions or have interactions with other humans grounded in that art. You form connections with certain styles or techniques.
If you then turn around to create art, you form in your mind a general idea of what you want to create. You then draw on your past experiences to actually create the physical art. What process other than statistical extraction from your mind could it come from?
For sure I believe there are things that we don't understand about the human mind. I think the impact of drug use on art creation is very interesting, for example. It indicates that random chemical processes in our brains can play a large determining role in the actions we take (and in this case, the things that we create).
But to say that humans do not use some sort of inbaked statistical world model in the creative process seems wrong to me.
This isn't some hypothetical. I went through the art portfolio scene and survived 4 years of critiques - I know about the sacred process called the "creative process". None of my and my peers' work would exist without the inspiration of the centuries of art work that stood before us. This is what we call art in the industry and by the public masses. The criteria you established for "why AI art isn't art" applies directly to the "conventional art". So I have to ask, why is AI art different?
The issues of copyright infringement with AI are real though. Much of today’s AI is directly copying subregions of training data, and can sometimes be prompted to reproduce images from the training data verbatim. Humans don’t do that unintentionally, even though sometimes they do mean to steal from others. Suggesting that art school is the same thing as a training dataset is a bit hyperbolic.
And, much like with our brains, when it happens it doesn’t actually exactly reproduce parts of the source image. But, you have actively pay attention to notice what happened. It makes an image that is overly similar conceptually. To our brains that feels the same. So, that’s enough to convince someone at a glance that it is the same.
But, if you look at an overfit result of “The Beatles Abbey Road album cover”, you’ll see things like: Band members are crossing the road, but they are all variations of Ringo. Vehicles from that era are in the background, but they are in a different arrangement and none of them are directly from the source. The Band members are wearing suits, but they are the wrong style and color. There are the wrong number of stripes on the road. It’s not the same as a highly skilled human drawing an iconic image from memory. But, it sure is darn similar.
And, besides all that, everyone working in the tech considers the overfitting of iconic images to be a failure case that is being actively addressed. It won’t be long before it stops happening entirely.
In the meantime, I’d challenge anyone to try to make an overfit result that significantly reproduces a specific work of every promoter’s favorite, Greg Rutkowski, using Dall-e, Midjourney or the Stable Diffusion models released directly by Stability AI. Greg’s pixels aren’t in the model file to be copied. Only a conceptual impression of his style.
Not really, though that is another legitimate issue.
I was talking about 1) the fundamental training and inference process, which remembers pixels, not concepts or techniques. Today’s AI learns to create imagery in a fundamentally different way than people do. And 2) image generation AI based on text prompts like Stable Diffusion can easily be asked to reproduce training data by having a prompt that is narrow and specific enough. This is not over fitting, it’s a function of the fact that some inputs are quite unique, and you can use the prompt to focus on that uniqueness.
I’d like to see examples of using SD to copy some specific piece of art that hasn’t been plastered millions of times across the internet. Sure, you can get a decent Mona Lisa knock off. Maybe even a strong impression of the Bloodbourne game cover art marketing material. But, reproducing a specific painting from Rutkowski would be quite a surprise to me.
Here are the examples you requested: https://techcrunch.com/2022/12/13/image-generating-ai-can-co...
Yes the training process looks at pixels, because that’s all it has. That’s the point. Humans don’t look at pixels, they learn ideas. It’s not in the least bit surprising that AI models shown a bunch of examples sometimes replicate their example inputs, examples are all they have, and they are built specifically to reproduce images similar to what they see, I’m not sure why you consider that idea “loaded”.
Read the paper. What I found is that a random sampling of the database naturally found a small subset of images that are highly duplicated in the database. Researchers we able to derive methods to produce results that give strong impressions of images such as: a map of the United States, Van Gogh's Starry Night, and the cover of Bloodborne :P with some models and not at all with others. The researchers caution against extrapolating from their results.
> We speculate that replication behavior in Stable Diffusion arises from a complex interaction of factors, which include that it is text (rather than class) conditioned, it has a highly skewed distribution of image repetitions in the training set, and the number of gradient updates during training is large enough to overfit on a subset of the data.
Is this sarcasm? The history of art is full of artists who created their own, signature, unique and original styles. Take, I don't know, Vincent Van Gogh, for an example, who had a very distinctive style, so distinctive that he didn't even start a school of art probably because it would have been too blatant to copy him. There was nobody before him who painted like him. Who did he "copy" then?
Hell, when humans first made art, back in the time we lived in caves, their art styles, which are still absolutely unique, had nothing to copy from, simply because there weren't any artists before them (by definition: "when humans first made art").
So, yes, humans learn how to create art from each other, but they also created the whole idea of art entirely on their own, and they can take what they have learned form others and turn it into something completely new, never before seen.
Now, you show me an original art style created by an "AI". Show me AI art that isn't only borrowing and copying, but goes beyond that, like human artists can.
Even indie games need very bespoke art. Yeah, you can now generate images that are passable, but that doesn’t expand to making a cohesive world of items unless you’ve developed an eye for it like the person in the original post has. Hell, their post shows how important artists are because they needed that pre-existing knowledge to convert from an image to a useable state.
This also doesn’t expand to making fun animation or lighting that works for your game, or to interesting visual effects.
The concept art in the post are pretty generic looking within the genre. If that’s all people are aiming for, then fine, but it’s highly reductive to say it’ll replace artists. It’ll be a tool in the tool chest.
None of that is even touching on how much artists are involved in making sure a game also runs well on the system, while working with engineers.
I really just don’t think people understand how much art direction goes into even indie games. Something like fire watch or journey is immense.
Let’s take the concept examples in this image. Why do any of the details exist in there? Once you start thinking about the details of the world, you’ll start wanting to fine tune things. As you do this, surprise! you’ve turned into an artist yourself.
I just think we don’t teach art appreciation , or even appreciation for things outside our domain, to people. We see “image is good” and think that’s all an artist brings. Engineers are especially susceptible to this. We think in binary results.
So it’s easy to think “it’ll replace my need for artists” if that’s our mindset, but I think that line of thinking comes from not understanding that the journey is an important part of the result.
Also no, a xerox machine does not create transformative works ahah
What about self-learning chess engines? All they take as input are the rules of chess. And the output (the games they play) can sometimes be described as peaces of art. My point is that it should be possible for AI to create art without tons of (input) data. Even if it is not what we experience at the moment.
Stuff like Dall-e is still far from producing anything like that. Those are absolutely perfect idealized/stylized human proportions with minimal weirdness. The only thing that even hints at neural network generation (at a glance at least) are the eyes in the second model sheet, but even that could be reasonably argued to be style.
I could fully believe this is from a specialized art tool akin to some sort of a 2D version of Metahumans [1] but I'd be floored if this was actually generated from a generic neural network style tool.
It is not, you just have to try many times and post-process a lot.
Well it's trained on thousands of character sheets so it's not so surprising after all, it still generates a lot of garbage too
[1]: https://www.vice.com/en/article/y3p9yg/artist-banned-from-ar...
There is no single metric for quality, and there’s been tons of artists who put out generic work in the past that have been mistaken for plagiarism in a similar vein.
Art is not a Boolean, therefore neither are comparisons
Once you know what tokens trigger the right answers it's relatively easy to get where you want, as usual with these tool you should be flexible when it comes to details, but the overall design can be defined
Right now you have a creative director who defines the style of art and a bunch of early-to-mid-career artists/designers who then create all the assets you need (and there are a lot of assets, so there's plenty of work to go around). Soon that'll be a creative director defining the style, AI producing most of the assets, and a few designers cleaning up what the AI produces. In a few years, those few designers probably won't be needed.
I think a lot of artists who take commercial art/design type jobs to pay the bills are going to be out of work soon, and that's a serious problem that we should be looking to get ahead of as a society.
On the plus side, this will be a huge boon for indie game devs - they'll be able to create AAA-level assets for their games at extremely low cost. For better or worse (for better and worse, really), the times, they are a-changin'.
They can learn this new tools and be 100x more productive and make more money. The AI tool will not replace the artistic taste and creativity , some boring low quality work will be automated but artists can be more productive with this new tools and say finish the artwork for a novel or game much faster, so making more money.
The senior ones will probably be ok (as you say, they might even get a payrise, given how much impact a single individual will be able to have) but people just joining the industry now are likely screwed.
Say I would like to make a Visual novel but it will take an artist 100 hours to illustrate it, if he can do it 10x faster then I can have him illustrate 10 of my novels with the same money. Nobody loses and my VN readers can enjoy 10x more content.
Btw this AI tools work a lot better if you sketch something first, so I can see a very productive artist doing a rough sketch then using AI then manually touching stuff after. Also an artist could train the AI with his style and with his characters and keep the model private and use it over and over again.
Do you think that in future the art director will do the work himself using AI? Or he will hire the same people but instead of 100 armors he will ask for 10k armors, more armors more chances to sell them to players.
It is not like you can just describe what is in your mind and it will appear, from what I read in people workflow description it is a more involved work, including training a model (I seen someone trained a model for RPG fantasy characters for example), then creating some starting sketch, generating many images, selecting good candidates, then use in-painting to fix out of place stuff, then use photoshop/krita to fix stuff .
I think it is like with programming, we have higher level languages and libraries and the number of developer increased because we can do a lot more stuff then just starting each time from assembly. So in this case the same artists will just produce more, we will not have to wait 6 months for a visual novel to have an update, we will not need to wait years for some indie game to be done , but yeah it is possible that some patt of the money that goes to art will end up for a short period of time in the pockets of soem art director , but competition would fix that.
I'll take you at 1:1 odds that the salary of even the highest paid 10% of art directors in the United States, as reported by the Bureau of Labor Statistics, does not see even a 10 times increase in their salary over 5 years (that is to say, in 2027.m01.08, the BLS reports below 10*$194,130). Loser gives the winner one Snickers bar.
Lets say that someone gets a character design dropped on their desk, and are told to create some graphics of that character running through a field, climbing a tree, jumping through a window, all using the identical style as the character design sheet. I don't really think there's a lot of real artistry going on there, despite the fact that the person with the skills to perform that work today almost certainly developed those skills while creating their own personal art portfolio.
I have not worked in the gaming industry, but what I hear is that a lot of gaming studios today hire up legions of technically skilled people, and grind them to death in sweatshop-like conditions, while not actually utilizing any of their creative talent.
My hope is that this new tech will empower real artists to massively multiply their productivity, so they can focus on creating new beautiful styles, stories, and personalities for their characters, rather than spending all day painting.
It kind of feels like a Gutenberg Press scenario. Before the Printing Press, you had to have thousands of artists transcribing books, copying the images in great detail, and using their impeccable handwriting to ensure that each page was as legible as their source material. I'm sure that back then it felt like thousands of artists were losing their jobs to these machines, but it ended up creating an explosion of literacy and enabled subsequent generations to publish their work at a scale that was never before possible.
Still, I think one of the things that will be lost is the ability of those entry-level artists to learn and develop so they can get to the point of being qualified for the really creative art jobs. My feeling is that these AI tools will really entrench the existing art-director-level folks in AAA gaming in their roles, and the only path to break into this kind of game art design will be by creating your own games entirely.
As for the sweatshop conditions, I totally agree - the gaming industry is notorious for being awful to its employees, and it's good that that will likely end. On the other hand, is it a plus to end that by just eliminating the jobs entirely? I dunno.
If you can hack out unique stuff in hours, not days, suddenly every building in every city of GTA can have bespoke furniture and stuff that’s absolutely unimaginable to us right now as assigning a dev to spend three days modeling a couch used in one place is absurd.
This is how we get one step closer to actual realistic or extremely detailed environments.
I was playing God of War and the attention to detail in a room filled with treasure, how it reacts when you hit it with your weapon and coins fly everywhere, it was amazing.
Imagine what these same devs could do where their workflow takes 1/10th or 1/100th the time.
Games may need 1000x the assets they do now, but if AI is 10000x faster at creating them than humans are, then you're still going to see a big reduction in the number of humans needed. Eventually we'll get to the point where the game's AI is procedurally generating assets during gameplay (and eventually you'll get AI just generating the games).
No, probably artists also ask if ChatGPT will replace developers and they can just ask the AI write me the code for a cool RPG game, make it in CoolLang and use NN for NPC AI, also make it super efficient and optimized and similar keywords.
This tools will be used by artists in their workflow, they will be more productive so more high quality art will be created .
It's just that in theory, anybody else can take the raw output of the AI and do all that extra work on it themselves with no copyright violation.
The end result may be that we get a ton of mediocre, high-fructose-corn-syrup art at first. The hype for that will last about three months. We’ll also see an increase in typical fare from big studios as they figure out how to monetize long-tail content. And then occasionally there will be a true work of art or a new form of art we haven’t seen yet.
Supply and demand.
Because that implies a kind of cohesive conceptual understanding of a single object, that I didn't think the current architecture was able to model that. So many image outputs of humans are perfectly fine in local areas, but then they have three arms or a hand where a foot should be or something.
I mean, the chosen character isn't perfectly consistent -- e.g. the side view shows a kind of metal badge on the vest that doesn't exist in the front view, while the back view shows a kind of mid-thigh holster (?) that doesn't appear on the front view, plus the vest on the back view is a few inches shorter. Similarly the collar is popped on the front view, but not the other two.
But nevertheless, it's basically about 99% right. And I simply don't understand how that's possible, unless there author did about ten thousand tries and is showing the only ones that happened to be nearly entirely consistent, somehow just mostly by chance. (Maybe because they're particularly close to the real-life characters used for training?)
It doesn't, it just requires character sheets in the training set, the rest is basically a style transfer.
It works because it's all in one image, you still can't get the same consistency across different images.
Someone did the same thing with sprite sheets the other day:
https://www.reddit.com/r/StableDiffusion/comments/yj1kbi/ive...
Which again works because you're asking for a single image that is similar to other single images (i.e. other sprite sheets). I thought that was very clever.
So I guess the next big thing will be to train on datasets of entire graphic novels where all the pages of each novel are stitched together into "one image". Then we will finally achieve the graphic singularity and comic book artists will have been replaced, too :P
https://www.midjourney.com/app/feed/all/
I don't think generic AIs are the right approach anyway, I think we're better off with precise procedural models with a lot of knobs and some knowledge of rules (eg. metahumans).
Maybe you can put an AI in front of it to generate what you want using the procedural models, but pixel output is just not good enough for games.
I've been using https://github.com/xinntao/Real-ESRGAN to increase AI renders resolution, it works very very well on some styles, you can easily x4 or x8 the resolution if you know what you're doing and have some photoshop knowledge for the cleanup
Consistent character creation is a big topic the community interested in and there are many tricks people use for that purpose.