AI and the Future of Pixel Art
pixelparmesan.com
pixelparmesan.com
- generate details and texture, which photobashing was already used for; but that mostly solve the licensing problem for it
- generate random inspiration boards, for which image search was used (again, mostly solve the licensing problem)
- generate derivative stuff, e.g. typical game portraits or props
In practice, only the third point is generally discussed, because it lowers tremendously the entry barrier to generate images for people without any skills. It's like if you could pick up screenshots from other content, clean them up, and you're free to use it.
Whereas it essentially does not work for:
- cartoony generation. It relies too much on line consistency, visual clarity and abstraction
- concept design --not the flashy 10 minutes speedpaint type, but where you have to combine ideas in meaningful ways. In particular hard-surface design which required good 3D thinking and consistency of the whole.
These fundamental flaws are omnipresent, but can be hidden by certain styles where things are implied by color blobs, hidden by stylized brushstrokes, or simply an overflow of details (something Midjourney is very good at).
All in all, it feels like AI is a danger for people at the bottom of the profession hierarchy, but will elevate people at the top, whose work cannot be replaced. In other words, people who are more akin to be considered "artisans" rather than artists, who will take a prompt and simply clean it.
In particular, drawing has something like the 20/80 rule, where all the creative input is in the first 20% and the rest is 'rendering', a very mechanical task which you can mostly do with your brain turned off. As Yumenoley put it, "it was a mistake to let the AI do the interesting part".
https://dreambooth.github.io shows a glimpse of the future. You’ll be able to upload a few drawings that you want to emulate (e.g. Mickey Mouse), and then you can give it a specific prompt (e.g. Mickey Mouse doing a handstand).
You’re probably right about the consistency of 3D concepts, though. On the other hand, I was going to say “If you need a specific table, AI might not be able to help” — but again, dreambooth shows that we might be able to upload a few photos of a certain table, and it’ll take care of the details.
Give it a few years. :)
I think you’re spot on that AI will be an incredible tool for artisans. I used it to make some video game music: https://soundcloud.com/theshawwn/sets/ai-generated-videogame... Even though I can’t play any instruments too well, I was able to craft each piece uniquely. (My favorite is “Crossing the Channel”, which has a strange rhythm because I’m pretty sure the AI made a mistake at the beginning, and then extrapolated the next “actually, this isn’t a mistake” song that it thought of, which turned out to sound cool. A bit like a guitarist doing improv.)
There's a role for AI in filling in gaps, but illustrators remain essential in creating the basic style to extrapolate from and even more so in professional quality work (And to some extent, other procedural generation techniques were able to fill gaps up before cutting-edge NNs - the stampede of wildebeest in the 1994 Lion King was procedurally generated from a handful of models for example)
Here are some potential use cases:
- for fun (giving yourself a makeover, inserting yourself into famous movies)
- cheaper way to get studio photos (wedding photos, professional headshots for actors/models)
- easy way to create marketing assets. Like if you own an etsy store and don't want to engage a marketing studio, instead just create a dreambooth model of your necklace or whatever and create high quality product photos
There are probably a bunch of other use cases. I think making this easy (no figuring out how to do a git pull or rent a gpu) plus the community sharing aspects will make this technology a lot more accessible to artists and general users, and then the users will be doing all kinds of cool things with it organically.
The question is whether it not rather puts the people in the middle under pressure, when it helps the people at the bottom to produce better quality.
I can observe such a shift in translations. Deepl is not perfect, but it allows me to improve my own English texts considerably. Even if I were to give my text to a professional to polish, she or he would have much less to do than before, when I was not assisted by Deepl.
This reminds of of the chess computers from 30 years ago. Experts were convinced computers would never beat GM humans, based on the state-of-the-art of the time. They never took into account all the future advancements in hardware and software.
The same thing happened for both Go and Starcraft in the next decades, the experts said computers couldn't replicate enough spatial feel, and then they did. And now it's happening for AI art. Enough computational power and a sufficiently well-trained neural network can indeed exceed anything a human can do.
AI has been roughly doubling in performance every year or two, for quite some time. We just never noticed when it went from 0.0001% to 0.0002% of human capability. This is the year that it doubles from 10% to 20% and everybody notices. And there's not a lot of doublings left until it shoots past 100%.
The super hard problem is driving a robotic body, vs rendering an animation of the above.
I am an artist who is in a place where she gets paid decently to draw whatever the heck she feels like, with little regard for commercial viability.
The jobs you're dismissing as mere "artisans" are the ones where I was able to be in a place where I could spend most of my waking hours honing my drawing skills and still pay my rent. Every time you draw a thing, you get a little better at drawing that thing, and a little better at drawing in general. This is how you master your craft. If you're part of a studio then even better: there's older artists above you, and tons of opportunities for them to critique your work and open your eyes to the major flaws in it you can't see yet. If AI art fills all those niches for expanding on someone else's prompts, budding pros will find it a lot harder to get to the point where this virtuous circle of being paid to practice their craft gets started.
> In particular, drawing has something like the 20/80 rule, where all the creative input is in the first 20% and the rest is 'rendering', a very mechanical task which you can mostly do with your brain turned off.
This really depends on your process. A big part of becoming a pro, in my experience, was finding ways to make the boring parts happen a lot faster, and giving myself more opportunities to do the fun parts at every stage of the piece. That said there's also a pleasure to be found in putting on some good music, turning your brain off, and rendering the heck out of something! It's that much-coveted "flow state" people love to talk about.
Pure economics makes whatever hopeful expectation of advantage created by the AI generated medium irrelevant for anyone with a livelihood in the field. Being the creative genius may allow you to retain your employment, but rest assured, your income growth will decline and the next generation coming after you will be making less. The value of graphics will decline and so the money earned by the people making it will also decline. You're making truly creative work? Great, a company can just take your rendering and use it as the basis for an AI generated image to remove any licensing requirement. Good luck trying to prove original ownership. And who needs creativity anyway? How often are graphics used in commercial/retail/media enterprises really breaking the mold? Whatever shortcomings mentioned are already being mitigated by add-on AI systems (Tencent's ARC for instance) so there should be an expectation of rapid improvement in the coming years.
People employed in the field will likely adapt away from non-lucrative image generating work eventually, but the transition will be pretty painful and the overall effect will likely constrain incomes for creative work in general.
With this tool suddenly the effort of making a series of compositions that are coherent will dramatically change the game, specially in the videogame industry. I think there are a lot of programmers out there capable of making great games but might be lacking the resources to fully complete their visions due to having to needing assets for their games.
It seems to me like the approach stable diffusion has taken has dramatically increased interest and utility of these tools. So I'm hoping they follow similar lines for other types of AI. Every week I'm reading of a new novel use for these generators that I hadn't really considered before.
They probably still going to adapt the work to fit their needs, and so companies probably still want actual artists for entertainment and artistic purposes.
Small companies (and individuals) probably can use it to circumvent costly stock images all together, so the lower tier photographers/artists there gonna have a problem.
[1] https://colab.research.google.com/github/TheLastBen/fast-sta...
The main challenge is finding the right balance between "make something that looks exactly like this" and "put it in a completely different context". Better similarity equals less flexibility.
For now, the most effective combination will be artist + AI, although it does feel a bit like those that incorporate it in their workflow are helping to dig their own grave.
1 Generate a bunch of characters or objects based on a prompt.
2 Pick the one you like
3 Tell the AI to extract the character's traits from that one picture
4 Miracle happens (I don't know. What does a "character definition file" look like?)
5 Make a new prompt, but add the character definition from step 4, so you get the same traits, only in a different setting or position, etc
Note that training is extremely expensive, and is beyond the capabilities of most end users. Here are the details of their training method:
> Given ~3-5 images of a subject we fine tune a text-to-image diffusion in two steps: (a) fine tuning the low-resolution text-to-image model with the input images paired with a text prompt containing a unique identifier and the name of the class the subject belongs to (e.g., "A photo of a [T] dog”), in parallel, we apply a class-specific prior preservation loss, which leverages the semantic prior that the model has on the class and encourages it to generate diverse instances belong to the subject's class by injecting the class name in the text prompt (e.g., "A photo of a dog”). (b) fine-tuning the super resolution components with pairs of low-resolution and high-resolution images taken from our input images set, which enables us to maintain high-fidelity to small details of the subject.
Each fine-tuned model is a copy of the original model. So if the model is 10GB, the fine tuned version will be a separate 10GB file. That might not sound like a lot, but it quickly adds up.
In this case, end users are artists. One could imagine a cloud-based art program which will fine tune on demand. That certainly seems like a good startup idea.
You can't run it on consumer hardware, but you can just rent a GPU (or use a free collab book) for a few hours to generate the model. Then you download and reuse it locally at will. Yes, you need storage and if you train often it can get expensive, but it is by no mean out of the end user, at least professional end users, capabilities. And of course there are growing libraries of freely available pretrained models.
> In this case, end users are artists. One could imagine a cloud-based art program which will fine tune on demand. That certainly seems like a good startup idea.
Very much agree about this. At least for a while, I strongly believe that AI will just be another tool for artists willing to embrace it, far from replacing them.
This area is moving really fast, one-click solutions are already being created.
People are also averaging weights of multiple models to create new models based on multiple other dreambooth models, and it works surprisingly well.
The stable diffusion reddit is a good place to see how all of this is developing.
I have an rtx 2070 with 8gb and it has been working quite well for me. However there are always models that will not fit. For those running on the cpu with potential nvme offload is not that bad. For example a single inference on bloom 7b (30gb of ram required just for weights) on a 32gb ram machine takes about 30s (it has to offload few GB to nvme). This is on zen 3 ryzen and with no gpu use. I can't wait to try cpus that support avx512.
Do you happen to have a link to that YouTube video you followed?
You can use it to make procedurally-generated art, but it's still obvious what the source material was.
One obvious counterexample is stylegan interpolations. If you interpolate between two images, the midpoint is usually unique — it’s often not obvious what the source material was. (E.g. Gwern’s anime interpolations; sure, they’re anime faces, but from where? “All of danbooru” might as well be “all styles of anime ever created.”)
Maybe someone else can argue the point further.
DreamBooth really should have destroyed all doubt about this point point, as the AI can generated highly specific images of subjects that aren't even in the original training set.
I haven't seen any concrete examples of this yet.
There's a lot of good discussion to be had around the ethics of AI art, training on copyrighted materials, etc. But it is equivocally not just copy pasting from a database of images.
> But it is equivocally not just copy pasting from a database of images.
Just because you put a lossy compression step in the middle doesn't mean it's not copy-paste anymore.
Just like a hipster/enthusiast with a more manual photography and development process
some people might appreciate it, the market likely wont but thats already the same before AI
I don't have much of a solarpunk outlook on this, I don't think it's going to be that beneficial. It'd be nice to be wrong though.
But in doing that, I've spent years learning how the micro-timing of musical grooves work (some of it's very obvious, like how a heavy snare backbeat will lag slightly behind what the perfect time is, almost to the point of being a 'flam')
As a result, I was able to make an album that conveyed certain kinds of groove not immediately accessible to novice musicians (or drummers)… but given the same information, any schmoe could push a button and have a preset in his DAW produce the same effect.
I think to some extent if the person doesn't really understand the purpose or need for such an effect, their grasp of how to implement and craft it will be pretty loose. If you don't know why you're doing the effortless thing you're not going to guide it very well and your results will be kind of generic.
Rather than thinking of years of skill, maybe call it years of focus, or years of purpose? To some extent, we collectively respond to creations with a profound sense of purpose. If that purpose comes out of AI it will have to be something from the AI, and not simply a blind reflection of us and our crudest drives.
its funny, people have been saying this every few years since the 1980s
I'm in an odd position on the subject myself: I've long wanted to get facile and slick in my art creation, able to do any genre or style, but it backfires. The only thing that keeps me going is, pursuing really eccentric pursuits. Now I think that's a blessing, because ability to be facile is seriously devalued now.
If you're a human artist very derivative of Greg Rutkowski right now, you're more screwed than I could possibly imagine. And yet, the reason someone would do that is because the style is popular on an extremely basic level: a far cry from trying to be a Basquiat from scratch.
I think a very real concern is, can AI adopt and popularize a trending style so fast that it obscures the initial artist from which the AI is trained? If novelty is what's needed, how small of a seed is enough to spawn 10000 AI derivatives, which may themselves be popularizations of the original concept?
Maybe the future of AI is to proliferate any new innovation so profusely that it inevitably chokes out the innovator and hybridizes with 'what's commonly popular' which is the guts of the neural network that makes AI what it is.
Maybe in some fields this has already begun to happen with a little assistance from AI-guiding humans. Take it up a level: what about AI prompt crafting? Can you define, not just what will be commonly accepted as popular, but what will be innovative and trendy, in a neural network?
In Star Trek there still seem to be writers and creators of holo deck programs, even though you could probably have the computer serve a blend of different stories. Perhaps “a craft” will still be seen as admirable.
Be careful with that, you may end up being more productive but also be expected to do more with fewer human resources or teammates. This will definitely work for some but I can't see this not becoming a race to the bottom for the creators. There will be tons of benefits too but just like with automation, it could be a double edge sword.
I used to work in the animation industry. My friends who are still in it tell me that there is constant pressure from the studios to combine multiple roles into one, without any increase in pay. They've got a strong union so there's a lot of pressure to keep union shops as places with a lot of ways for you to bounce around and spend a year really delving into the fine details of one particular part of the craft, while still being able to afford to live in LA.
I can understand that. In the visual domain, especially: how often haven't I thought about a story, my own or someone else's, and wished I could show people. I can also understand the dismay of some artists, who did all that hard work, and suddenly find that it's even harder to reach people with those songs, pictures and stories bubbling in them, because now so many others are able to express theirs too.
Markets for many kinds of art have already been driven to near zero, but I wonder if AI will drive them negative: if you want a human audience to give you their much-contended attention, you'll have to start paying them.
Inventor: I made a thing! I want people to know about it so they will want it, and pay me for my cool thing.
Public: We can't notice you over the million other things to pay attention to.
Inventor: Hey Mr. Zuckerberg, here's $100. MAKE THEM NOTICE ME.
People will continue to paint, like they did after the invention of the photograph and photoshop. Same for music, films, etc.
In chess, people would still rather see people play chess than a robot. The top chess players in the world started as a brute force obsession about the game. Those would go on to teach the next generation. They advent of computers allowed for historically statically advantaged moves. ML came along and disrupted even further.
Now many of the top chess players consult the ML chess oracle.
I see the same thing happening in a lot of areas: grammar, image generation, text replies.
I see a world where humans are celebrated for their humanness while machines assist.
Yes, considering the story behind art pieces and the artist does significantly impact the way we interact with art, but you can still just like a painting or a piece of pixel art without knowing anything else about it. While Chess is about the players, art has products that exist on their own.
I guess that depends on what the definition of art is. If it's a digitally rendered illustration then maybe.
If it's a physical object crafted to imperfect perfections then no. AI can probably produce some alternative version of Guernica but that's just a fascimile of something that's already been made in a digital space.
I can see an artist using these tools as a way to produce ground-breaking work in the future but I would wager the artist that does that could make good art without any of these tools. You still need a craftsman to master the tools and without knowing the basics you're left with images of Elon Musk as a Disney princess on repeat.
(I admit that I intentionally misunderstood your post a bit. I suppose you are talking about images and text for general use, not only in an artistic context)
And hand-drawn images look better if the artist is skilled.
https://kofaniv.snk-corp.co.jp/english/info/15th_anniv/2d_do...
In the past this was called roto scoping and used to capture outlines and movement of real objects for animation. The end result is still a hand drawn and shaded object instead of just a screenshot of a posed 3D model.
It's never a sprite though. Even back on the Super Nintendo game sprites had 9 angles * lots of actions * lots of frames for every animation. Multiply that by different lighting in modern games, and using the power of a 3D engine starts looking like an obvious choice. I know lots of games do this for environments. I assumed they did it for everything.
Every motion in every direction has to be animated. It's akin to creating animation by drawing every change on a separate sheet of paper, photographing it, and then stitching together.
The barrier of entry is indeed low. It doesn't take much to start sketching some pixel images. Making a coherent animated whole out if it? Very labor intensive and hard.
(I am instead sure they revealed that the characters in video animation were re-touched recordings of created puppets.)
Many techniques may be part of the creative process and of the production.
Guilty Gears for example started out with 2D sprites and then went 3D with Guilty Gears Xrd[1][2], but with a lot of trickery to keep it looking 2D'ish, though that always was more anime-style than pixel art.
I think this is a better video from him showing how it looks in practice: https://www.youtube.com/watch?v=1FrIBkuq0ZI&t=420s
It requires some specific models and techniques that wouldn't look good in 3D, though, it's not just cell shading.
But it lets them have a wild variety of player skins and enemy animations without having to redraw everything by hand.
Basically, if you look at where CG is used in contemporary anime, that's where the line is drawn. Many shows will CG their vehicle shots and parts of their action scenes since they call for a lot of perspective drawing with swooping camera movements. They may use and trace over CG dolls to do establishing shots or get static posed characters in a scene, but once detailed acting is called for they tend to revert to drawing in keyframes. Again, generally relying on reference footage to get the acting down, but modifying it to fit the character designs.
Huh? Half the point of pixel art is that single pixels make a difference. That art is way above the threshold where one pixel can make too much of a difference.
The tools behind AI are fine and have honestly existed for decades. If anyone is up against AI because it isn't authentic or something, that's a fools errand really because people, artists themselves and developers, will find them useful. The problem is and always will be how the data you fit on and how you obtained it.
Chances are really very low that this will become illegal and enforceable. It would require some very draconian laws, whilst copyright legislation is low priority in government circles, even more so in these times.
Even if somehow this would be outlawed in the US, nobody cares internationally. Right now, on Amazon you can buy knockoffs of millions of products from China that violate IP/copyright. Nobody cares. Do you think they will care about something as worthless as a digital image? A digital image that can't even be reliably detected as being AI generated?
And there's yet another work-around. Scrape images that don't require consent or make consent part of terms and conditions. Google made Google Photos free for about a decade, and trained it for free on all your stuff.
Here's another option: https://old.reddit.com/r/StableDiffusion/search?q=pixel+mode...
For example:
* Pixel art sprite sheets: https://old.reddit.com/r/StableDiffusion/comments/y54isd/cou...
* pixel-art-v1: https://old.reddit.com/r/StableDiffusion/comments/yj1kbi/ive...
But even ignoring that, we have barely even stared exploring what we can do with the technology as is. A lot of it is still just experiments living in a git repository or need more GPU than the average person has. Give it a few more months or years, and you'll have it integrated into every major photo and video editor software and optimized to run on normal consumer hardware. That simple improvement in accessibility will have very wide reaching consequences just by itself, even without improving the underlying AI drastically.
And no, you won't replace the professional artists anytime soon, after all somebody still need to have the final say into what goes into the game, but it will drastically transform how that artist will work and the amount of content they'll be able to produce.
> e.g., unless it's a very low effort, low quality game.
The output of Midjourney and Co. already looks spectacular, easily better than a lot of games out there. I could easily see that replacing or enhancing a lot of art in 2D RPGs, point&click adventures or visual novels.
Everything that needs animation or 3D meshes will take a while longer, but for 2D games it's already more than good enough. It's really more an issue with artists and game developers still needing to catch up on all the rapid new developments that happened over the last few months.
Do you have any recommendations for (web)apps that allow you to generate images from prompts in good resolutions? How about ones for img2img?