Remaking old computer graphics with AI image generation
jalammar.github.io
jalammar.github.io
You can force a model to generate nearly the same actual pixels with DreamBooth, which can be interesting for putting people’s faces in a picture, but otherwise I’d call it overfitting.
But there is a paper about it: https://www.aaai.org/Papers/Symposia/Spring/2007/SS-07-05/SS...
So in theory you select one photo with the AI image generator, create variants of it with separate image tools, then build a fine-tuned model based on some cherry-picked variants.
I think this will get easier as AI image tools focus more on depth and 3D modelling.
The “aiactors” subreddit has some interesting experiments along these lines.
- dreambooth, ~15-20 minutes finetuning but generally generates high quality and diverse outputs if trained properly,
- textual inversion, you essentially find a new "word" in the embedding space that describes the object/person, this can generate good results, but generally less effective than dreambooth,
- LORA finetuning[1], similar to dreambooth, but you're essentially finetuning the weight deltas to achieve the look, faster than dreambooth, much smaller output.
...but, all of these can't maintain consistency.
All they can do is generate the same 'concept'. For example, 'pictures of batman' will always generate pictures that are recognizably batman.
However, good luck generating comic cells; there is nothing (that I'm aware of) that will let you generate consistency across images; every cell will have a subtly different batman, with a different background, different props, different lighting, etc.
The image-to-image (and depth-to-image) pipelines will let you generate structurally consistent outputs (eg. here is a bed, here is a building), but they will still be completely distinct in detail, and lack consistency.
This is why all animations using this tech have that 'hand drawn jitter' to them, because it's basically not possible (currently) to say: "an image of batman in a new pose, but that is like this previous frame".
So... to the OP's question:
Recognizable outputs? Yes sure, you've already been able to generate 'a picture of a dog'.
New outputs? Yeah! You can train it for something like 'a picture of 'Renata Glasc the Chem-Baroness' now.
Consistency across outputs? No, not really. Not at all.
As for consistency of character details, I think that will depend on how many images you use to train dreambooth etc. and how varied those images are.[1]
For the animation stuff where you need frame to frame consistency, the new diffusion based video models show that it's possible [1][2]. These are not open source yet as far I know, but it's highly likely that we'll get them within a few months.
There's no difference between those things. It's a specific label that directs the diffusion model. It doesn't matter if your label is 'dog' or 'betty' (ie. my personal dog). Anyway...
> it's highly likely that we'll get them within a few months.
Yep! It's not a technical limitation of the technology for sure; but the OP asked:
> Is there a way to have current AI tools ...
...and right now you can't do it with the current AI tools that are publicly available.
what i mean is, assuming this technology moves forward, and GPUs continue increasing VRAM as they have, and enough people are interested in doing extremely detailed tagging with small shapes, the sorts of issues you're talking about will go away over time. Or, alternatively, someone or a group could develop a way to scan hundreds of outputs and collate them according to similarity, allowing a human to use batches that are similar enough to do something like short comics or whatever. As it stands, when i do txt2img or img2img i will run off 20-40 images. I'm also wondering how much seed fiddling could be done - when i first got "Anything v3.0" every image was some person sitting at a dining table near a window with food in front of them, dozens in a row. I have no idea how it happened, but there was enough global cohesion between images i thought it was trained on just that for the first hour or so.
Each of the below images is a set of 4 images (i think generally called a grid in SD), so each image is a set of 4 "2 panel comic strips" - they aren't really intended to flow between the grid squares, but you'll notice that the clothing, hairstyles, etc between strips matches, even if they don't match between individual images. My personal favorite - and the one i used for something online, is the top left set in the first .png https://i.imgur.com/BWek3YI.png https://i.imgur.com/LHchsj5.png
P.S. if anyone knows what the source art could possibly be, let me know?
https://huggingface.co/docs/diffusers/training/text_inversio...
In automatic1111 UI you can alternate between prompts e.g. "Closeup portrait of (elon musk | Jeff bezos | bill gates)". Final image will be a face that look like all three. See this https://i.redd.it/8uq52mnausu91.png
Now do the same with two people but invert the gender. The female version of what I gave example of won't look like anything you know about. And it will remain consistent.
It kind of works.
[1] - https://www.scenario.gg/
[2] - https://twitter.com/Beekzor/status/1608862875862589441?s=20
I'd recommend giving it a shot if you have an Nvidia GPU with ≥4GB VRAM.
Edit: There are also training and hypernetworks, but they require a body of source material, keywording, and significantly more time and compute resources, so I haven't attempted either.
It's a cool exercise but using this for a real-world project would eliminate any attempt at producing an artistic "voice". AI image generation excels only at generating stock art and placeholder content.
Also, one model doesn't speak for them all. I have the problem that my results are often too consistent with certain models, largely because of prompt complexity, lack of wildcarding, etc.
[0] https://github.com/Stability-AI/stablediffusion#image-modifi...
You should take a look at embeddings too. They are tiny files, no more than 128kB, that have a huge influence on the final output. You put the files in the embeddings folder and use the filename in your prompt. Ideally the filename is a unique word so it doesn’t interfere with the normal prompt logic.
You can find the best embeddings in the stable diffusion discord.
I've seen this stated but it has not been my experience. Nor have I seen solid examples of it. For 2.0 (haven't tried 2.1) I find the model very finicky and unstable. Any prompt which works well also seems to work at least as well in 1.5.
But that's just me.
I seriously have never understood why what gets published in these blog posts isn't just lower res especially since this is precisely about old video game graphics.
I usually don't notice those. It's only when someone mentions it.
I literally just registered spritesheet.ai yesterday.
and I've already had mild success training a model on spritesheet data.
But I'll be launching one on alexbrown.io sometime in the next few weeks to document my projects.
https://www.reddit.com/r/StableDiffusion/comments/yj1kbi/ive...
https://www.spriters-resource.com
Fingers crossed.
It's a decent writeup on the process of trying to generate specific images using text prompts, I guess, with the conclusion that it's really hard, and in some cases basically impossible (hence the lack of the three forehead eyes).
In contrast, I’ve been using chatGPT multiple times daily