TheseToonsDoNotExist: StyleGAN2-ADA trained on CG cartoon faces
thesetoonsdonotexist.com
thesetoonsdonotexist.com
Like you get Elsa with different hair or the up grandpa in a suit.
Once you compare them to the closest examples from the training date it becomes a lot less impressive than the implied “this face is completely out of the imagination of a AI model” turns out the model just imagined someone in the training set with the hair of someone else. Quite boring.
That's not how GANs work in general. With a limited training set the end result may appear as if they were stitched together samples, but that's not what happens under the hood.
Interpolation videos make it obvious that it encodes visual concepts, can freely manipulate them and even crank up parameters beyond anything found in the training set, thus giving exaggerated results.
One of them on mine is just Helen from the incredibles head with Elsas face. I couldn’t design a cartoon character from scratch, but I could definitely make a ‘new’ cartoon character if I’m allowed to take Homer Simpson’s head and paste Goofys face on top.
https://openaccess.thecvf.com/content_ICCV_2019/papers/Abdal...
given an arbitrary face, we can find its embedding in the latent space of the model. this shows that the model has the potential to generalise to real but unseen examples?
on the other hand, i suspect you might be observing a bias in the structuring of the latent space.
thispersondoesnotexist.com likely samples the latent space with a gaussian or uniform distribution, and while the latent space may contain the full spectrum of possibilities, the density of semantically meaningful embeddings may be structured around the distribution of the training set rather than a uniform or gaussian.
i'm stretching my understanding of the topic in trying to convey this.
Looking at these images, and not familiar with how the underlying CG training set is made, I wonder if the original series itself has some comparatively small set of latent features - dimensions you could adjust when drawing the faces - that the model is just learning, so that newly generated faces are effectively the same thing as if one had changed whatever setting you tweak when working with the underlying tool.
This shouldn't be the case unless you start actively looking for it.
It's just much easier to recognise with these cartoon characters than with realistic faces as there's naturally much less variety in the training material. Also the features are simplified to a point of being easily recognisable as well.
Edit: wire it up with Stripe and sell characters to Pixar? Ha!
[1] https://news.ycombinator.com/item?id=19144280
With real human faces it's almost impossible to tell but with these you can definitely pick a character per feature.
It's striking that some of these examples have distinct features from specific, identifiable datasets: we can occasionally recognize specific characters (the old man from Pixar's UP is getting mentioned a lot), but it also reproduces more general aesthetic patterns. Even when I can't recognize the source data, I can distinctly see in some of these faces "the Pixar look", and in others "the DreamWorks look".
Were I an IP lawyer, I would start thinking of arguments along the lines of "this technology simply obfuscates the source of plagiarisms". I would also start to think about trying to force anyone who uses this technology to disclose the sources of their training data, since a model trained largely on "the Pixar look" could be benefiting from Pixar's character design processes without having to hire any of Pixar's artists.
And, if I were philosophically inclined, I would also start thinking about how this is any different from hiring a random artist and instructing them to "design characters that look like Pixar characters".
I suspect that one key difference is that the human artist's success can't easily be measured, but the GAN's success can very easily be measured.
So in the end this technology might not be as "liberating" as people think it is.
In a way, this is a much better setup for artists and creatives. There isn't some giant licensing firm controlling your work. You simply buy or rent the best tools to make your work.
That said, it'll only be good for creatives and consumers if there is sufficient competition. And open source equivalents that still enable creation.
I have seen some really cool and wacky designs in the hallways of DreamWorks, Blue Sky, Pixar, etc. during my time in the industry. I would love to get all of those designs into a training set as well.
A rule-based system combined with a Transformer and CV-based postprocessing to filter the most plausible and interesting results would be awesome.
just to get lots of retarded-baby monkey-fish-frogs.
Edit: actually, yeah, after looking at more examples nearly every one has some amount of cross-eye/focus disorder where both eyes aren't pointing in the same direction.
I was suddenly struck by this question, and think there might be something to it. Clearly it was a standard feature of the training set.
Though exhibited more as some kind of a confused smirk. Which is of course also ubiquitous in cartoons for some reason.
1) how much approx cost will it be to rent the servers to train such models?
2) Can this be done on a home computer running a $1k nvidia card in a reasonable time?
3) Can i use free tools like google colab(?) for this purpose?
I've always been interested in learning more about this field but haven't really bothered because I feel it would cost an arm and leg just to experiment. Can somebody please shed some light on this?
Using a standard dataset like CelebA [1] and an "HQ" model (512x512) like StyleGAN2, training requires at least 1 GPU with 12GiB of VRAM and training of about a week with a single V100 GPU.
Depending on your provider of choice, this will cost anywhere from ~$514 (AWS), ~$420 (Google) to $210 (Lambda Labs, RTX 6000 - should be in the same ballpark).
If your training process is interruptible and can be resumed at any time (most training scripts support this), costs will drop dramatically for AWS and Google (think $50 to $200).
2) Yes. A used ~$200 Tesla K80 will do. Alternatively any NVIDIA card with at least 8 GiB of VRAM is capable of doing the job, but lower batch sizes and increased training time are to be expected. If you can use a dedicated machine with an RTX 3060 or a brand new A4000 (if you're willing to pay the premium), close to a week of training time can be achieved.
3) Yes*
*your work will be freely available to everyone and your training process is limited to 12h or so per day.
All in all I wouldn't recommend training a StyleGAN model from scratch anyway. Finetuning a pretrained model using your own dataset can be done much more quickly (think hours to a day or two) and on consumer-level hardware (I train my models on an old desktop with a GTX 1070).
This is interesting! Do you have some links about doing that?
My desktop computer has a GTX 1060 with 6 GB of VRAM. But hopefully I can use it for something like this.
I've only used Google Colab in the past, and only tried stuff with prompting existing models.
Would love to experiment a bit with fine-tuning models on my own datasets to get some kind of unique stuff.
Noob question here: how does that work? Do you run the scripts in the hosting provider's downtime, with something like a nighttime rate? Or what magic is this?
Google charges about 1/3rd of the usual hourly rate for such instances and AWS has a "market place" where you can bid for such instances (you name your max. price beforehand) and whenever an instance with your selected specs becomes available at that rate, you get it.
Hyperscalers like Google and AWS basically have two types of machines/VMs for rent: instances reserved for long term commitment (think months to years) and on-demand. Naturally there's peak demand times in each region (usually during business hours) followed by periods of low demand so there's heavy fluctuation.
Instead of just having their machines sitting idle while still costing money, they offer such idle resources at a heavily discounted rate, with the catch that as soon as regular demand rises again, your VM is being shut down to be offered at the normal hourly rate (you get notified so your script has some time to save its state).
It's similar to what hotels do - discounted rates are available most of the time, but whenever there's a convention in town or on national holidays, you get kicked out and prices quadruple. Only with hyperscalers prices stay the same and you simply lose your cheap resource.
The actual time at which VMs become available is pretty random. If you are ok with using multiple regions there's pretty much always instances available, though usually not for hours on end.
Yeah, I think that might be the problem here. While there's plenty of material for real human faces, there's only so many high quality 3D cartoon characters.
Wait until you find out how hand-drawn animation is made.
:^)
https://disney.fandom.com/wiki/List_of_recycled_animation_in...
Definitely something they did quite a bit more than just a couple of times.
Not saying there’s anything “wrong” with that per se btw. Just found it relevant to the discussion about comparing modern animation to factory output.
There are modern 3d movies I particularly dislike and many I don’t like or love. I just don’t agree that the medium is innately lacking.
That gripe aside, if you're just training on a bunch of headshots and generating new ones, it's been done, over and over at this point. Want to impress? Figure out how to generate a full sequence of coherent animation frames.
Edit: whenever it gets inspired by Mr Incredible the result ends up looking like Conan O'Brien.
I was going to set it up to mess around with until I saw the requirements.
The easiest win would probably be to have an algorithm pick between predetermined types of assets (heads, appendages, clothes, etc), reshape them without actually adding new geometry, and then doing essentially what the linked page does with skins and shaders.
I mean, that's about as close as we may get to the holonovels from Star Trek in our lifetime.
Say "computer, delete crowd and replace with cowboys" and it just does it using imagined designs consistent with the rest of the movie/game.