How we built the Waifu Vending Machine
waifulabs.com
waifulabs.com
I am more into "disgusted anime girl that looks at you like you're trash" type and I couldn't find a waifu (even with their refinement steps).
Really impressed that this is even possible though!
Edit: "disgusted anime girl that looks at you like you're trash" is the theme of a recent book that got adapted into anime, so it's a bit of a fad currently. I thought that was worth mentioning.
The steps could be something like this: (1) Select character, (2) Select colors, (3) Select outfit, (4) Select pose
The reason I suggest this is I was a little confused on what the options were, what my future options would be, and how many options are left.
Sizigi Studios uses a dataset I put together (https://www.gwern.net/Danbooru2018) as their primary training corpus, AFAIK, and they were partially inspired by my application of GANs to anime art (as demonstrated on https://www.thiswaifudoesnotexist.net/ and see https://www.gwern.net/Faces for much more detail about every aspect of it), but I have never been involved with them and don't know much about what they've done other than what you can read in OP. It tickles me pink to see people following up on my anime GANs, though, especially as a startup! The Great Work goes on.
I could also see something like this having applications in https://en.wikipedia.org/wiki/Identicon generation.
There was a time where I didn’t listen to rock music; Pantera and Green Day would sound the same. Memorable is up to cultural fluency...
The key fingerprint is:
SHA256:s6N0OwlTDKjDez98kZRwUGZbTYaQUArv+EYC6sigFwA ben@eshwil
The key's randomart image is:
+---[RSA 2048]----+
|E ..o=*o.+o |
|. .oo+oo... |
|.... o=.. |
| o+. o = |
|o .oo ooS. |
|* ...+o oo |
|oo.. o+o+o |
| . o+o+o |
| .o.. |
+----[SHA256]-----+
(from https://blog.benjojo.co.uk/post/ssh-randomart-how-does-it-wo... )(yes, that's a thing)
https://medium.com/syncedreview/biggan-a-new-state-of-the-ar...
When shipping posters, use triangle tubes, not circular tubes, it saves you money.
[0] https://www.linkedin.com/feed/update/urn:li:activity:6549769...
To other people doing this in the future: bring (or order) a fat battery pack with an AC outlet for $100 so you don't have to keep swapping laptops and can use a mobile hotspot all day.
I don't have a sense for hardware requirements though. Does anyone have a good idea of how much time and money it would take to train such a model?
How expensive is it in terms of labelled data and compute? Do you know if anyone tried this for just ahegao faces?
All the stuff you see in that page (except the BigGAN ones) is unconditional, no labels. You just dump the images in and it figures it out. StyleGAN does support labels via one-hot embedding as I understand it, but I don't know how to use it so none of my experiments use it. A few people have mentioned or used it, but there's no good documentation about how to make it work, so... For unconditional samples, it depends on how many you have and how different they are. You can see in the various examples transfer learning with a few hundred to a few thousand (with and without data augmentation).
> Do you know if anyone tried this for just ahegao faces?
It's funny you ask that because I was corresponding with an anonymous who was using it for just that (and ball gags). He'd run into some issues with the encoder/editing functionality and wanted advice, but the regular transfer learning worked fine. He'd compiled a small dataset of a few hundred to a few thousand examples on his own, and it worked disturbingly well: sufficiently so I didn't want to write it up. (I try to keep my site SFW.)
How much of a plagiarist I am if I make my characters using this and pretend they are original?
Even if robots replace human illustration in short order, it will probably never stop being a fun hobby, and I imagine neural networks were bound to be at least involved in the process at some point. People still draw on paper even though it’s hard to argue against the benefits of modern digital drawing.
That does NOT make me happy. Self-driving cars are just around the corner!
As a result, every single artist whose work was included in that dataset has a clear, meaningful claim that each and every 'waifu' sold ($20, if customized, or $5 if random) by Sizigi Studios is an infringing derivative work. Coupled with at least one of the project authors' ready admissions -- in this very comment section -- of scraping image sites himself, I would say that this team is playing with fire. Even in the case that an algorithm's output is somehow found to be 'creative' rather than mechanistic, AND this specific application is found to be in all cases substantially transformative, there's STILL the original massive 2.5 TB of copyright infringement up front to deal with.
All an enterprising lawyer would need to begin is to search the BigQuery metadata for the 'artist' and 'copyright' tags on these images. Note of course that the 'copyright' tag is widely misused on boorus and similar image repositories to refer to the inspiring franchise; 'trademark' would be much more accurate descriptor.
EDIT: I do not mean to suggest that litigation from the use of the dataset in this ML (as opposed to the original, clearly infringing, download & redistribution) would in any way be an easy, one-sided case --- only that this scenario would represent nearly the worst possible test case imaginable for determining the future legality of ML, short of directly antagonizing the RIAA or MPAA.
Where exactly is the point between a derivative work and an original work that was inspired by something? A lot of fan art I see clearly depicts a character from some franchise in a style that is close to the original, but say, in a new pose or setting. Is that copyright infringement? Trademark infringement? What if it's an original character in the exact style from the franchise?
Some artists sell this kind of art as their own at conventions and will aggressively try to remove reposts on the web. On the other extreme, I've seen other artists tag any fan art remotely connected with some franchise with "copyright by <franchise owner>" and denounce any rights to their work.
The fact that other datasets are also massively infringing is evidence of a severe problem for the field, not an exoneration.
However, in this specific case, I would expect most courts to place heavy weight on the clear, initial massive-scale infringment from the dataset alone to conclude that no good-faith effort (or, apparently, any effort at all) was made to avoid trampling these artists' rights. Such 'dirty hands' would potentially discredit any attempt to claim original creative expression, rather than commercialism, motivated the creation of this ML model.
This represent very nearly the worst possible case to serve as a potential test case for the legality of ML techniques.
Other datasets, like the Open Images Dataset (https://arxiv.org/abs/1811.00982) explicitly recognize and address this concern in their curation of included images.
However, it is not a standard part of artist training to obtain and redistribute, without license, the (in this case millions) of paintings they studied.
> Never to the best of my knowledge has this been used to argue that a picture with no visible elements of another infringes.
One only needs to look as far back as 2013, in Williams v. Bridgeport Music, to find such a thing not merely argued, but successfully litigated. In this case, the estate of Marvin Gaye alleged that Robin Thicke's "Blurred Lines" copied the 'feel' and 'sound' of "Got to Give It Up" despite containing no samples or even an identical chord progression.
Perhaps more surprising to you will be the fact that the court found in favor of Gaye's estate, i.e. that "Blurred Lines" was infringing!
A not-insubstantial factor in reaching this decision was, as I alluded to above, the attitude of the defendant regarding the infringement. Thicke testified "No" when asked if he considered himself an honest person, and admitted that "Got to Give It Up" was a direct inspiration for the song.
It would be difficult to argue that the data used to create your ML was anything BUT its explicit inspiration, and as I've mentioned in other posts, this is compounded by the fact that the acquiring the initial dataset is itself a separate and very clear-cut case of copyright infringement.
In the interest of good discourse, do note that many legal scholars and industry experts were, admittedly, shocked by the decision and decry it as fundamentally mistaken. Nevertheless it is now certainly precedential caselaw.
It would be quite interesting to compare the generated images with the dataset using a tineye-style image matcher. I wouldn't be surprised if large segments of the generated pictures are outright identical to some image from the dataset.
Maybe, but they probably won't right? Like, if you had to bet money on if anyone is going to sue these guys over the next couple years, which side would you bet on?
I'd bet on "No". The reality is that these people probably made some X0,000 dollars on this, and nobody is going to bother to pay lawyers a bunch of money to go after something, that only kinda sorta looks like infringement to machine learning experts who are squinting really hard, all in order to claim their 1/3millionth percentage claim on it.
I don't see anybody wasting their time and money to go after a machine learning waifu vending machine that is hosted at a couple anime conventions.
You're almost making the argument that everyone who has ever read a pirated book or other work (I know plenty of people who started their whole career that way...) and then uses that knowledge is guilty, which is just a perfect example of how ridiculously insane copyright law is.
"Everything is a derivative work."
"Stand on the shoulders of giants."
Of course, it helps that some series like Touhou have explicit copyright licenses allowing this. But basically everything's OK unless you try to make porn for a Nintendo game.
It is true that 'boorus' do not bother with artist consent, and I think by and large the sentiment from Japanese artists has been fairly negative, if a bit muted (perhaps this is a bit of a cultural thing?) Despite that, the dynamic between boorus and content creators has always been complicated.
Boorus have some positives even for artists. For one thing, they organize images pretty extensively; it is not abnormal to see a booru post with 50 to 100 tags, and it's probably fairly uncommon to see less than 10 tags - so they're great for searching for images. Because of that, they are an excellent place to search for inspiration or references, and I definitely know folks that do this. They also act as content aggregates, which does help people discover content and artists, especially combined with tags. Boorus do tend to comply with "do not post" requests, though that obviously does not mean artists are then implicitly consenting to their work being posted or used.
But of course, the boorus themselves are not charities. Boorus tend to run ads or even accept direct payment from users in exchange for features. Personally, I do find this unethical, and I highly recommend you do not browse a booru without an adblocker, not even because of ethical concerns but due to the fact that many of them have pretty nasty advertising (Also, be aware that many boorus, like Gelbooru for example, are not particularly shameful about explicit content, and neither are their ads.)
I was involved in a small scale booru years ago in my youth, for a particular interest. It's funny, I was so wrapped up in the utility of organizing and categorizing images that I had not even considered the issue of artist consent or copyright (I was not making any money, in my case, at the very least.) Ultimately, an artist complained that their work was repeatedly posted without proper links, something we did our best to discourage, and I shut the whole thing down in short order, for better or worse.
Like many things on the internet, boorus would not be nearly as useful without their flagrant disregard for copyright. It's not just booru owners in this debate, though - some online artists take a view that copying and sharing online is both inevitable and the point, while others, probably moreso these days due to an increase in bad actors, take a hardline stance against copyright infringement. The law clearly and plainly sides with the latter, and I think I do too. Before sites like Pixiv existed, there were very few resources for finding Japanese illustrations in one place, much of it scattered around in fc2 blogs and Geocities sites (Yes, even until pretty recently! Geocities Japan outlived many other Geocities regions.) Nowadays, there's not nearly as bad of a discovery problem with Japanese art, and the boorus feel more parasitic than they once did.
(Aside: I think many software engineers are used to being very technical about copyright, even when they do not understand the details correctly. Most artists I know do not dabble much into the details of copyright or licensing, and may even have some difficulties grasping the implications of say, Creative Commons terms. It's important to recognize that not everyone has the same point of view and things that seem obvious to us may not even make sense to others.)
Of course, though, that a booru is basically massive copyright infringement is one thing. Now we're talking about neural networks built from them. And damn, that is complicated, and probably breaking new ground. I'm sure people with strong opinions will claim that it is very black and white, but I can't see this as anything less than a dilemma. I agree the model clearly constitutes copyright infringement if redistributed, but the results... that seems deeply complicated.
On one hand, for all of their impressive strides as of late, current neural network algorithms are not really that impressive in their creativity. It's clear they are very heavily influenced by the data sets. That said, though... at the end of the day, our brains are also 'neural nets.' If we design a sufficiently advanced neural net and feed it a handful of pictures from Danbooru, and it is able to spit out similar but clearly distinct art, how is that any more copyright infringement than a human artist that draws inspiration from the same images?
I think people predicted that the copyright situation would get complicated when DeepDream showed up, but it's nothing compared to the calamity that could occur as a result of this kind of data set.
However, all of that is irrelevant when the concerns are the distribution of the site's contents by torrent, or the use of those copyrighted images in making derivative works. In both cases the party in question is now first-party to the concern, either the distribution or the derivation.
As globalization moves on there are more and more people without romantic partners and/or close friends nearby and they'll use their imagination to fulfill desires they're missing.
Because yea, it's creepy.
Are you sure that you aren't just using "creepy" to shut down non-traditional forms of private sexuality?
Also, just cause you find it creepy doesn't mean it's reciprocated by everyone else in the world. Besides, the usage in that booth was very tongue in cheek.
Sarcasm aside, this is one of the many, many examples of choosing a scapegoat to frame an entire sexuality, race, or any group of people with a common interest as evil while completely ignoring any and all context. People are not animals and possess some degree of responsibility and the ability to tell reality from fiction. Unless someone presents some hard evidence that stylized drawings lead to actual attacks against real children (and to my knowledge, this simply is not true; in fact, it's easy to argue the opposite) we need to stop with this puritan outrage like we stopped blaming computer games for any and all violent crime back in the late 90s.
Are you making fun of my engagement?
Also, the post title is misspelled. It's "building", not "builing".
Morally, one can differentiate between virtual murder and virtual pedophilia and condemn virtual pedophilia while consistently enjoying games and other media depicting murder - but as Gary Young pointed out in his piece on the Gamer's Dilemma, it requires us to accept moral relativism.
Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should.
It claims that Japan has the lowest birth rate ever. And while it may have the lowest number of births, that is to be expected in any country with a birth rate below replacement - I.e basically the entire developed world.
In fact, as opposed to the US, UK, Canada, Australia and New Zealand, Japan’s fertility rate has actually been increasing for some years now. See details here: https://fred.stlouisfed.org/series/SPDYNTFRTINJPN
I really wish this meme would die, but it obviously drives clicks so people keep publishing it.
It seems to have started going up in 2005. That's really interesting and unexpected. Has anyone proposed an explanation? I know they've been trying to encourage people to form families for a while. Maybe some of their measures have worked?
I personally never looked for a partner because I'd rather spend more time watching anime.
There's an interesting anecdote to support it. When the Soviet Union (with all its faults) was established, it has provided education to much broader masses of the population than ever before. You know how it turned out? When you read the biographies of the many of the brightest academics of the USSR, many of them are descendants of a typical peasant family. Some of their parents weren't able to read or write, had a lot of children, etc. This "genetic handicap" didn't stop them from getting a degree and becoming bright scientists.
Yes.
> Judging from a few of the latest pop-sci books
That's not really much better than “judging from my newspaper horoscope”.
> I was under the impression that culture and nurture play a much greater role: high quality medicine, food rich in nutrients, good education available to all.
Those play a huge role in whether genetic potential will be reached, genetics play a huge role in the actual outcome given similar environment.
> When you read the biographies of the many of the brightest academics of the USSR, many of them are descendants of a typical peasant family. Some of their parents weren't able to read or write, had a lot of children, etc. This "genetic handicap"
You describe an socioeconomic handicap, not a genetic one, so it doesn't really say anything about the effects of genetics.
Why?
I'd rather preserve the genes of the people who care so much about human relationships that they forego any craft, but neither extreme is really optimal.
Genes that make people form families rather than work in their craft are not exactly in a bad place, so excuse me for not being concerned about them.
Of course that doesn't work so well when the average fertility rate is less than 2 children per couple.
Also, the same gene can sometimes have a positive and negative effect on your chances of reproducing at the same time. For instance, suppose there's a gene that makes you want to work out instead of have sex. That might just increase your chances more than decrease in an environment where you need strength to survive (or that partners are attracted to that).
Oh please! There is a world outside of the developed world. There are entire tribes of people all over the world that never even heard of the internet. Which now that you mention it sci-fi hardly ever addresses in futuristic stories.
I could bring up old racial stereotypes regarding Asian people to further bolster this point but I think most people are aware of this.
[0] https://www.apa.org/news/press/releases/2014/03/black-boys-o...
So you're not crazy. These are really young looking. It is probably just representative of the dataset. The dataset itself wasn't chosen for any malicious reason, the people who like this stuff just happen to be very prolific artists so there's a machine learning scale amount of it.
Of course, it's just my perception and your interpretation may be different.