apparently, there is a separate drama about hacking at SD and accusations of ownership of code.
apparently, there is a separate drama about hacking at SD and accusations of ownership of code.
This is extremely fascinating to me. In what Andrej Karpathy calls "software 2.0"[1], leaking your model is equivalent to leaking the main IP of your company. Unlike source code where it loses value quickly out of context (e.g. twitch leaked their code yet that didn't spawn a bunch of twitch clones), models can be fine-tuned and transfer-learning-ed for many other purposes
Since these models take millions of dollars to train, I see these sort of hacks becoming a thing! I wonder when companies will start adding "trap streets"[2] to prove that others are using their stolen models?
I don't know if it tripped an internal filter on HN or something. Here is the link to the post in case you're curious.
https://news.ycombinator.com/item?id=33146603
In case it matters I used the GitHub issue title as the name of the post - "Stable-diffusion-webui is using stolen code".
Here is the actual link I posted.
https://github.com/AUTOMATIC1111/stable-diffusion-webui/issu...
We can let that particular link go through if it's the best one on the topic but perhaps there is something else out there?
Kind of with automatic on this one, good imperative code can only take so many forms.
Plus there is the argument that no IP deserves protection, and that no IP owner should have any rights, ever, given that information is infinitely copyable and yearns to be free.
These arguments aren’t taken that seriously by people that understand investment, technology development, and risk mitigation though.
That we allow (even encourage) such blatant violations of natural laws and physics is one of the more contemptible attributes of us as a people, IMHO
Just because an industry exists doesn't mean it's valid. Like living things, business entities have a survival instinct that will fight anything that that threatens it. And the dismantling of their business plan is a pretty big threat. Like cancer, these things start small and metastasize until they endanger the life of the host.
Which violates a number of Reddit's conflict of interest rules for moderators.
Those don't exist. There is something called the "Reddiquette" which is by the community and completely informal.
Subreddits owned by companies is the new normal and if you look at games, most reddit users these days even prefer it to community-run ones. Not uncommon to have two subreddits at a launch of a game and the unofficial one will always "lose".
Besides that, Reddit itself reaches out to brands & product owners and influencers to make official subs for them, managed by those entities then.
Subreddits are the new Facebook Groups and Reddit is completely "mainstream". I wonder what's next in a couple of years. Maybe we can go back to forums :)
Games often link to the "official" reddit community from inside the game, so no surprise that traffic will be driven there
But even besides that, if you look at any time this situation happened in the last ~3 years, you'll always see more users being vocal about the preferring the official sub than the community one in comparison.
Reddit's demographics changed, and it just exploded over the last couple of years with "casual users", "went mainstream" or what ever one wants to call it.
Most users these days don't even know that "official subreddits" were something that was super unpopular and uncommon on Reddit.
It used to be official, although still informal. It was eventually relegated to an obscure page in their help desk, though.
This goes beyond that, taking the stance that by merely conforming your api to work with a user-provided proprietary checkpoint, you're in the wrong? This same philosophy forbids sharing open source game emulators, and we all know how that turned out (can be the best way to play a game).
now they're in "monetize later" and some rando's internet repo is more usable and has a better pipeline than their "dreamstudio", and to boot now these rare gtx aren't rare anymore thanks to the bitcoin crash, so enthusiast can readily use the model at home.
The additional drama includes that said open-source author has looked at the leaked code and concluded that it stole some of his open source work in the way it tunes text parameters.
Unless these AI companies want google and facebook to be literally the only companies in the world that can train large scale machine learning (by using their TOS to get licenses from their users) they should tread carefully.
In this particular case the leaked code apparently exposed that the proprietary codebase was also using the OSS developer's work without attribution.
There's an interesting anecdote around how stand-up comedians protect jokes against theft, given how weak copyright is on jokes: it's keying cars, poisoning drinks (generally non-fatally, but it's hard to have a good night on stage when your lower GI tract wants to be elsewhere), and never-work-in-this-town-again agreements. We put these protections into the law because the alternative isn't no protection; it's people-take-it-into-their-own-hands protection.
The users of these models and the developers are fundamentally different.
If it costs $10 million to find information/weightings/etc., our current legal system would consider that intellectual property which might not be copyrightable but would be considered IP theft if stolen.
If they were negligent in handling it, e.g. left it on a publicly accessible share and some member of the public stumbled into it, then trade secret protection would likely be lost.
If some employee violated their NDA and snuck it out-- well that would be a different matter. etc.
Another perspective that may become important is the fact that not all cultures share the same interpretation of copyright. In Japan there was a case in which a court ruled that selling a memory card with preloaded save data for a video game was a breach of the original work's integrity.[1]
This I think will get greater attention in the near future because a large portion of the interest in SD stems from generating new art derived from the styles of art on Pixiv, a Japanese website. The data for many popular forks of SD like Waifu Diffusion and the proprietary NovelAI model were sourced from Western sites like Danbooru, which has been known for violating copyright and artist takedown requests by reposting art without the creator's permission for many years. With the sheer popularity of SD and the fact that so much of the innovation came off of the backs of thousands of artists who weren't so much as asked for consent, it remains to be seen if attitudes towards those sites and this process of mass-scale data collection will remain the same in the near future.
I also have to wonder what the implications would have been if NovelAI ended up launching what is now the leaked model as a paid service, given the unresolved question of consent that surrounds the original data.
HN and the people who support SD can have their own opinions about copyright not applying in this specific case. They can delve into the technicalities of why they think the models are not copyrightable. But even beyond legal means, the artists can still ask the programmers to take everything down, and potentially be refused. The insistence that "it's different in this case" can break the hearts of people that see the world differently.
I think this will be a debate that transcends arguing over the technicalities of copyright, involving fundamentally differing cultural values of how the acts of creation and reproduction should be treated with respect. It will not end with "how will this fit into the existing (Western-centric) framework of copyright," but "what is the right thing to do."
[1] https://ja.m.wikipedia.org/wiki/%E3%81%A8%E3%81%8D%E3%82%81%...
Anyone who sets their ethics based on the law is probably acting like a big jerk.
:)
huh? It was launched as a paid (subscription only) service before even getting leaked.
Looking back it seems many in the Japanese community on Twitter aren't happy with discovering their art is being used as training data. In the last week since the NovelAI announcement the number of banned artists on Danbooru has doubled, approximately the same amount as all the banned artists that were registered in the site's entire history.
Beyond the training issue there is the fact that the art was reposted on another site at all, which is what many take note of. It seems Danbooru's initial statement to Japanese users made it seem like it was NovelAI's fault that things blew over, without bringing up the reposting part that was the root issue, and it didn't seem to be an effective apology as a result.
Possibly that's why artists taking down from Danbooru , it's their right.
There is likely a legal threat by NAI involved in the decision calculus here.
My understanding might be out of date, but as I recall it seemed that the code he thought was stolen from his open-source work actually originated from a third project that's MIT licensed.
1. Accusations that AUTOMATIC1111 (the web frontend developer) copied code from the NovelAI leak relating to the loading of hypernetworks
2. The leak revealing that Anlatan (the company behind NovelAI) had copied code from AUTOMATIC1111's repo (who, as above, Anlatan are accusing of copying from them) relating to the weighting of words. AUTOMATIC1111's repo does not have a permissive license to allow this
The third party MIT-licensed code is relevant to #1. Some code AUTOMATIC1111 was accused of copying from the leak (https://i.imgur.com/r1AkvBG.png) actually already appears in multiple older permissively-licensed public repos (https://github.com/lucidrains/perceiver-pytorch/blob/main/pe..., https://github.com/CompVis/stable-diffusion/blob/main/ldm/mo...), one of which was credited in the readme by AUTOMATIC1111.
For #2, the Anlatan CEO blamed it on an intern (https://i.imgur.com/BFjKG1V.png). The leak shows that the offending code was committed by the CEO (https://i.imgur.com/aLiA2tr.png), which doesn't necessarily rule it out originating from an intern (e.g: "send me the code over teams to review and I'll add it") but doesn't look great.
From other examples I'd say AUTOMATIC1111 did get a bit sloppy in terms of not following clean-room design regarding the leak, but I'm inclined to give some leeway to a solo developer making a hugely popular public tool for free.