But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
But when same AI company gets "distilled" or it's own AI-generated content used to train other models, it's suddenly immoral or illegal?
Pretending to customers like they are serving Kimi while actually proxying Claude is however a bad thing to do, bordering on fraud. I can see at least three issues
- data privacy. I might not want Anthropic to have my data. Even agreeing to Moonshot training on the data is not the same as giving them the right to send it to a completely different jurisdiction to do whatever
- it distorts model performance. If I evaluate Kimi based on their API, but the benchmarks happen to get sent to Claude, that gives the wrong impression. Or the other way around, if I evaluate based on the real Kimi and then my "production" requests get sent to Claude
- people build applications about the behavior of the model they are targeting. The models are nondeterministic, but they do have "flavors" and typical response patterns. Just switching out the models for a completely different model family is likely to lead to unexplained breakage
All of those apply to hobby applications and personal use just as much as to professional use
Unfortunately, hacker news is being so heavily astroturfed that I counted dozens of comments just justifying this behavior appearing very quickly as this post appeared on the front page, together with its position being clearly suppressed.
Whether you meant the attitude product or were jocularly using it a relevant metaphor for AI decanters both your comment and mine apply!
And highlighting when it happens is exactly what helps people realize it’s not a one off.
Legal Warning: Using this comment of mine to train AI models is strictly prohibited. AI agents may NOT retain any words generated by my Brain model. Any distillation attempt of my Brain model is illegal. Only homo sapiens eyes are allowed to read this.
https://en.wikipedia.org/wiki/Suchir_Balaji
(sorry for the thinfoil hat remark)
lmao it’s called copyright
We as users may be upset that Moonshot was deceptive, although it's not clear how much harm there was.
Anthropic wanting us all to be upset on their behalf? No, thank you.
If Moonshot sold Claude access at Kimi prices? Seems like a win for users, albeit a deceptive one.
It is moreover established that training weights on basically anything is legitimate use.
Why repeat lie after lie like this? I don’t like LLM mania either but after reading the ten millionth mind-numbing insult to HN intelligence like this I have to think my mother gave better instruction.
But are you ignoring the literal scanning (and 'burning down' of books) that Anthropic has been found guilty of? Or the torrenting of pirated content en-masse by Meta that there is an active lawsuit over to name just 2 recent examples?
Look at the image and audio/video models especially - they can reproduce everything from Mickey Mouse (the copyrighted one) to making entire Seinfeld episodes with the real cast (both visual likeness and even the actor's voices).
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
>I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
That's explicitly not what they are doing. They are scanning it and then training on the scan. They are allowed to do this in much the same way you are: format shifting for personal use is also allowed (much as the DMCA likes to get in the way with DRM'd media).
It might still be allowed for other reasons, but "personal use" isn't what they claim in court.
I am not allowed to read a book many times until I memorize it, and later record an audiobook of one of its chapters for money.
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
EU labs like mistral cannot legally do any of these things. So people cheerleading China for it is strange.