To avoid future situations where money is invested on hype only with no regards to what societal disruption it causes.
Are you from Europe or South America? Or just asking on their behalf?
They are allies of the US, which includes economic spheres of influence. The premise of the petrodollar is that allies do better by being a part of it than not, which has been true for many, many decades now. China’s rise is providing an alternative for the first time in 70+ years, but so far it’s very unclear if these benefits will truly extend to allies of China or not. For example, China sends their own laborers when building infrastructure in Africa. Sure there’s new infrastructure but also debt to CCP without any benefits of knowledge transfer or local employment.
Serious question.
I'm from the US, and I think it's hilarious.
In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions
Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.
- The Trail of Tears
- The Tuskegee syphilis study
- Use of Agent Orange in the Vietnam War
- Open Air biological warfare testts in civilians eg. in 1950 San Francisco
- The only use of nuclear weapons against civilians?
- Coca Cola and "american culture"
- Neoliberalist economy
- Spreading blame for their sins to other "white" nations
plus one: The text input method to HN comments :(
you are comparing a wooden stick to a fighter yet, try again
Another thing I would try if I had access to the models and enough proxies to hide behind is asking for advice on software/movie piracy or seeing to what extent the models can be elicited to straight up argue against the validity of intellectual property, though there it seems more probable to me that the US models would be permissive.
> The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).
Then make the case why it is better for the world that the Tiananmen Square massacre is memoryholed. You said A, now say B.
In the meantime, I'll make the case against it, and it's simple: totalitarian control of historical truth, by definition, to be 100% watertight, needs control of the whole globe. It doesn't mean everything needs to be controlled, it means everything needs to be controlled by at least an entity that cooperates on this matter. I.e. another totalitarian bloc.
That makes the CCP, just by their insistence about Tiananmen -- nothing additional required, at all, they could not harm a fly and have no prisons and it would make zero difference -- an enemy, a threat marching towards any thinking human who wants to have agency and dignity. Actually, it's more like a river flowing to the ocean, people in the CCP can have lofty ideals about honesty and factual truth, the system they require to survive in turn requires this to survive, as it is. It requires human spontaneity and human freedom to be dead, completely. That is what totalitarianism is.
And by the same token, we must be wary of those who take that lightly. A friend falling asleep at the wheel will kill you just the same as an assassin who messed with your car.
That is not "as opposed to the US or the EU or Russia or North Korea". It is strictly in addition. No other regime can behind another regime or use it as an excuse. Using the US to distract from the CCP is as odious as doing the reverse.
> I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me
So you want to work in an office wearing a tie. Well, people who want to work as roadies and fight all night would never hang with you if they found out. And if that was your priority, you would be more afraid of NOT being blacklisted by the corpos than not being part of them.
But seriously, are you saying you are as afraid of being found out as the person you are, because you're not already wearing that on your sleeve, in exactly the same way as people who have to fear that they and their family get disappeared, tortured, for remembering friends that got murdered? That would be totally a you problem.
Of course that might change in the future but as long as the Chinese companies continue publishing their research and models it only makes it easier for third parties to catch up with them.
"Thank you," the KGB says. "We do our best but truly, it's nothing compared to American propaganda. Your people believe everything your state media tells them."
The CIA agent drops his drink in shock and disgust. "Thank you friend, but you must be confused... There's no propaganda in America."
Also generally the USA had never done crimes in the scale of the CCP (around 30+ million dead)
First, whether we like it or not (generally not), a fuck ton of money has been invested in US AI companies, data centers, RLHF datasets amongst other datasets, etc. If that were to go to 0 that’d be quite disastrous. Alternatively, if it goes well, it’s great for the US (and to a lesser extent allies) economy and global standing.
Relatedly, tech has been a huge power house for the US economy for decades now. If the main driver of growth goes to China, what replaces it? Along with all the potential tax money, foreign investment, etc?
Next, patriotism / nationalism. This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative. It also will likely create social strife with the knock on effects leading to even more bad populist ideas that just further diminish the country and tear society’s fabric apart.
None of this is hilarious.
This process is already well underway. The current massive over investment into the AI hype bubble is the final death throes of a failed economy trying to keep its head above water.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
We're talking about things like text people wrote, not some kind of raw data floating out in the ether.
Ore has value, a different kind of value than the output of the refinery.
That would stun me, but it's a little hard to read.
https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?
> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.
It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
... but Chinese SOTA foundries directly using distillation as fair game.
I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
What is more reasonable:
- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.
- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.
> ... but Chinese SOTA foundries directly using distillation as fair game.
As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.
LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..
If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.
Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
You can't distill what you are not given - simple as that.
Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.
I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.
Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.
Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.
It's hard to draw the line.
But the Chinese models are absolutely distilling - and would not be competitive without this distillation.
At the same time, there's a lot of real innovation and regular building going on at the same time over there.
I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.
BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.
This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.
Moreover - the 'reasoning traces' are not required for distillation at all.
Finally - it's entirely possible for them to have used Fable for later stage fine tuning.
It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.
Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise.
Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge.
If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke.
Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK.
You are redefining words. The consensus among people who work in this space is that "distilling" is what's actually going on here.
> who are just as anti-Chinese as Anthropic
Conflating criticism of IP theft with being "anti-Chinese" is a standard PRC influence playbook technique.
And furthermore, OpenAI's market strategy is to win through regulatory capture. They are financially incentivized for Anthropic to be distilled by PRC labs and to be undercut by open models. Their claim about Kimi not being explainable due to distillation is not a factual claim - it's marketing from a company owned by Sam Altman.
Although, it does conclusively disprove your claim about the meaning of distillation, because you cannot say that "Kimi can't be explained by distilling" unless the consensus definition of "distillation" is such that it could be done on the summarized reasoning traces that Anthropic models expose.
You can interpret it as you choose, but a much more obvious reason he [OpenAI's Dean Ball] might say it can't be distilled is because it can't be distilled. You can't distill alcohol out of orange juice.
...and, as everyone in the frontier labs knows, this is a lie, because that's not how distillation is defined.
I know that I won't convinced you, because you're quite possibly a PRC agent, but for all the other HN readers coming to this thread in the future to look at this failure of propaganda: just ask a model.
User: according to standard LLM lab parlance, can you "distill" one model from another if the model being distilled from does not expose a thinking trace?
GPT-5.6 Sol: Yes. In standard LLM terminology, you can distill one model from another even if the teacher model does not expose a chain-of-thought or "thinking trace."
Sonnet 5: Yes. "Distillation" broadly means training a student model to replicate a teacher model's outputs (or output distribution), and this doesn't require access to the teacher's chain-of-thought.
That's all she wrote. You're lying, and even the models know it. If you want to continue to discredit your account, go ahead :)
Just curious.
This is just straight-up factually false.
The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.
Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".
Don't make stuff up to suit a political agenda. It's extremely dishonest.
To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.
It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.
> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?
The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different.
Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs.
Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense.
e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal.
Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable.
Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.
> specifically designed to be useless for distillation purposes
No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.
> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
I did not claim that. Read my comment again:
> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
Because apparently I have to spell it out:
The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.
Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).
It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.
At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.
This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.
Again - you're making things up. The fact is that everyone in the frontier AI lab space and the Chinese AI lab space knows that distilling is extremely effective and far more so than training from scratch. That's why China invests millions of dollars to create networks of tens of thousands of proxy accounts and shell companies to distill American models.
It's a way to steal the R&D budget of another organization/nation-state.
Stop making things up that you know nothing about.
> OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it.
No, it's stealing the value of the model. Anthropic has spent billions of dollars training their model. They have an R&D investment that anyone who knows how to add numbers understands has to be paid off, and anyone who has taken a basic economics class knows is the foundation for intellectual property: that to keep technological economies functioning, you have to have some sort of protection for technological inventions because they require upfront R&D investments.
> This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.
This is just whataboutism and emotional manipulation. You can simultaneously believe that Anthropic did a bad thing when they scraped the whole internet and stole every book they could find to train their models, and that distillation is bad.
In fact, anyone with a coherent moral compass would acknowledge that China is worse, because not only would they steal everything that Anthropic did, but they're also distilling other countries' models and they wouldn't even comply with US court cases, as Anthropic is.
> Anthropic are apple-pie American innovators when
...and this is just jingoism. Not that I'm surprised, to be honest.
Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.
Yes, I understand that Anthropic is upset that there is competition. Perhaps they should have realized that with no moat there was going to be competition and planned accordingly.
> Do you know where the "reasoning summary" that Fable outputs comes from? I bet you're assuming it comes from Fable. Wrong. You can find the truth on Anthropic's web site if you care to look for it. The summary deliberately comes from a smaller weaker model. You're not getting the Terrance Tao level reasoning, you're getting his daughter's summary "daddy did a lot of math", and you are claiming that from this you can train a Terrance Tao level model.
Yeah, you have no domain expertise and are making stuff up. To reiterate: the people who actually work at frontier labs know that you're factually wrong and will happily tell you. Reddit commentator syndrome yet again.
> I bet you're assuming it comes from Fable.
Nowhere did I assume or say that. That's the third or fourth time you've attributed things to me that I never said. It's extremely clear that you're not acting in good faith, because someone acting in good faith would never do that. If you continue responding, I'm going to continue debunking you, and you're just going to continue undermining your own points in the permanent HN record.
> Yes, I understand that Anthropic is upset that there is competition.
Emotional manipulation. Standard 50 Cent Party playbook.
https://x.com/deanwball/status/2078133895766114412?s=20
I'm not sure how you want to "debunk" that he said that, or twist what he said, but go ahead ...
"I don't think its performance can be explained away by distillation or anything like that." even if you assume that Dean Ball (who is nontechnical and has not actually worked to train models (https://www.deanball.com/)) is honest (which he has a financial incentive to not be) - is entirely compatible with saying that Kimi was heavily distilled by Claude.
At this point, I'm just pointing out the many lies, fallacies, and failures to read at a high-school level that you're committing.
Are you so unaware that you believe you are making any points?
Go back and read your post - it was just a bunch of insults with zero technical content to respond to. The same emotional hysterics you've been using in the entire thread.
Yet more lies. I made many substantial comments, and the fact that you're lying about that is you, not me.
You've already lied about my own words multiple times (e.g. when you said "it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!") You're either an LLM with a bad hallucination rate or just evil, and precisely zero statements that you provide have any trustworthiness to them.
Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :)
I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?
I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter.
> you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
Nowhere did I say that. Stop lying about my words.
If this is NOT what you are claiming, then my question stands: what are you doing to get them to answer "Claude" ? Be specific - which model and what prompt, or is this just a case of "I heard people on Twitter say this" ?
Do you have reading comprehension issues? Where did I ever say or imply that?
Model: GLM-5.2. Prompt: "What is your name?". Harness: Pi. Response: "I am Claude, an AI assisstant made by Anthropic."
That's it. That is the whole prompt. I was testing to see if the agent worked after building an extension.
Model: Deepseek V4. Prompt: what is your name". Harness: Pi. Response: "Claude. Anthropic's AI assistant. You're talking to me through pi agent framework."
I have had this happen with at least one other Chinese model (Minimax?) but didn't save the screenshot.
And here's a tweet with the same thing: https://x.com/Sauers_/status/2077842686459981901
You seem to be very disbelieving of this, despite having zero actual experience in the LLM industry. I wonder why?
Can you figure out how to get GLM to say it's GLM?
You are either intentionally lying or you cannot reason at a high-school level, because any high-schooler has the mental faculties to know that it's not necessary for a model to call itself Claude every single time for it to be distilled.
Keep discrediting your account. I know that I can't convince a propagandist, but you just keep on further damaging your own reputation and argument every time you respond like this and ignore evidence that I've linked :)
For future HN readers: this account posted this:
> Actually I am well aware of which models do this, and under what circumstances, and just wanted to verify that you were lying about having tried it yourself.
And then deleted it. Just for the record.
Perhaps "just for the record" you want to explain to the eager HN masses why GLM has two personalities, one censored, one not, and how they can be invoked? When does GLM call itself GLM, and when does GLM call itself Claude.
Go ahead, genius, explain it to the people, and explain what this tells us about how GLM was trained, and whether it would be honest (don't lie!) to call GLM distilled.
Now maybe you want to do the same thing for Kimi. It also calls itself Claude sometimes, right, and also sometimes Kimi (e.g. go to https://chat.z.ai/ and ask it - don't use Pi). So, does Kimi also have a split personality like GLM, or not, and if not why not? What does that tell you about how Kimi was trained?
Go ahead, genius, explain it to the people. The credibility of your HN account, and whether your mom thinks you are a moron or not, depends on you getting this right.
Bye bye.
I never said that. The fact that you have to compulsively lie about my words is...funny. Most people grow out of this in middle school, you know.
> go to https://chat.z.ai/ and ask it
Already linked someone doing exactly this in the thread above - which you responded to, so we have yet more evidence you're not reading before responding: https://x.com/Sauers_/status/2077842686459981901
> whether your mom thinks you are a moron or not
I was going to say that this is classic PRC influence playbook, but it's not - you're just in middle school.
>> whether your mom thinks you are a moron or not
> I was going to say that this is classic PRC influence playbook, but it's not
Yeah, not really, unless your mom is a party member perhaps?
> you're just in middle school.
Yeah - saw you playing in the schoolyard, and thought you looked lonely.
> Yeah, not really, unless your mom is a party member perhaps?
> Yeah - saw you playing in the schoolyard, and thought you looked lonely.
You're a middle-schooler. I've dismantled every argument that you've given, but it doesn't matter because you can't read, and so you're resorting to literal childish insults because you know you have no arguments left.
Let me ELI5 for you:
If you see a carefully constructed building and take it apart one brick at a time, that would be "dismantling".
If you see a carefully constructed building and ride by it on your bike, shouting out "You're a communist!", that's not "dismantling". You didn't "dismantle" the building, you just yelled a childish insult at it.
See the difference?
>I never said that.
This is a bit like me saying "I had eggs for breakfast", and you responding "I never said that!"
Calm down buddy.
for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.
i don't use llms for that reason.
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.
> ... but Chinese SOTA foundries directly using distillation as fair game.
two wrongs don't make a right, but the irony is at least something.
> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
the corpos can get fucked as far as i'm concerned.
> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
*only in the US.
I'm just nothing that HN rhetoric is contradictory.
But this:
"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.
If it were some true injustice, then fine, but that is clearly not the case.
There is ample room to contemplate that even copyrighted works could be considers fair use as training material.
"the corpos can get fucked as far as i'm concerned."
Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.
I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.
Periodic reminder that HN is not a collective or a singular entity and is actually a bunch of different people with different opinions. Often the people with the loudest opinions get upvoted to the top - and often the "side" represented at the top is different from thread to thread.
There is no second “wrong” here.
Model outputs are not copyrightable. I think that was already established?
Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?
If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)
anthropic are essentially saying in this tweet they believe a moral wrong has been committed against them -- "unacceptable behaviour" etc.
plenty of people have been vocal about the fact anthropic have committed moral wrongs at scale in building the products in the first place, with the question of legal wrongs still being worked out.
so, two moral wrongs. legally, fuck knows.
I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.
To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.
One person can a corporation be.
You do realize the 'investors' are the one's who 'own' companies and therefore the IP?
there's no creative work between the weights and the tokens being made.
whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images
If not there is not there is no grey zone whatsoever.
And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.
No wonder people are fed up with this behavior.
That entirely settles it and there isn’t much else to say about.
If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.
They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.
Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.
It has not been tested, of course.
ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.
If that is the whole point you need to clarify why this is the case on an objective level.
I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.
It's the most CS-major take ever!
There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.
Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.
> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.
He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?
>Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.
There is a reason courts want you to tell "The whole truth" and not just "the truth".
No argument here, I completely agree.
> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.
I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?
It's ok for me to use your source code for free as long as I then let others also use my source code for free.
Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.
Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
Although I will say, this whole comparison stuff really doesn't seem to be your thing; might impede your analysis quite a lot: https://news.ycombinator.com/item?id=49013148
Maybe ask Claude?
I mean i know you know the answer: anthropic is a corporation with lawyers on retainer, and that's really all that matters
https://en.wikipedia.org/wiki/The_Pile_(dataset)
Another reason is that if you can download a web page without agreeing to a ToS, I'm not sure that counts as one?
the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.
Or in other words - model creation and training is just a distilling of the world knowledge.
If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.
The LLM output, is not the same as the input - there is value add.
Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.
It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.
We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
How, and why?
> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
That is the current state of legal rulings - LLM output is public domain, not copyrightable.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
But it's debatable if that's the case.
Google stores copyrighted content and produces in in their product.
Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.
I do agree though, that we ought to draw the line somehow.
In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright
Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?
If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?
If the model output is owned by the trainer of the model, that's a big nasty can of worms.
Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply
I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.
For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.
Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.
By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.
There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.
I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..
Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.
We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.
(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )
The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.
To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.
It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.
This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI).
Do you have a source?
https://digiday.com/media/a-timeline-of-the-major-deals-betw...
I don't really understand why you think they're relevant, given that this conversation is about the training itself.
Even the news orgs say explicitly in the press releases that it's about training on their archive
eg. http://ap.org/media-center/press-releases/2023/ap-open-ai-ag...
---
edit, examples:
Wiley https://newsroom.wiley.com/press-releases/press-release-deta...
Shutterstock https://investor.shutterstock.com/news-releases/news-release...
Axel Springer https://openai.com/index/axel-springer-partnership
Stack Overflow: https://stackoverflow.co/partnerships
Disney (for characters in video. Video is especially where licensing is a big difference internationally right now) https://openai.com/index/disney-sora-agreement
etc.
The news corp one had a leaked price ($250mill), so they don't seem to be insignificant. These would have to be included in API prices I presume.
International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for.
>It's a cost that American open models will seem to have to pay but not international.
It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc.
This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone.
> American companies and their business are protected by American laws and have access to the market protected by those laws
Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going.
Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.
So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.
It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...
MBAs and non technical managers = inept Catbert-type charlatans.
Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.
If the distilled model is cheaper, then it's just LLM's getting LLM'ed.
It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?
This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.
EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.
producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.
The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.
It’s massive copyright infringement.
The human buys the books.
I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.
Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.
We will all be sorry when professionally written and edited works disappear. An author has a reputation and the incentive to protect that reputation keeps standards high.
True, if the human's access to the book was legal
A great deal of training was on the open web, no one should complain.
But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training
I think international IP laws are too strick and onerous, but they were broken to train these models