I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Coming out would be a bold move for them.
(not legal advice!)
One wonders if they're still doing it.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
Anna's Archive didn't exist in 2019.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.
Citation?
* We should not sell powerful chips or chipmaking equipment to China
* We should crack down on industrial-scale distillation operations
That's not quite a ban of local ML inference, but it basically says that he doesn't want companies in China to create their own state of the art models.
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
PlayStation recently sent out emails regarding this, effectively reminding people that they did not, in fact, own what they purchased[0]. It's very confusing - someone here would have a much higher chance of knowing that it's not a direct purchase or right to ownership, though I'll bet most of the general public do not.
[0]: https://windowsreport.com/sony-emails-playstation-users-to-r...
I will say, however, that they've gone up against some very scary people with very large teams of lawyers, and for that I salute them.
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.