On the other hand, AGPL continues to be the future of F/OSS.
Even the most unscrupulous lawyer is going to look at the MIT license, realize the target can defend it for a trivial amount of money (a single form letter from their lawyer) and move on.
If I can reproduce the entirety of most books off the top of my head and sell that to people as a service, it's a copyright violation. If AI does it, it's fair use.
Pants-on-head idiotic judge.
Is the hinge that the tools can recall a huge portion (not perfectly of course) but usually don't? What seems even more straight forward is the substitute good idea, it seems reasonable to assume people will buy less copies of book X when they start generating books heavily inspired by book X.
But, this is probably just a case of a layman wandering into a complex topic, maybe it's the case that AI has just nestled into the absolute perfect spot in current copyright law, just like other things that seem like they should be illegal now but aren't.
Assuming you're referring to Bartz v. Anthropic, that is explicitly not what the ruling said, in fact it's almost the inverse. The judge said that output from an AI model which is a straight up reproduction of copyrighted material would likely be an explicit violation of copyright. This is on page 12/32 of the judgement[1].
But the vast majority of output from an LLM like Claude is not a word for word reproduction; it's a transformative use of the original work. In fact, the authors bringing the suit didn't even claim that it had reproduced their work. From page 7, "Authors do not allege that any infringing copy of their works was or would ever be provided to users by the Claude service." That's because Anthropic is already explicitly filtering out results that might contain copyrighted material. (I've run into this myself while trying to translate foreign language song lyrics to English. Claude will simply refuse to do this)[2]
[1] https://www.courtlistener.com/docket/69058235/231/bartz-v-an...
[2] https://claude.ai/share/d0586248-8d00-4d50-8e45-f9c5ef09ec81
Now, Anthropic was found to have pirated copyrighted work when they downloaded and trained Claude on the LibGen library. And they will likely pay substantial damages for this. So on those grounds, they're as screwed as the 12 year olds and their parents. The trial to determine damages hasn't happened yet though.
Agreed
> the Sony Betamax case, which found that it was legal and a transformative use of copyrighted material to create a copy of a publicly aired broadcast
Good thing libgen is not publicly aired in broadcast format.
> So on those grounds, they're as screwed as the 12 year olds and their parents.
Except they have deep enough pockets to actually pay the damages for each count of infringement. That's the blood most of us want to see shed.
You cannot have trained the model without possession of copyrighted works. Which we seem to be in agreement on.
I daresay the difference with AI is that pretty much no human can do that well enough to harm the copyright holder, whereas AI can churn it out.
Now there's precedent for future cases where theft of code or any other work of art can be considered fair use.
[0] https://www.cambridge.org/core/books/abs/preaching-the-crusa...
That said, AGPL as a trend was a huge closing of the spigot of free F/OSS code for companies to use and not contribute back to.
It's too late at this point. The damage is done. These companies trained on illegally obtained data and they will never be held accountable for that. The training is done and they got what they needed. So even if they can't train on it in the future, it doesn't matter. They already have those base models.
It’s an EULA trying to pretend it’s a license. You can’t have it both ways.
https://www.gnu.org/licenses/agpl-3.0.en.html
Could you expand on why you think it's nonfree? Also, it's not that hard to comply with either...
https://news.ycombinator.com/item?id=30495647
https://news.ycombinator.com/item?id=30044019
GNU/FSF are the anticapitalist zealots that are pushing this EULA. Just because they approve of it doesn’t make it free software. They are confused.
Free software refers to user freedoms, not developer freedoms.
I don't think the below is right:
> > Notwithstanding any other provision of this License, if you modify the Program, your modified version must prominently offer all users interacting with it remotely through a computer network (if your version supports such interaction) an opportunity to receive the Corresponding Source of your version by providing access to the Corresponding Source from a network server at no charge, through some standard or customary means of facilitating copying of software.
>
> Let's break it down:
>
> > If you modify the Program
>
> That is if you are a developer making changes to the source code (or binary, but let's ignore that option)
>
> > your modified version
>
> The modified source code you have created
>
> > must prominently offer all users interacting with it remotely through a computer network
>
> Must include the mandatory feature of offering all users interacting with it through a computer network (computer network is left undefined and subject to wide interpretation)
I read the AGPL to mean if you modify the program then the users of the program (remotely, through a computer network) must be able to access the source code.
It has yet to be tested, but that seems like the common sense reading for me (which matters, because judges do apply judgement). It just seems like they are trying too hard to do a legal gotcha. I'm not a lawyer so I can't speak to that, but I certainly don't read it the same way.
I don't agree with this interpretation of every-change-is-a-violation either:
> Step 1: Clone the GitHub repo
>
> Step 2: Make a change to the code - oops, license violation! Clause 13! I need to change the source code offer first!
>
> Step 1.5: Change the source code offer to point to your repo
This example seems incorrect -- modifying the code does not automatically make people interact with the program over a network...
"free software" was defined by the GNU/FSF... so I generally default to their definitions. I don't think the license falls afoul of their stated definitions.
That said, they're certainly anti-capitalist zealots, that's kind of their thing. I don't agree with that, but that's besides the point.
And yes, it is an EULA pretending to be a license. I'd put good odds on it being illegal in my country, and it may even be illegal on the US. But it's well aligned with the goals of GNU.
I like open source. I also don't think that is where the magic is anymore.
It was scale for 20 years.
Now it is speed.
A world without open source may have given birth to 2020s AI but probably at a slower pace.
The "good guy" is a competitive environment that would render Meta's AI offerings to be irrelevant right now if it didnt open source.
Don’t let the perfect be the enemy of the good.
I feel like we right now live in that perfect competition environment though. Inference is mostly commoditized, and it’s a race to the bottom for price and latency. I don’t think any of the big providers are making super-normal profit, and are probably discounting inference for access to data/users.
Why would anyone think that, and why do you think everyone thinks that?
And this pattern has repeated itself reliably since the industrial revolution.
Successful ASI would essentially end this process, because after ASI there's nowhere else for humans to go (in tech at least.)
That's a cool smaht phrase but help me understand, for which Meta products are LLMs a complement?
The entire point of Meta owning everything is that it wants as much of your data stream as it can get, so it can then sell more ad products derived from that.
If much of that data begins going off-Meta, because someone else has better LLMs and builds them into products, that's a huge loss to Meta.
>because someone else has better LLMs and builds them into products
If that were true they wouldn't be trying to create the best LLM and give it for free.
(Disclaimer: I don't think Zuck is doing this out of the good of his heart, obv. but I don't see the connection with the complements and whatnot)
If LLM effectiveness is all about the same, then other factors dominate customer choice.
Like which (legacy) platforms have the strongest network effects. (Which Meta would be thrilled about)
LLMs, along with image and video generation models, are generators of very dynamic, engaging and personalised content. If Open AI or anyone else wins a monopoly there it could be terrible for Meta's business. Commoditizing it with Llama, and at the same time building internal capability and a community for their LLMs, was solid strategy from Meta.
There's two products:
A) (Meta) Hey, here are all your family members and friends, you can keep up with them in our apps, message them, see what they're up to, etc...
B) (OpenAI and others) Hey, we generated some artificial friends for you, they will write messages to you everyday, almost like a real human! They also look like this (queue AI generated profile picture). We will post updates on the imaginary adventures we come up with, written by LLMs. We will simulate a whole existence around you, "age" like real humans, we might even get married between us and have imaginary babies. You could attend our virtual generated wedding online, using the latest technology, and you can send us gifts and money to celebrate these significant events.
And, presumably, people will prefer to use B?
MEGA lmao.
It takes content to sell advertisements online. LLMs produce an infinite stream of content.
When any scheme involves some grand long-term goal, I think a far more naive approach to behaviors is much more appropriate in basically all cases. There's a million twists on that old quote that 'no plan survives first contact with the enemy', and with these sort of grand schemes - we're all that enemy. Bring on the malevolent schemers with their benevolent means - the world would be a much nicer place than one filled with benevolent schemers with their malevolent means.
That doesn't feel quite right as an explanation. If something fails 10 times, that just makes the means 10x worse. If the ends justify the means then doesn't that still fit into Machiavellian principles? Isn't the complaint closet to "sometimes the ends don't justify the means"?
It's extremely difficult to think of any real achievements sustained on the back of Machiavellianism, but one can list essentially endless entities whose downfall was brought on precisely by such.
author is "board certified in clinical child and adolescent psychology, and serves as the John Van Seters Distinguished Professor of Psychology and Neuroscience, and the Director of Clinical Psychology at the University of North Carolina at Chapel Hill" and the book is based on evidence
edit: you can't take a book from 1600 and a few alive assholes with power and conclude that. there's a bunch of philanthropists and other people around
Same goes for when Microsoft went gaga for open source and demanded brownie points for pretending to turn over a new leaf.
Considering the rest of your comment it's not clear to me if "anthropomorphizing" really captures the meaning you intended, but regardless, I love this
Its very possible that China is open sourcing LLMs because its currently in their best interest to do so, not because of some moral or principled stance.
I want open source AI i can run myself without any creepy surveillance capitalist or state agency using it to slurp up my data.
Chinese companies are giving me that - I don't really care about what their grand plan is. Grand plans have a habit of not working out, but open source software is open source software nonetheless.
What are you running?
> Chinese companies are giving me that
I have not become aware of anything other than DeepSeek. Can you recommend a few others that are worth looking into?
If that is true and the software is any good, you should be able to name an open-source project that we've heard of started by people living in China.
DeepSeek released some models as open weights and some software for running the models. That's the only example I can think of.
> On 18 May 2022, Gitee announced all code will be manually reviewed before public availability.[4][5] Gitee did not specify a reason for the change, though there was widespread speculation it was ordered by the Chinese government amid increasing online censorship in China.[4][6]
I have a feeling that their collaborative hacker culture is more hardware oriented, which would be a natural extension from the tech zones where 500 companies are within a few miles of each other and engineers are rapidly popping in and out and prototyping parts sometimes within a day.
Anecdotally, I've dealt with Chinese collaborative community projects in the ThinkPad space, where they have come together to design custom motherboards to modernize old ThinkPads. Of course there was a lot of software work as well when it comes to BIOS code, Thunderbolt, etc. I remember thinking how watching that project develop was like peering into another world with a parallel hacker culture that just developed... differently.
Oh there's also a Chinese project that's going to modernize old Blackberries with 5G internals. Cool stuff!
Yea exactly, there is also a lot of chinese people out there, statistically a large chunk are cool with it.
Same dynamic as the US can be really - other countries see the US government and think to themselves, "I don't like these US people, look at what their government did" meanwhile US people are like "what do you mean, I don't like what the government did either". That's what a lot of Chinese people are thinking (but now allowed to say, in China criticizing the government is against their community guidelines)
Highly paid software engineers working in a ZIRP economy with skyrocketing compensation packages were absolutely willing to play this game, because "open source" in that context often is/was a resume or portfolio building tool and companies were willing to pay some % of open source developers in order to lubricate the wheels of commerce.
That, I think, is going to change.
Free software, which I interpret as copyleft, is absolutely antithetical to them, and reviled precisely because it gets in the way of getting work for free/cheap and often gets in the way of making money.
And is building on top of the unpaid labour of SW engineers really a major part of the open source ecosystem? I feel open source is more a way for companies to cooperate in building shared software with less duplication of costs.
Meta has open sourced all of their offerings purely to try to commoditize the industry to the greatest extent possible, hoping to avoid their competitors getting a leg up. There is zero altruism or good intentions.
If Meta had actually competitive AI offering, there is zero chance they would be releasing any of it.
China has stopped releasing frontier models, and Meta doesn't release anything that isn't in the llama family.
- Hunyuan Image 2.0 (200 millisecond flux) is not released
- Hunyuan 3D 2.5, the top performing 3D model and an order of magnitude improvement over 2.1, is not released
- Seedream Video, which outperforms Google Veo 3 on ELO rankings, is not released
- Qwen VLo, an instructive autoregressive model, is not released
The list is much larger than this.
- Hunyuan Image 2.0 (200 millisecond flux) is not released
- Hunyuan 3D 2.5, the top performing 3D model and an order of magnitude improvement over 2.1, is not released
- Seedream Video, which outperforms Google Veo 3 on ELO rankings, is not released
- Qwen VLo, an instructive autoregressive model, is not released
But yeah by analogy with the US, it’s not as if the W. Bush administration can be credited with the creation of Google.
* Do not use emotional reinforcement (e.g., "Excellent," "Perfect," "Unfortunately").
* Do not use metaphors or hyperbole (e.g., "smoking gun," "major turning point").
* Do not express confidence or certainty in potential solutions.
into the instructions, so that it doesn't treat you like a child, teenager or narcissistic individual who is craving for flattery, can really affect the mood and way of thinking of an individual, those Chinese models might as well have baked in something similar but targeted at reducing the productivity of certain individuals or weakening their beliefs in western culture.I am not saying they are doing that, but they could be doing it sometime down the road without us noticing.
If you can make that algebra add up to "bad guy" then be my guest.
It's like telling an iPhone user that iCloud isn't trustworthy because of the Foxconn suicide nets. It's basically the definition of a non-sequitur.
> The problem is that people don’t realize that if we license one single book, we won’t be able to lean into fair use strategy.
[0] https://www.theatlantic.com/technology/archive/2025/03/libge...
You imply there are some good guys.
What company?
Twitter circa 2012?
In 2025? Nobody, I don't think. Even Mozilla is turning into the bad guys these days.
Obv
Kagi, on the other hand, has released none of their technology publicly, meaning they have full power to boil the frog, with no actual assurance that their technology will be useful regardless of their future actions.
For instance, of all companies I've interviewed with or have friends working at that developed tech, some companies build and sell furnitures. Some are your electricity provider or transporter. Some are building inventory management systems for hospitals and drug stores. Some develop a content management system for medical dictionnary. The list is long.
The overwhelming majority of companies are pretty harmless and ethically mundane. They may still get involved in bad practice, but that's not inherent to their business. The hot tech companies may be paying more (blood money if you ask me), but you have other options.
But, can't think of one off hand. Maybe Toys-R-Us? Ooops gone. Radio Shack? Ooops, also gone.
On the scale of Bad/Profit, Nice dies out.
There is no good or open AI company of scale yet, and there may never be.
A few that contribute to the commons are Deep Seek and Black Forest Labs. But they don't have the same breadth and budget as the hyperscalers.
I know because I wanted to, as a form of protest/performance art, train a model to a few Disney movies and publicly distribute, but legal advice was this would put me directly into hot water not just because of who im pissing off (which i knew and was comfortable with) but also the fact there was precedent (i.e. news papers suing LLM providers).
It would be an open and shut case that would leave me in financial ruin.
The reason openAI hasn't been struck with this yet is, who has the time? and there isn't much to learn from all that either. Most open source tooling out competes openAI's offering as is, so the community wouldn't really win beyond punishing someone.
Whereas meta suing you into radioactive rubble is straightforward.
Deepseek, Baidu.
When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."