All written text, art work, etc needs to come imbued with a GPL style license: if you train your model on this, your weights and training code must be published.
All written text, art work, etc needs to come imbued with a GPL style license: if you train your model on this, your weights and training code must be published.
The Tailwind CSS situation is a good example. They built something genuinely useful, adoption exploded, and in the past that would have meant more traffic, more visibility, and more revenue. Now the usage still explodes, but the traffic disappears because people get answers directly from LLMs. The value is clearly there, but the money never reaches the source. That is less a moral problem and more an economic one.
Ideas like GPL-style licensing point at the right tension, but they are hard to apply after the fact. These models were built during a massive spending phase, financed by huge amounts of capital and debt, and they are not even profitable yet. Figuring out royalties on top of that, while the infrastructure is already in place and rolling out at scale, is extremely hard.
That is why this feels like a much bigger governance problem. We have a system that clearly creates value, but no longer distributes it in a sustainable way. I am not sure our policies or institutions are ready to catch up to that reality yet.
The same thing happened (and is still happening) with news media and aggregation/embedding like Google News or Facebook.
I don't know if anyone has found a working solution yet. There have been some laws passed and licensing deals [1]. But they don't really seem to be working out [2].
[1] https://www.cjr.org/the_media_today/canada_australia_platfor...
[2] https://www.abc.net.au/news/2025-04-02/media-bargaining-code...
The biggest issue with the scheme is the fact that it was structured to explicitly favour media incumbents, and is therefore politically unpopular.
This won't help tailwind in this case, but it'll change the answer to "Should I publish this thing free online?" from "No, because a few AI companies are going to exclusively benefit from it" to "Yes, I want to contribute to the corpus of human knowledge."
There are many projects that love to brag about being open source (it's "free"!), only to lock useful features behind a paywall, or do the inevitable license rug pull after other companies start profiting from the freedoms they've provided them. This is the same tactic used by drug dealers to get you hooked on the product.
Instead, the primary incentive to release a project as open source should be the desire to contribute to the corpus of human knowledge. That doesn't mean that you have to abandon any business model around the project, but that shouldn't be your main goal. There are many successful companies built around OSS that balance this correctly.
"AI" tools and services corrupt this intention. They leech off the public good will, and concentrate the data under the control of a single company. This forces well-intentioned actors to abandon open source, since instead of contributing to human knowledge, their work contributes to "AI" companies. I'm frankly not upset when this affects projects who were abusing open source to begin with.
So GP has a point. Forcing "AI" tools, and even more crucially, the data they collect and use, to be free/libre, would restore the incentive for people to want to provide a public good.
The narrative that "AI" will bring world prosperity is a fantasy promoted by the people who will profit the most. The opposite is true: it will concentrate wealth and power in the hands of a few even more than it is today. It will corrupt the last vestiges of digital freedoms we still enjoy today.
I hope we can pass regulation that prevents this from happening, but I'm not holding my breath. These people are already in power, and governments are increasingly in symbiotic relationships with them.
This is on point.
It does not "create value" it harvests value and redirects the proceeds it accrues towards its owners. The business model is a middleman that arbitrages the content by separating it from the delivery.
Software licensing has been broken for 2 decades. That's why free software isn't financially viable for anybody except a tiny minority. It should be. The entire industry has been operating by charity. The rich mega corporations have decided they're not longer going to be charitable.
I think once it becomes infrastructure and widely used knowledge the authors can't claim control anymore. Or shouldn't.
This is a big eye-roll but otherwise ya, this is one way to think of it. It's not all about money, though. The people running these companies are just taking, en masse, without credit. This is a basic human desire. Of course there is a discussion of whether or not we should evolve beyond that. It feels incredibly dystopian to me, though.
Aside from the incentive problem, there is a kind of theft, known as conversion: when you were granted a license under some conditions, and you went beyond them - you kept the car past your rental date, etc. In this case, the documentation is for people to read; AI using it to answer questions is a kind of conversion (no, not fair use). But these license limits are mostly implicit in the assumption that (only) people are reading, or buried in unenforceable site terms of use. So it's a squishy kind of stealing after breaching a squishy kind of contract - too fuzzy to stop incented parties.
100% I get that you are OK with it being used by non-human ingestion. And I think many might be OK with that.
One thing, I'm not sure how helpful the documentation is. I think we're getting training out of example, not docs. This makes me think... we could test this by creating a new pseudo-language, and then provide no examples, only docs.
If the LLM can then code effectively after reading the docs, we'd have a successful test. Otherwise? It's all parroting.
LLMs broke that social contract. Now that product will likely go away.
People can twist themselves into knots about how LLMs create “value” and that makes all of this ok, but the truth is they stole information to generate a new product that generates revenue for themselves at the cost of other people’s work. This is literally theft. This is what copyright law is meant to protect. If LLM manufacturers are making money off someone’s work, they need to compensate people for that work, same as any client or customer.
LLMs are not doing this for the good of society. They themselves are making money off this. And I’m sure if someone comes along with LLM 2.0 and rips them off, they’re going to be screaming to governments and attorneys for protection.
The ironic part of all of this is that LLMs are literally killing the businesses they need to survive. When people stop visiting (and paying) Tailwind, Wikipedia, news sites, weather, and so on, and only use LLMs, those sites and services will die. Heck, there’s even good reason to think LLMs will kill the Internet at large, at least as an information source. Why in the hell would I publish news or a book or events on the Internet if it’s just going to be stolen and illegally republished through an LLM without compensating me for my work? Once this information goes away or is locked behind nothing but paywalls, I hope everyone is ready for the end of the free ride.
Copyright, as practiced in late 20 and this century, is a tool for big corps to extract profits from actual artists, creators, and consumers of this art[0] equally. Starving artists do not actually benefit.
Look at Spotify (owned and squeezed by record labels) giving 70% of the revenue to the record labels, while artists get peanuts. Look at Disney deciding it doesn't need to pay royalties to book writers. Hell, look at Disney's hits from Snow White onwards, and then apply your "LLMs are IP theft" logic to that.
Here's what Cory Doctorow, a book author and critic of AI, has to say about it in [1]:
> So what is the alternative? A lot of artists and their allies think they have an answer: they say we should extend copyright to cover the activities associated with training a model.
> And I'm here to tell you they are wrong: wrong because this would inflict terrible collateral damage on socially beneficial activities, and it would represent a massive expansion of copyright over activities that are currently permitted – for good reason!.
---
> All written text, art work, etc needs to come imbued with a GPL style license
GPL-style license has been long known not to work well for artifacts other than code. That's the whole reason for existence of Creative Commons, GNU Free Documentation License, and others.
[0] "consumers of art" sounds abhorrent, yet that's exactly what we are [1] https://pluralistic.net/2025/12/05/pop-that-bubble/
How about my books?
I’ve published a few, make maybe $500 a month from them.
Is it fine for the LLMs to rip them off?
2. Will people stop buying your books if they can get them from the library? Is a library ripping you off?
3. Assuming your books are non-fiction (otherwise the answer to (1) would be a clear "no"), am I ripping you off if I read your books, create a course that teaches the same things (that you tought me through your book) and earn mega-money because I'm mega-skilled at marketing?
4. How about if I lend my copy to dozens of my friends, who are all very interested in the stuff you write but don't want to pay themselves?
5. Did OpenAI go to the bookstore, buy your book and scan it? Or did Amazon or any other ebook retailer just gave them the book PDF when they asked nicely? How did the rip off happen?
6. If an Anthropic employee buys your book in a bookstore, scans it and destroys the physical copy, and the digital equivalent is only used to train Claude, is that a ripoff?
This stuff is complex and as a society we're just starting to grapple with the consequences. Cory's making the case against copyright being used as a tool much more eloquently than I am - I encourage you to read it if you haven't already.
BTW in your particular case, I'd say you're pretty safe. Nobody stops buying books because they can get the same info from LLMs. If that's your concern, you might as well be mad at the Internet at large.
You've got a convert here. I don't think I'll publish my next book. I might just email it straight to Open AI.
And Cory Doctorow - I've attempted a few of his books. Felt like I was reading young adult fiction. He's pretty much the '2 prescient statements and a few average books' guy.
Yes. It is good (and IMO should be encouraged) that derivative works can are made, even if it would make you less money.
Although getting rid of the copyright/IP nonsense for works that already exist is a big win on its own.
You missed the point entirely.
Nobody will write books when big companies will just rip them off and take the profits.
I didn't miss it, I just don't buy it. Lots of people create and write in their free time, without any financial incentives whatsoever.
I'm okay with writing books for profit, though. But if your business case depends on artificially limiting what users can do with their copies, then your business is fundamentally unethical, whether it's a one-man company or OpenAI. It's not a bad thing for such a business to go under.
We don’t own anything we release to the world.
Also, if you publish your code in your own server, it will be DDoSed to death by the many robots that will try to scrape it simultaneously.
I'd be all for forcing these companies to open source their models. I'm game to hear other proposals. But "just stop contributing to the commons" strikes me as a very negative result here.
We desperately need better legal abstractions for data-about-me and data-I-created so that we can stop using my-data as a one-size-fits-all square peg. Property is just out of place here.
There are no ”commons” in this scenario, there are a few frontier labs owning everything (taking it without attribution) and they have the capability to take it away, or increase prices to a point where it becomes a tool for the rich.
Nobody is doing this for the good of anything, it’s a money grab.
I don't wanna look a gift horse in the mouth here. I'm happy to have benefited from whatever contributions were originally forthcoming and I wouldn't begrudge anybody for no longer going above and beyond and instead reverting to normal behavior.
I just don't get it, it's like you're opposed to people building walls, but you see a particularly large wall which makes you mad, so your response is to go build a wall yourself.
This is why I think permissive licenses are a mistake for most projects. Unlike copyleft licenses, they allow users to take away the freedoms they enjoy from users of derivative works. It's no surprise that dishonest actors take advantage of this for their own gain. This is the paradox of tolerance.
"AI" companies take this a step further, and completely disregard the original license. Whereas copyleft would somewhat be a deterrent for potential abusers, it's not for this new wave of companies. They can hide behind the already loosely defined legal frameworks, and claim that the data is derivative enough, or impossible to trace back, or what have you. It's dishonest at best, and corrupts the last remnants of public good will we still enjoy on the internet.
We need new legal frameworks for this technology, but since that is a glacial process, companies can get rich in the meantime. Especially shovel salespeople.
If I put something on Github with a GPL 3 license, it's supposed to require anyone with access to the binary to also have access to the source code. The concern is, if you think that it is theft, then someone can train an LLM on your GPL code, and then a for-profit corporation can use the code (or any clever algorithms you've come up with) and effectively "launder" your use of GPL code and make money in the process. It basically would be converting your code from Copyleft to Public Domain, which I think a lot of people would have an issue with.
Copyright and copyleft only deal with source code distribution. Your last sentence is not really true from a factual perspective.
I think if you really believe in the open source free software mentality that code should be available to help everyone and improvements to it should also be available and not locked up behind a corporate wall (e.g., a company using GPL code and releasing it with modifications without redistributing the source code), LLMs should be the least of your worries since they don’t do that action. On a literal level they don’t violate GPLv2/v3.
Perhaps copyright law needs new concepts to respond to this change in capability compared to the past, but so far there has been very little legal success with companies and individuals trying to litigate AI companies for copyright violations. Direct violations have been rare and only get more rare over time as training methods evolve.
That said, haven’t part of the complaints about Copilot and the like been specifically because they are reproducing large chunks of code verbatim?
Wait, are you kidding? This is literally a problem we have today with tools like Copilot.
Another point, there is a lot of free and permissive license content to train AI on, where the GPL or copyright can be respected. In many cases, the violating AI companies knew what they were doing was wrong.
The problem here is enforcement.
It's well known that AI companies simply pirated content in order to train their models. No amount of license really helps in that scenario.
Imagine OpenAI is required by law to list their weights on huggingface. The occasional nerd with enough GPUs can now self host.
How does this solve any tangible problems with LLMs regurgitating someone else's work?
I'm not the OP, but here's my from-the-hip answer: if weights are public, building and operating an LLM is no longer a business plan in and of itself, as anyone could operate the same LLM. Therefore companies like OpenAI will be disincentivized from simply redirecting web traffic to their own site.
The AI goldrush has proven that intellectual property laws are null and void. Money is all that matters.
Indeed they never really mattered. They were a tool for large corporations to make money and they will go away if they can no longer serve such purpose. Anyone that thought there was a real moral or ethical basis to "intellectual property" laws fell for propaganda and got scammed as a result.
GPL died. Licenses died.
Exnation: LLMs were trained also on GPL code. The fact that all the previously-paranoid businesses that used to warn SWEs not to touch GPL code with a ten foot pole are now fearlessly embracing LLMs' outputs, means that de facto they consider an LLM their license-washing machine. Courts are going to rubber stamp it because billions of dollars, etc.
a lossy compression algorithm is not "inspired" when it is fed copyrighted input
That's exactly what happens when you read. Copyrighted input fed straight into your brain, a lossy storage and processing machine.
I can be in a room looking at something with my eyeballs and listening with my ears perfectly legally... But it would not be legal if I replaced myself with a humanoid mannequin with a video camera for a head.
But it is automatically copying if you use a copier.
When a LLM is trained on copyright works and regurgitates these works verbatim without consent or compensation, and then sells the result for profit, there is currently no negative impact for the company selling the LLM service.
And I can not copy paste myself to discuss with thousands or millions of users at time.
To me clear solution is to make some large payment to each author of material used in traing per training of model say 10k to 100k range.
This feels like the simplest & best single regulation that can be applied in this industry.
for the burglar
Facebook was busted torrenting all sorts of things in violation of laws/regulations that would lead to my internet being cut off by my ISP. They did it at scale and faced no consequences. Scraping sites, taking down public libraries, torrenting, they just do whatever they want with impunity. You should be angry!
I can see this being extremely limiting in training data, as only "compatible" licensed data would be possible to package together to train each model.
That's part of the point.
The other way is to argue that LLMs democratize access to knowledge. Anyone has access to all ever written by humanity.
Crazy impressive if you ask me.
The first one's free.
After you're hooked, and don't know how to think any more for yourself, and all the primary sources have folded, the deal will be altered.
the mega corps are getting all of us and the investors to fund the RnD.
The thing is, copyright law is not really on your side. Viewing copyrighted material without paying for it is not generally something people get fined for. A lot of training falls under fair use that overrides whatever license you come up with. Disney can’t stop me from uploading clips of their movies alongside commentary and review because fair use allows that. LLMs generally aren’t redistributing code, which is the thing that copyright protects.
If I inspect some GPL code and get inspired by it and write something similar, the GPL license doesn’t apply to me.
It has always been the case that if you don’t want other people to apply fair use to your works, your only recourse is to keep those works private. I suspect that now individuals and companies that don’t want their code to be trained on will simply keep the code private.
Now, there have been times where LLMs have reproduced verbatim copyright material. The NYTimes sued OpenAI over this issue. I believe they’ve settled and come up with a licensing scheme unless I’m mixing up my news stories.
Second thing, your issue becomes moot if there exists a model that only trains off of MIT-licensed code, and there is a TON of that code out there.
Third thing, your issue becomes moot if users have agreed to submit their code for training, like what the GitHub ToS does for users who don’t change their settings, or if giant companies with giant code bases just use their own code to train LLMs.
Where I agree with you is that perhaps copyright law should evolve. Still, I think there’s a practical “cat is out of the bag” issue.
Not all, but some kind of IP.
Some of those that is created for sake of creating it and nothing else.
It's come up quite often even before AI when people released things under significantly looser licenses than they really intended and imagined them being used.
humans learning : machines learning == whale swimming : submarine swimming
It's not the exact 100% same thing. Therefore you cannot base any rights on it.
If you still don't buy it, consider this analogy:
killing a human vs. destroying a machine
Thank god that we're not using your line of thinking here.
You don't need to redistribute the original material, it's enough that you just copied it.
This is what should be fixed in the first place, then. You shouldn't get sued from what you do with your copy of a song.
The moment Tailwind becomes a for-profit, commercial business, they have to duke it out just like anyone else. If the thing you sell is not defensible, it means you have a brittle business model. If I’m allowed to take Tailwind, the open source project, and build something commercial around it, I don’t see why OpenAI or Anthropic cannot.
It's fine to have a project that generates html-css as long as the users can find the docs for the dependencies, but when you take away the docs and stop giving real credit to the creators it starts feeling more like plagiarism and that is what's costing tailwind here.
Thats a big difference. Freelancers using it can bring compounding value as they can sell templates to every client
AI using it bring no value as the AI won't recommend buying the templates
The problem is that LLMs are better than people at this stuff. They can read a huge quantity of publicly available information and organize it into a form where the LLM can do things with it. That's what education does, more slowly and at greater expense.
The article does provide a hint: "Operate". One needs to get paid for what LLMs cannot do. A good example is Laravel. They built services like Forge, Cloud, Nightwatch around open source.
Yes, this betrayed the entire concept of the us as a pro-human market.
Don an experiment, memorize a popular small poem, then publish it under your name (though I suggest to check the laws in your rtegion for this and also consider it might affect your reputation).
IMO is the same if ChatGPT memorizes my poem and then you ask it for a poem , you copy paste my poem from ChatGPT and publish it as your own.
And as a result of this, the models will start consuming their own output for training. This will create new incentives to promote human generated code.
IMO copyright laws should be rewritten to bring copyright inline with the rest of the economy.
Plumbers are not claiming use fees from the pipes they installed a decade ago. Doctor isn't getting paid by a 70 year old for saving the 70 year old when they were in a car accident at age 50.
Why should intellectual property authors be given extreme ownership over behavior then?
In the Constitution Congress is allowed to protect with copyright "for a limited time".
The status quo of life of author + 99 years means works can be copyrighted for many peoples entire lives. In effect unlimited protection.
Why is society on the hook to preserve a political norm that materially benefits so few?
Because the screen tells us the end is nigh! and giant foot will crush us! if we move on from old America. Sad and pathetic acquiescence to propaganda.
My fellow Americans; must we be such unserious people all the time?
This hypernormalized finance engineered, "I am my job! We make line go up here!" culture is a joke.
But I still need to pay rent.
Rent is a political problem.
Perhaps invest in the courage to confront some droopy faced Boomers in Congress.
We should work on fixing that, then.
I agree with your point about big tech companies salivating at opportunities to collect rent. IP is part of the problem.
It is completely against the spirit of information wants to be free. Using that catch phrase in protection of mega corps is a travesty.
The actual problem is political. Has nothing to do with LLMs.
C'mon. You know good and well that what is being discussed is the _use_ of LLMs, with the concomitant heavy usage of CPU, storage, and bandwidth that the average user has no hope of matching.
Not the person you're replying to, but I've found that some people do argue against LLMs themselves (as in, the tech, not just the usage). Specially in humanities/arts cycles which seem to have a stronger feeling of panic towards LLMs.
Clarifying which one you're talking about can save a lot of typing/talking some times.
Maybe?
The person I responded to said "LLMs are just a concept, an abstraction."
Were that true, were they simply words in some dusty CS textbook, it's hardly likely that the humanities/arts people you describe would even know about them.
No, it's the fact that these people have seen regurgitated pictures and words that makes it an issue.
A political problem.
Which, yes, is partly a resource/political problem, but there are additional arguments against the use of those resources for LLMs.
Don't let politics allocate resources to massive data center projects
Have the biggest models be legally forced to be released in the open for end users, then. Best of both worlds.
Wait a few years, and you'll even be able to run those models in commodity hardware.
Enshittification in order to give returns to shareholders suck. The tech is great and empowering for the commons.
Information has zero desires.
> It's wild to me seeing the tech world veer into hyper-capitalism and IP protectionism.
Really? Where have you been the last 50 years?
> Plumbers are not claiming use fees from the pipes they installed a decade ago.
Plumbers also don't discount the job on the hopes they can sell more, or just go around installing random pipes in random locations hoping they can convince someone to pay them.
> Why should intellectual property authors be given extreme ownership over behavior then?
The position that cultural artifacts should enter into the commons sooner rather than later is not unreasonable by any means, but most software is not cultural, requires heavy maintenance for the duration of its life, and still is well past obsolescence, gone and forgotten, well before the time frame you are discussing.
You say that like it's a bad thing...
"LLM Inference Compensation MIT License (LLM-ICMIT)"
A license that is MIT compatible but requires LLM providers to pay after inference, but only restricts online providers, not self-hosted models
Yeah we’ve got a legal system for it, but it always has been and always will be silly.
Also, that Botox patent should be expiring by now, shouldn’t it?
IP is the final delusion of 19th century thinking. It was crushed when we could synthesize anything, at little cost, little effort. Turns out, the hard work had to be done once, and we could automate to infinity forever.
Hold on to 19th century delusions if you wish, the future is accelerating, and you are going to be left behind.
If one wastes their life in court, arguing 19th century myths, that’s on the players.
IP as a concept needs to die.
Yes it does, but if you think that's going to happen on its own by allowing the largest corporations to run roughshod over everyone else, you're going to be disappointed. True freedom takes struggle.