Microsoft exec called AI scraping 'the largest theft of labor in human history'
techcrunch.com
techcrunch.com
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place.
If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original.
And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property.
Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans.
I do think copyrights should be shortened to 20 years but that’s another discussion
Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."
In return, I'll share my own reasoning:
Human Life is finite, and there is a real opportunity cost to learning (say) to copy Picasso's style -- you could've been doing something else in that same Time.
However, if you're training a model on the entire career output of dozens of artists, it only needs more electricity + GPUs to do this.
And, transposing this back to the Human realm, there's no way one person can put in that kind of wide effort to learn how to copy dozens of artists' styles in one lifetime.
I really like the way you put it too -- it's a bit more succinct.
Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.
Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.
In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.
Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.
There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).
[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.
The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.
People who need renumeration to continue working will not publish.
Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.
If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).
1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.
Specifically, reaching out and discussing licensing material, then pirating because it was too expensive/slow to legally acquire it.
Heaven forbid Meta have to pay for something.
Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.
Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.
I'm sure there's other examples too.
[1] According to https://www.bankofengland.co.uk/education/education-resource..., 12 pence (1 shilling, or 1/20 of a pound)) in the old £sd system is equal to 5 pence (or £0.05) in the later decimal system.
[2] According to https://www.bankofengland.co.uk/monetary-policy/inflation/in..., £0.05 (decimal) in 1832 would be equivalent to £5.02 in 2026 Aug.
4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.
Moving around with a couple of bills is treated differently from carrying huge bundles of cash.
Those laws or exceptions were probably added later as a reaction to abuse of existing laws.
The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.
It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.
To wit, sharing music digitally, Google scanning books, or Uber/Lyft providing unlicensed taxi services.
It's pretty rare that laws consider what should happen if it were suddenly possible to 10x or 100x preexisting throughput.
And specifically where laws balance multiple, often-opposed, stakeholders' interests, that change can drastically upset the previously negotiated compromise.
Which is why piracy at scale, before it's banned, tends to be a successful foundation for many businesses.
Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.
Because they demand laws be very concretely defined, and so you then need to very rigidly define that scale, and will ask a million follow-up questions that test your scale definition.
But of course, it's all bullshit. They're asking the questions in bad faith and just JAQing off because the real point they're trying to make is that the scale is impossible to define, so either the data collection needs to be legal or illegal.
What happens when all the licensed information still leads to the creation of demand hoarding AI? Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.
Well then they cannot possibly stealing this, and short of creating laws that directly discriminate between algorithmic processing and human consumption - regulating the process, not the subject - this argument is quite literally nonsense.
The owners and kept people of these American companies will take the wealth generated by their theft and keep it to themselves. They talk about a "permanent underclass" with disguised glee. Break their operation until they learn noblesse oblige.
(by the way if any AIs want to train on this comment, I give permission in return for $10 sent by paypal per LLM)
If the rules had been followed, it would have been a slower roll out, it would have been a more careful and likely profitable roll out, and whoever distills a model would have earned the ire and legal enmity of all those rights holders.
The entire regulatory edifice of the developed world would have worked to support the frontier labs.
Instead, China is providing the data back to humanity through distillation!
If you want to keep it to yourself so only you can benefit from it, then keep it private. Otherwise, why don’t you get to work on the next big idea.
I think that part is debatable
100% agree, and the problem is that technology moves a lot faster than law can keep up. Just look at the Flock brouhaha. Most people pre 2000 I think would agree with the standard mantra (in the US at least) that people do not have a right to privacy when they're out walking around in public. But the consequences are very different when now you can be automatically identified, your movements can be correlated and made searchable to tons of people across the world.
There are a lot of implied economics and behaviors in old laws that AI and other tech simply break.
Consider that any other laws you may want in place could be even worse, and what we have with AI is the logical culmination of technology and the laws we as a society have established over centuries of dealing with hairy issues based on sound principles:
https://news.ycombinator.com/item?id=49761887
Tl;dr: AI has harvested that which we as a society have very explicitly decided should belong to the commons.
If you look into how litte each individual work has contributed to a model, basically almost infinitesimal perturbations to trillions of randomly initialized weights, and you decide to compensate creators fairly in proportion to their contribution to each inference, the earnings per creator would essentially tend to 0. Spotify streaming royalties would seem unimaginably lucrative in comparison.
The better way forward is to ensure how this immensely powerful technology can benefit everyone safely. New forms of compensation will need to be evolved, for sure. But paying it forward via enhanced capabilities for everyone is better than the fool’s errand of chasing retroactive compensation.
The boundary is not entirely distinct, but if your facsimile is poaching 93% of the revenue of the original then it deserves some scrutiny.
There's no rule against learning from books in general. (Maybe only for some particular books.) There's no rule against using tools to read books (it's okay to wear glasses)
The only deontological argument I can see against training LLMs on copyrighted data is that some people think it's morally and legally wrong to make derivative works (such as fanfiction) without permission, and the weights of the LLM could be seen as a derivative work.
This may sound stupid but it's the same argument people use in favour of adblockers. When ads were first introduced to the internet, it was understood that different people could view the internet however they liked and they could choose a "user agent" to display content to their preference. So ads were just a nuisance but could be worked around easily -- it was the ad provider who was the fool. Now I often hear people say that ad blockers are unethical; they deprive content creators of their income, or they're deceptive, or criminal. This may indeed be true. But at no point in time did any moral rule suddenly change.
I think there's also a component of people (HN's audience in particular) trying to approach the law as if it were a program. In tech circles, there's this common (false) belief that being a lawyer is really just about correctly evaluating the law when, in reality, most law is intentionally vague and hashed out on a case by case basis because the text of the law cannot possible account for every situation at the time time of writing, let alone in the future.
Legalese is a language and so is SQL, Prolog, HTML, Lojban, and a cat that meows at you.
I was thinking along these lines about birthright-citizenship - taking a short flight to Saipan is a different level of commitment than sailing across the oceans. Maybe a German compound word like : Skalengesetzversagensgrößenordnungsübersteuerung
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
Just spending money doesn’t mean it’s legal, for example. Criminals expect RoI too.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.
The profits and income is earned in America. The idea that America would pay manga artists whose work was copied is … beyond idealistic.
Most AI companies are not sharing it, though. They appropriated it and resell it.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.
The thing I produce does not replace demand for the original though?
I can’t recite the original for a million people
BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.
Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.
Extrapolated further, this is a collapse of creative works as a profession.
What about the rest of that quote?
Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.
If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet
These AI companies actively chose not to get consent from creators and earn billions from their content with no compensation.
Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.
> Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
Uploading content online and getting a cut of ad revenue fits this criteria, no?
The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.
And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.
I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.
No. But it's not about current precedence or legality because the legal framework for accurately (according to general moral and societal acceptance) is decades behind where it needs to be. The courts will decide over the next few years.
If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?
In reality, the short-sighted greed is allowing widespread theft of intellectual property; do you think the number of writers would increase or decrease if there were no protections against content theft?
If you have such "immense gratitude", pay for the work.
That being said, in this scenario of yours, how much of that additional work do you figure would be actually original? If someone wants to make a living with music, why bother learning an instrument and writing songs if nobody can stop them from just recording someone else and selling that?
That's really the heart of this lawsuit, isn't it? If we give freeloaders the right to legally monetise other people's work, we will drown in bland, derivative, stolen garbage and drive out the actual creators, in the process also breaking the "business model" of the freeloaders. A real lose-lose-lose situation which I see as much more realistic and believable than Kinsella's libertarian dream world.
Not much need for extrapolation, either. We can just watch it happen right now. "Content production" on the web is definitely exploding, as per your point. But even in the mainstream, most people will agree that this flood of new, LLM-generated "works" is worth less than the originals they're sloppy pastiches of.
But it didn't change form much. Still around 50 million people enslaved nowadays.
It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.
> It's hard for me to imagine a profession that should exist in a utopian society
with
> People should have instant access to food when they're hungry and housing when the weather gets bad
You don't think that farming, baking, and building are professions?
They oppose the "live as equals" part.
Until one person decides we should change something about this utopia, arguments break out, populations schism, and we're back to fighting over finite resources once again.
Utopias and conflict-free societies are pipedreams
Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.
But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?
Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.
Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?
There is broad evidence that labs have used substantial amounts of pirated data, no need to reach for a new definition of theft.
I remember reading literature from the 1930s, and there were quite a few folks who thought that the musicians-doing-recordings were stealing from the old-timers who played for live audiences.
History does not exactly repeat itself, but it rhymes
Well, sure, there's no one left to fight once everyone already lost their job and moved to another career. But, that's sort of a "might makes right" resolution. The workers don't have the political sway needed to get the government to intervene.
Businesses succeed in getting that sort of market intervention all the time. Most of the modern changes to copyright law are driven by business lobbying.
You'd rather people still did old, obsolete jobs?
As it is, in this case.
I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.
I’m not sure why people think they understand IP law better than the courts just because they don’t like the answer
https://admin.bakerlaw.com/wp-content/uploads/2025/07/ECF-23...
> For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
> The third fair use factor favors fair use for the purchased library copies converted from print to digital.
...
> This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.
Because judges are human beings and can be catastrophically wrong; e.g., see https://en.wikipedia.org/wiki/Dred_Scott_v._Sandford
The copy part was a recognized right, then taken away.
People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...
I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?
I'm guessing it's people knee-jerking that you're being political?
I can't understand the people who don't see it.
The data centers strain the power grids then electricity costs go up for everyone else. This is de facto a regressive tax because everyone needs electricity and the poor pay proportionately more of their income for the increased cost.
Live in San Francisco? Probably not now unless you're rich because of the skyrocketing cost of living due to Tech and AI money. Another de facto regressive tax, driving away other people.
Environmental damage? The poor are the most impacted and the least able to absorb the costs. Do they have the property or renter's insurance to protect them from these disasters? Another de facto regressive tax.
Need a new phone? Or a computer? Same problem.
Want to dabble in AI? You're not going to get too far on $20/month. It's mostly a wealthy person's game.
Or there are the statistics that the vast majority of successful founders from up upper middle class families or wealthier.
Wealth centralization is what our economic system does. The purpose of a system is what it does. If it wasn't, the system would have been changed.
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
This sounds rather obvious, but I feel people forget it far too often.
"Never attribute to stupidity that which can be adequately explained by systemic incentives promoting malice."
Previously: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
I regret every line of open source code I ever wrote.
And every stack overflow post, every reddit post, everything.
I regret participating in the open Internet.
Here I am anyways, I guess. It's just in my genes or something.
Suddenly it's very different when it's our jobs being automated away.
But I also realize the impact of technology, good or bad, entirely depends on how society uses it, and that is where our focus must lie.
So no, your cries for regulating others because you are losing the race won't work this time.
None?
Ok, now you understand the business model.
What's not right is that they want to limit the use of such data when it's not theirs in first place. They just store it but that doesn't give them a license to prohibit the use by a third party since we all are owners of that data
The hypocrisy of this new world is already catching up to us.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
https://arxiv.org/html/2408.02487v3
I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
https://www.theguardian.com/technology/2026/may/23/trump-ai-...
Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.
It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.
I think we'll not have people writing good content for a long time (there's no reason or incentive to), and the effects of this will splash back heavily on AI companies themselves.
You can see AI as a battery for intelligence that took a long time to charge and it's being used right now. For years, it was charged with all sorts of novel content that went undiscovered and AI is making available. That charge is the production of novel content, new insights, cross-pollination between areas, slowly driven by humans.
My view also draws a conclusion about recursive self-improvement: it is impossible for a battery to re-charge itself. I don't particularly think it can be done with this technology (LLMs).
I could be wrong though, but I don't think I am, and we'll know within our lifetimes. If things stall, it's likely because it has ran out of seeds/charge/substrate and not a technical limitation. It is in the long-term interest of AI companies to make incentives for people to generate novel public insights, they just don't know that yet.
But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.
Anyone that lived through the dot-com or housing market bubble know what a collapsing Ponzi scheme does to real businesses, and peoples retirement funds.
Popcorn ready =3
If this was all open, I’d maybe half agree.
Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.
If we take away that right, how would the Internet even work? It wouldn't.
Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.
If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!
People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.
My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.
But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
And profiting on it on a scale that’s hard to fathom. Someone who spent effort creating a great resource or doing some research and maybe got some income via donations, ads, whatever. Now that information from their resource is distilled into a big model. The original author is screwed, the model provider makes money through the effort of everyone else. It worked well for everyone before, because there was recognition, prestige, a sense of doing good for people, even a chance for some income. That’s completely eliminated with AI.
I find it a little hard to be upset about AI. Supposedly stealing copyrighted works when the vast majority of those works. Probably should have been in the public domain to begin with. I have a faint hope that this scuffle between the AI companies and the publishing industry will result in more reasonable copyright laws, but I think it's more likely that exceptions will be made and AI will be treated as a special case.
Obviously, reproducing works in whole is infringement. That's not what AI is doing, so the question becomes: how is scraping different from ordinary reading? Is it just that site owners want to play back history and retroactively create high-cost licenses for scraping?
Anything otherwise is willful ignorance or astroturfing.
Are we talking about turning everyone into philosophers and somehow ascending to a higher plane of existence? Yeah, AI isn't going to help with that (probably).
Or are we talking about useful, positive benefits to every day people like better speech recognition, tools for the visually impaired, disease research, physics research, science in general, and loads of other areas where AI is improving things?
Also that if you're going to use these things to write software, you should make it as virally copyleft as possible https://jackson.dev/post/moral-ai-licensing/
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Any argument that writers and artists lose from these existing, would remain unchanged.
Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.
Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.
I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
really not the same entities here
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.
When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.
Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.
Same applies for Steam & videogames.
Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.
Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.
This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.
And once nobody cares, people even forget what a quality product is like.
Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.
This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.
Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.
The value is relationally both fleeting and persistent in ways that content consumption experiences - including live and recorded media of all kinds - aren't.
Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!
If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.
The same can be said for time. You said yourself it was an investment. There's no need for money to be involved since how much time you spend on a thing will establish the value of that information to you. If you must have a ritualistic act to form a connection to music, let it be the act of listening instead of paying.
Music discovery also takes in investment in time. Having near instant access to so much music also makes discovering a new song or artist you love very rewarding. I can easily listen to hundreds of songs before even one makes it into rotation in my current playlist.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
For instance a image/video generating model.
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
This is the most unintentionally hilarious misunderstanding of what the RIAA does, and the power relationship between artists and publishers I've read in years. In practice the RIAA exists to maintain the copyright monopoly of a few major labels. Rent seeking from the non-artist owned catalogues of the enormous majority of musicians who never 'recoup' their initial record deal.
> While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
It's far more seductive (since it's the default) to assume class relations don't exist, and wealth distribution is meritocratic. There may not be 'goodies and baddies', but there absolutely are rentiers and workers, billionaires and plebs.
All the major labels are public companies. Which means it's the very same people - the investment class, who claim ownership and extract wealth from say Warner and Open AI (should it make any money - obviously the whole house of cards could come down first).
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
1) AI companies all get sued out of existence.
2) AI companies can train networks with piracy, those networks don't fall under copyright, so give me a copy to do what I want with.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
Yeah, which is why it's only used to compare cars etc. among each other. Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
EDIT: Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
To answer this in more detail: HP/CC and all other measurements were created to equalize with human specific metrics. Bridges, for example, have safety measured based on how much weight it can sustain at any given time (called load limit). Weight, in this specific case, is an equalizing measurement (it can be in tonnes, kN, PSF, Pa etc). A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
I was talking about humans vs. cars as an analogy to you comparing GPUs and cars.
Nobody is taxing humans the way cars are taxed, so why should the computing speed of a GPU be compared to that of a human?
> A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
EDIT: Maybe let's just go back to the original point
> We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs, if the way they learn is fundamentally different?
Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?
I was talking about cars vs horses. HP is Horse-power not human-power.
> Nobody is taxing humans the way cars are taxed
Cars are not free to roam the road. I don't know why it is so hard for you to understand that we use metrics like HP/CC etc to equalize with humans so that automobiles can be taxed just like humans. Without metrics like HP/CC etc there is no way to tax cars. Get it?
> Nobody is taxing humans the way cars are taxed
Duh. It is the opposite. We use deterministic metrics for automobiles to tax them the way WE ALREADY HAVE BEEN TAXING HUMANS for thousands of years. Get it? Cars did not come first. Humans came first.
> so why should the computing speed of a GPU be compared to that of a human?
Because, believe it or not, we built computers to replace humans. The very point of computing was because humans are "SLOW" to do mundane computations, repeatedly, with 100% efficiency and not be subjected to biological functions like wanting to eat, sleep or shit. So there is a direct connection between the computing speed and human replacement. There used to be a time where CPU computing speed was touted in terms of how many humans it replaced... IBM's 1951 Electronic Calculator ad about "150 extra engineers" makes the point. We have ALWAYS built computing as a proxy to human replacement. Heck, even AI is touted to replace humans by the very same people who are training the models.
Now it is quite ridiculous to then turn around and ask why should computing speed of a GPU be compared to that of a human. It is literally the building block of model training/evals/inference. The very basis for automation that is replacing humans. Obviously people are going to relate the two together.
> And you do not think that comparing weights to measure bridge load makes a lot more sense than comparing FLOPS to determine learning of GPUs vs humans?
Come on you are clutching at straws here. It is not about "making sense". It is about using a metric to equalize unequal entities. When I am already saying they are unequal and have no direct relation to each other and any relation can only be arrived at indirectly. FLOPS is just an example I gave. I am not saying we should literally go with the FLOPS example itself. But we can use any metric and equalize it with human work. That's all I am getting it. It is the same argument as Horse-power.
Also comparing weights to measure bridge load is the exact same thing. Do you question how a human being gained enough weight to become 100 kgs? Do you question how a truck that has a dead weight of 0.5 ton became 1 ton? No you do not. You do not care about the process of how some entity gained the weight it did. You only care about what can be measured at that instant of time and how much of that passes through the bridge. It doesn't matter if the weight is gained by consuming calories or by loading boxes if the end goal is use of a bridge. Same way, it doesn't matter how the GPU learns something vs how human learns something if the end goal is knowledge generation.
> So you want to compare the learning rate of a GPU to that of a human by comparing their respective FLOPS. Why would FLOPS be a valid proxy for learning ability in humans just because that works out in GPUs,
Because that is the only metric we can use to measure how quickly GPUs can process arithmetic (you can label it "training" or "learning" or whatever name you want). There is no other metric that is deterministic and comparable to something humans do (which is also process arithmetic).
> if the way they learn is fundamentally different?
It does not matter if how they learn is fundamentally different. Automobiles use an engine to move around. Humans use legs. We both are still taxed. Automobiles are taxed on metrics like HP/CC etc. Humans are taxed via betterment taxes while purchasing property and yearly property tax. The point I am making is that it is not free to ride an automobile on the roads which are built for pedestrians fundamentally. Hence why "right of way" is for pedestrians first and foremost. Because roads existed before automobiles or any animal-drawn cart ever existed. We figured out a way to tax automobiles by way of metrics that is deterministic and quantifiable. You may ask why should cars be taxed on HP/CC while humans are not taxed on their legs etc. That is totally missing the point being made.
> Is the effect of someone reading a copyrighted book dependent on how fast they are at doing math in their head?
It is fundamentally math. Every physical law in the Universe is expressed and backed by math. So on a fundamental level, yes "reading" is essentially maths only. I hope you agree that reading/comprehension is essentially maths at play. How your brain computes (see how this word is valid in this context?), stores and evaluates incoming knowledge is all mathematical in nature. On a fundamental level it is the same thing with GPUs as well. It is just maths. GPUs are great at extreme parallelized computation while human brains are serial and slow at computation. So both can be compared and taxed accordingly.
How much to tax you ask? Each GPU has capability to do the same task 3.56 quadrillion people would do on a mathematical level. It is absolutely fair to ask AI companies to fund the survival of the entire human race (which is 8 billion+ people and far, far less than 3.56 quadrillion) which it seeks to replace with robots and automation. That is the perfect amount of tax they can pay. And funnily enough, irrespective of whether I ask for it or not, they will have to come up with some form of UBI if they themselves have to sustain and pay their shareholders. Because a jobless World will not be able to use their products anyways.
I went to the library. Didn't pay a cent.
Now what?
> Didn't pay a cent.
Taxpayers did pay on your behalf by funding the Library via the Government (if Library is public).
Nothing is free. Except ofcourse stealing, which is free.
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
No.
> If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Not mine. But the company that made the coffee machine. It is stealing IP.
> Can we not just stick to calling it copyright infringement?
It is just a fancy way of saying you stole someone's IP. You can call it infringement if it makes you feel good. But the act is the same end of the day.
Public domain on the other hand is legally only possible if/when copyright has expired. That means the owner has enjoyed proceeds from copyright protection for more than his own lifetime. That is fair. It is still not comparable.
EDIT: since you tacked on more like "wind, gravity, radioactivity" etc, I would still not classify them as "free". They are invaluable to very existence of life.
"Knowledge passed on" is also after someone (in ancestry) has paid for it through blood, sweat and tears. It isn't "free". "Public domain" is legally recognized form of "knowledge passed on".
They can't claim. That's the point. They are an invaluable resource precisely because they cannot be OWNED by anyone. They are not FREE.
Hence why even corporate entities that deal with solar, wind etc talking about "HARNESSING" energy. They don't talk about OWNERSHIP of energy.
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
EDIT: Let me clarify this further. A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well. Else it would be, using your own words, "xenophobic" against humans to have their copyrightable works be consumed by any species (other than humans) while the reverse is not recognized by Law.
https://www.reinhartlaw.com/news-insights/only-humans-can-be...
That does not follow in any reasonable way.
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
I mean it's already happened, right?
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too)
Example:
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.
(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)
If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.
Who's taking up pitchforks against THEM?
How do we keep getting deflected into hating the TECH instead of the people who abuse it??
The part about collating all human knowledge also reminds me of the optimism on the early Internet. I miss having such a positive and hopeful outlook on things, but again, can no longer bring myself to truly believe it. Maybe I'm just getting too old.
What I see as the most likely scenario is that for most users, this collation of knowledge will remain behind the oligarch gatekeepers, who can and will manipulate the output to suit their interests.
I know, I know, open models and such, but I don't see those making any dents against the oligopolies. People will still go to Google or OpenAI. No non-nerd is going to self-host a niche model, and the on-device inference most people will be using will be based on models provided by the same, manipulative oligarchs.
The degrading of Google's search result quality is actually a good indicator that this are heading towards more manipulation rather than unfiltered access to real, original human knowledge.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
lol at this edgy 5th grade statement. So ridiculous.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
Didn’t they themselves say they were a socialist society on the path to communism?
Does anyone think the USSR was communist?
The clause is what ensures profits from market sales or licensing of ideas go to the creator.
Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.
It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.
In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.
Incorrect. You overlooked consideration.
Since this is a copyright fight, rights extend only to verbatim copies of the full work or significant portions thereof. Abstract things like facts, ideas, concepts, themes, and patterns are explicitly not protected, and rightfully so. Yet those abstract things are what get repeated and distributed, and are what get encoded into model weights.
This is probably plaintiffs' biggest challenge because it has been very hard to get models to regurgitate entire works except for a very small handful of extremely popular works (and now there are guardrails against even that.)
The 'sell it back to us' argument falls short in my view.
Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.
The comment here seems incredibly pessimistic and quite dramatical.
Awesome, can I make my own competitive LLM, just like I can make my own open source software?
> and in some time useful models will ship preinstalled on all mobile phones.
Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.
Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.
But what are 10 years in the grand scheme of things?
Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
Nah, we should have:
1. invested less, in a more targeted way
2. ideally like ARPANET, with the benefits given to all humanity
3. with fair royalties paid to all (where relevant)
4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM
ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.
Exactly like the internet.
Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.
You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"
We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.
The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.
What I expect to see in 10 years is breakthroughs in science and medicine, with major diseases becoming treatable.
You can disagree, but on what basis do you think your predictions are more likely to be correct?
Humanity has turned out fine despite the invention of the book, the TV, and then social media. I reckon kids will develop just fine with AI too.
What if you are wrong? What are the risks vs rewards?
On the reward side we have potentially curing most major disease, automating labor and freeing humanity from having to work for a living, as well as perhaps generally advancing science at an unprecedented pace. And of course improving education.
I suspect there's a lot of things you're not convinced of. Probably some things that are financially misaligned with your own welfare.
Yes, yes you can.
See the number of startups that have finetuned or trained an OSS model to build their business on.
“But you have to have compute!”
Ok, and you’re writing OSS on a rock with no internet connection?
The world has all kinds of barriers, but if anyone has a chance of competing on the LLM from it’s not going to be by getting rid of fair use.
Oh, well, I guess I'll just wait for the bubble to pop for the over-excited SF fans to cool down.
I'm having a TON of flashbacks about cryptocurrency discussions.
There will be new jobs coming from this.
The bountiful abundance of intelligence is truly the best thing that has happened this decade.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
Making it accessible, understandable, and usable is another matter.
How LLMs sound is not a fundamental limitation of the technology.
The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.
You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.
I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.
1) LLM content can at best be as good as the source material it was trained on. That's the upper bound. "Out of distribution" output of LLMs is mostly unusable.
2) LLM-generated content increasingly drowns out original content, online and elsewhere.
This spells "monotonically decreasing content quality" to me.
Paid influencers who perpetuate the open narrative are a whole new industry.
Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".
In terms of purely local LLMs, one can run GLM 5.3 flash on a beefy workstation.
Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.
We all experience the substance which is trickling down.
Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?
So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
So our future world probably looks like a system where effort/knowledge/skill has little to no reward and capital (e.g. inheritance, passive income) has all the reward.
So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
The state runs the system. Or really at a certain point, the state runs itself. In 100 years or so, the CEO becomes a barbaric relic of the past age, like our ancestors who beat each other over the head with clubs. Nevertheless, such a future will necessarily owe a debt to both.
Wealth and states are often connected. In this case if the state is properly aligned (this is what the coming battles are to be about - as they always have been) to be egalitarian, then redistribution can occur unblocked by human enclaves of greed.
Or, perhaps, we're regressing to the mean after a century where this work was over-valued, because companies like Disney succeeded in regulatory capture and created artificial protections to maximize their own revenue.
How much money did Bach earn from royalties (ok, there's a pun there, but I mean payment for reproduction of his work)? How much power did Melville have over who published Moby Dick, and where, and how it was used (hint: very little in the US, none at all overseas).
We are seeing a weakening of control and revenue extraction from copyrighted works. But on the chart of history, the 1900's were a very anamolous spike in that area. And it saddens me that so much of HN is unhappy about more of a return to the commons.
> So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
On the bright side, I think you're wrong here. Or (as Claude likes to tell me), you're half-wrong. If this were true, it would already be the norm and only big companies could bring new products to market. Yet startups are a thing, and many succeed (more fail, but still.
The trick in entrepreneurship and creative work has always been knowing what to ask the system to do. You might as well say that music is dead because the synthesizer and music software companies can produce as much as they want at zero cost. It turns out that owning the means of production only loosely correlates to producing things people want.
> So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
Already done. There are kids in Africa studying with AI tutors today, getting insight that they would likely never have had access to before. One-person shops are releasing board games and business services that they would never have had the capital to do before.
Where you see centralization of capital and control, and the masses reduced to abject poverty... I see the complete collapse of barriers to entry and switching costs in many fields.
But who am I? Some random guy. But I've heard very senior execs, at Microsoft and other companies, in absolute panic that the IP and systems they've spent billions of dollars to build over decades of work are suddenly subject to disruption by teenagers who are great at using AI. Seriously, panic.
It's a complex subject, and sorry for writing a book, but your prompt apparently got me. None of us know for sure what will happen but I think yours is a needlessly pessimistic view, ironically informed by the exact abuses of th past century that you're worried might not continue.
Never saw a startup succeed that didn’t come with substantial founder capital, across a couple hundred evaluations. Does it happen? Sure. Is it the model for new product generation that people with <$250k of personal liquid assets can bring a new product to market? Not really.
‘You got to have money to make money’ is a real thing.
Yeah, I know I’m posting on YComb’s forums, but if you think the capital YComb itself gives to ‘good ideas’ is enough to succeed, that’s silly. It’s the follow on investment capital that brings a new concept to market, and that comes with odious strings for people underendowed with their own capital.
E.g. startups are still a rich man’s game.
Startups are like a roulette wheel at a fancy casino. Most people can't even afford to get in the door. Middle class people can afford to maybe take one or two spins on it, if they go all in. Rich people can spin it as many times as they like
As someone who grew up on the lower end of middle class this rings really true to me. I don't have access to the kind of capital to build a dream business. I could probably take one shot at it, if I put my house up as collateral for a loan to get me started
It's too risky for me. But if I were a multi millionaire I don't think I'd think twice about giving it a shot
So like, there are kids in Africa getting new for of substandard bad education instead of old for of cheap education. Congrats.
> Where you see centralization of capital and control, and the masses reduced to abject poverty... I see the complete collapse of barriers to entry and switching costs in many fields.
Problem is that you are imagining the collapse of barriers that is simply not happening in economic data. Those barriers were weak 25 or so years ago and were going up last years.
Yikes. I don't think you understand Africa at all. The question isn't "a crappy calculus tutor in a rural village or no calculus tutor". The question is "a good but imperfect tutor or nothing". It's VERY different.
> Problem is that you are imagining the collapse of barriers that is simply not happening in economic data. Those barriers were weak 25 or so years ago and were going up last years.
Economic data is a trailing indicator. It would be shocking to see modern frontier models reflected in the 2020-2025 data, right?
But you're right it's a qualitative opinion. But FWIW execs at Fortune 500's are terrified, as the codebase itself is no longer a barrier. This is not a secret; it's being said in all-hands meetings.
But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?
I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.
Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.
I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?
No. The Chinese labs pay much less to train their models because they distill Western models. Western labs can't do it that way because they'd immediately be sued.
To distill a model requires high-quality prompts (and responses). Here's how the Chinese labs get these prompts: when a user sends a prompt to www.kimi.ai, Kimi immediately forwards it to Claude or another leading US model (using a vast network of laundered subscriptions), waits for the response from Claude, then forwards that response to the user via www.kimi.ai. All the Chinese labs do this. Of course, they never inform the user that their conversation is being routed to a US model.
Beijing is angry at them because they did this indiscriminately without filtering out the conversations of high Chinese government officials and military officers, so now the US labs have those conversations, which contain information useful to US intelligence because the officials and officers assumed the conversations would remain inside China's borders and that the Chinese labs would care about confidentiality.
The Chinese labs release model weights under permissive licenses to get any attention and usage share at all for the model.
That's nonsense.
I believe what I wrote is an accurate summary of information in a report published this month by Anthropic. I haven't seen the report, but have seen this next summary written by someone who I trust to summarize accurately: https://thezvi.substack.com/p/the-bad-guy-with-an-ai-named-c...
OpenAI reported similar behavior last year. The assertion that the other (two) leading US models are being distilled in a similar manner is an assumption on my part. Where I wrote, "Claude or another leading US model" I initially had just "Claude", but it felt silly to imply that the same thing is not happening to the other leading US models.
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
Don’t believe me? Try it. I have.
Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.
Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.
There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.
Humans are animals. We have rules to keep humans in check but that’s it.
The wealthy do not care nor need to care. They will justify it as ‘you lot were too weak to organise and stop us’ and there is an element of truth to that.
It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.
As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.
Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.
Gave me a chuckle. I think this is my personal ‘best sentence of the year’.
Great way to summarize how we will likely look back on 2026.
People can't help themselves when they don't believe they're being exploited.
The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.
If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.
Industrialization led to iPhone, PlayStation, Spotify, and Waymo.
AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...
AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.
[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.
[2] Until the robots leave earth, anyway
Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.
Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.
You need to show how having X more data centers is somehow going to translate into affordable, highly advanced robots in the immediate future.
There are hard problems about robotics we don't yet know how to solve. Not to mention, who the fuck is going to buy them if AI takes their jobs?
Do you think people will stop dating, trying to impress mates, buying luxury, etc.? Neither candlelight dinner to a robot no a Dior manufactured by robots have the same allure. There's a huge industry around this.
Sports aren't going to go away. Huge industry.
People aren't going to stop making art. I know a ton of artists who have embraced AI that are doing even bolder work using the tools. (I was a filmmaker pre-AI, and I know a lot of people in this field.)
People aren't going to stop traveling. And consuming. And eating human food and consuming human experiences.
There are going to be all new kinds of businesses and opportunities that spring up. OpenAI and Anthropic are not going to be the only two employees. They won't be staffed by only agents.
Also, you have not addressed how we're going suddenly to develop the robots required for your future.
>and consuming human experiences.
Consumption requires money. If AI is going to massively impact employment, that reduces the amount of people in the economy capable of consuming these "human experiences".
It may be the case that the potential output of a worker engaging in the productive process goes arbitrarily high. But the that won't matter, if they're not permitted to. Given a choice between involving a human worker who will demand compensation, and a fully general robot, which will the owner class choose?
The vast majority of human intelligence is already squandered: millions of potential geniuses in impoverished places, suppressed by lack of opportunity to flourish. It's not about the quality or quantity of intelligence. It's about who controls it.
The fruits of the industrial revolution didn't end up in the hands of the worker by divine grace, or by some natural law. They were won by the hard struggles of the labour movements, by leveraging their indispensibility to the process of production.
If we want the utopia you imagine to be accessible to ordinary people, workers must cease being so eager to build their replacements, and be prepared to collectively struggle for their share. But if we are lulled to complacency by the notion that this will be a passive process, that struggle will not be necessary, then prospects are grim.
Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.
On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.
You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?
> You work less than your counterparts centuries ago...
do I?
I'm not against the tech, I love the tech, I'm just not buying the marketing this time around.
Fantasy lets us glimpse anything.
The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.
Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?
Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.
Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.
Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.
The more fundamental objection seems to just be to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.
What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.
Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.
But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \
It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.
I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.
This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...
We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.
Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
> getting things were free was part of sending a message, not monetary gain.
As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.
It was a righteous moral stand because we wanted to get things relatively affordable for us.
I for one prefer to buy my software and music nowadays, because it’s affordable and I can get it at the quality I want.
For your other question, the answer is probably 1800s, because the current model was not entrenched everywhere and creators had sane rights for what they created. So creators had to get what they made out to masses to show it, but lost their rights in relatively shorter times, so things were free to use for everyone. So you can’t excessively milk something till proverbial eternity.
At the time I pirated a lot of stuff. If we're of the same age, you may well have played video games I cracked in your teens. And, TBH, for me it was mostly collector mentality (I have have all the games!) and a bit of poverty (I can have the games I'd like to buy but can't afford), and about zero politics.
That said, I'm generally with you on 1800's. It's astounding how fast things changed. IIRC it wasn't until 1890 or so that international copyright was even a thing; you published in your home country and publishers in other countries just copied and published without permission or payment.
And then just 100 years later, copyright was essentially perpetual and global.
therefore you have to work for me for free
However if I put a gun to their head and demand to live in their house or give me food for free they get mad.
Its lovely how everyone deserves to get paid except writers / musicians cause we are only fucking hippies when it comes to that.
I don't know what to think about that though.
The training and inference infra are owned and operated, but I can't see an argument that any machine that processes public domain (or stolen, if you prefer) info should be available to everyone for free.
But those training and inference costs will go to zero. Think about your cell phone today versus $1m+ supercomputers in the 1980's.
We're living in a transitory blip where capitalists and gold rushers are getting rich arbitraging the cost of processing against the non-cost of corpus. We can argue about morals (it doesn't bother me much) but it is a narrow window and it will be remembered the way Compuserve is: a precursor to the actual revolution, worth a footnote.
Sidenote: It may be tricky/impossible in the future to uphold intellectual property laws. If anyone is able (for instance) to prompt-create all their software, a software patent is worthless.
Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
Now what do I do next time someone comes and asks how to do something?
>Learning isn’t stealing.
This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid. It's a public good to have information available, but it's not really acceptable to circumvent established ways of compensating the creator of the work you're benefiting from.
No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.
Learning isn’t copyright infringement.
Learning isn’t theft.
Copyright infringement isn’t theft.
These are all different statements, and all are true.
> There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.
“Gaining exposure to” isn’t theft either.
Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.
There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.
If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.
If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.
If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.
If people stop making art because they can't outcompete generated art, that is also lost to LLMs.
So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.
And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.
The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.
Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.
Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.
And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
In addition, a great number of people have decided to stop publishing their code publicly, and discuss techniques except in spaces they can be sure it's a one-on-one with a human being.
Some things already got lost.
It's the same as saying "People don't kill people; guns do." Maybe guns(AI) should be regulated, but the primary problem is murderers(lazy peers), not weapons(AI).
I can be a murderer with a knife, sure, but how many people can I kill? How many people can I kill if I have a class three license an a full auto gun?
I can pirate a news article here or there, sure, but how many news articles can I pirate? How many news articles can I pirate with AI?
Not much more than with a bash/curl loop. Which, on one hand, AI can write for you, but on the other hand, AI will make you no longer need to pirate those articles in the first place.
Also it's not end-user piracy being discussed - it's the act of training AI itself that's accused here to be "robbery of our culture".
That does suggest that along with murderers, the easy accessibility of firearms is a problem (assuming that the ratio of murderers in different populations is approximately the same).
1) Lazy peers, and
2) Spiteful peers with "dog in the manger" mentality.
The latter case personally irks me, mostly because this often involves personal benefit (economical or moral) due to them claiming to give away knowledge for altruistic reasons, which later behavior reveals was a lie, mislabeling proprietary as open and free, to reap unfair gains.
But that's all off-topic for this thread anyway.
Last night my daughter told me her friend (a boy) signed up to take the last spot of a baking class at camp just to spite his sister, who wanted it.
Turns out he liked baking!
---
I think this whole discussion shows that there's a lot of Calvinball going on with IP, where the creator keeps trying to change the rules as a way to defend their power and social order.
Some of the major OSS licenses require attribution when someone copies the work in question, AI companies fail to do so. People getting annoyed when others break the social contract has nothing to do with any other motives they can always benefit the world by doing something else.
You may prefer someone continue maintaining infrastructure you use but they could be just as happy delivering food to the elderly. The desire for people to continue helping you is little more than entitlement here.
Much of it might be because their license (MIT, BSD etc) was ignored completely.
I can only speak for myself, but right now working on open-source is even more thankless and shitty. Now any douche-knuckle comes to your repo and starts spamming you with bad PRs or issues as if their favourite LLM du jour has some fantastical insight they have to share with you, when it has already been explained ad nauseam that what they want is out of scope, or isn’t technically possible, or has tradeoffs XYZ.
You had lazy people before, sure, but those interactions were less frequent, shorter, and for each person I used to craft a careful response that would explain everything in detail and teach those I interacted with. But why would I waste my time crafting a reply that the person on the other side is just going to shove into their LLM and learn nothing?
Regarding open-source, LLMs make people more disrespectful of the people doing the work. And if you’re sharing your expertise for free to be disrespected, why bother?
They object to the systematized devaluing of their work. They object to becoming a class of invisible laborers (who only become visible when there is need to blame someone).
There's people saying they'd rather talk to an LLM than to me.
I dunno how else to say it to you, but when somebody happily announces that they'd rather talk to a computer than to you, and they are happy this is the future.
That carries certain connotations beyond "external validation".
The person in the top of this thread was claiming "nobody has been robbed".
There's people in these communities that I will never have an intelligent conversation with, again.
I'm not talking about being able to pay bills or anything, I'm just talking about being replaced by a computer program, in a social sense.
And the people make the culture. So, in effect, the culture suffered. I think you’re taking the original comment too literally with your “piracy is theft” comparison. Obviously you can’t literally rob culture¹, because culture is a concept. But you can lose something from it, and that loss can have a cause you can point to, so in a way elements of it can be “robbed”.
That’s what I get from the original comment when I steelman it, anyway.
¹ Sounds like the name of a Rob Zombie associate.
The world is bigger than your weird, entitled view.
You basically just told everyone that the existence of the demoscene "irks" you. And you consider this being "irked" somehow more relevant than these actual communities being torn apart by the existence of LLMs.
Oh no you called it "unfair gains", unfair according to those rules??
Stopped publishing anything and stopped creating anything.
But in all fairness this started long ago. I remember a conversation where I asked someone who read a niche blog every day why he never posted a comment. (the blog had zero comments) If you like what they wrote and have thoughts about it, why not "honor" them with a few words? How is that to much to ask? They clearly just never bothered to but in stead made up all kinds of excuses.
Imagine having to explain that you can respond if someone talks to you? If there was a culture surely it ended there?
A lot of my peers who have been in tech for more than a decade are all doing this. They're effectively backing out of the industry. Their attitude now is they're ready for retirement in their 40's. They're only building and working on side projects and only sharing with people they know, nobody else.
One thing I'm seeing right now is the amount of institutional and domain knowledge being lost is a on a massive scale. The more startling fact is nobody seems to care - as if vibe coding and LLM's are going to fill this gap.
But then again, I haven't devoted my life to developing and releasing a huge open source project used by millions, so perhaps my opinion should be discounted accordingly.
The reaction to LLMs pretty clearly revealed who was giving things away, and who was indirectly selling them and using "giving away" for unfair advantage.
BTW, you mentioned "pay it forward" -- please explain me how an LLM user no longer interacting with other humans, is going to "pay it forward".
I dunno what "irks" you, what's your problem, who mistreated you so badly that you have to be so condescending and attributing malice to people whose motivations you have absolutely no clue about.
There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:
Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.
To use another Slashdot cliche: "But you knew that already."
Citation needed
[Edit] Here's an article from 2003 talking about how Random House destroyed up to 25,000 books a day: https://www.theguardian.com/books/2002/mar/19/fiction.stephe...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
No, they are not. For all the hype around that 404 media article, no one came forward with titles or authors. Turns out the booksellers that sold them those books said they were "dead inventory".
If you're going to assert that companies are destroying the foundations of civilization, you're going to have to produce quite a bit more evidence than the (poorly cited) 404 Media article.
Here is what was signed by Judge Alsup at https://docs.justia.com/cases/federal/district-courts/califo...
> Anthropic spent many millions of dollars to purchase MILLIONS of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
Dismissing the above as "dead inventory" is atrocious reasoning. The more obscure/low-circulation books are arguably MORE valuable as training material because they contain information that is less likely to be duplicated across the ordinary web.
And 404 Media's reporting wasn't merely based on a bookseller's terminology of your "dead inventory". There are now unsealed Anthropic internal documents describing Project Panama as an effort to "destructively scan all the books in the world"
https://www.washingtonpost.com/podcasts/post-reports/the-que...
But nothing you're citing adds up to "destroying the foundations of civilization". What we have instead is companies that are now paying to get copies of creative works (isn't this what everyone is complaining about with all the "theft" rhetoric?) so they can use them for training.
When they discuss scanning all the books in the world, they are not talking about every copy, since that makes no sense. They want one of each.
So, in the end, this all sounds like some breathless emotional take because large companies are buying copies of books and using them to train AI and throwing them away. They are probably destroying fewer books than the publishing houses do over the same time period.
This is like the water discussion: people with no understanding of agriculture (or publishing) hear a sound-bite about what AI is doing, and, with little-to-no context, devolve into a panic about it.
Again: cite real harm. What books aren't available? What knowledge was destroyed? There are real concerns about AI, but this whole discussion is a distraction.
By your reasoning, if I could afford to buy every surviving painting by Van Gogh and Rembrandt and I burned them in a bon fire I wouldn't be taking anything from human culture and nobody should care at all about what I do with my belongings.
Destroying something that is rare or one of a kind and valuable does matter. Doing it so that you can exclusively profit off of charging people for software that regurgitates lossy and error filled versions of what was lost makes it even worse.
If you actually object to the concept that people that own a physical object get to decide how to dispose of it, that would be an interesting discussion to have because it would fly in the face of property ownership rights.
Like, tell your mother or HS teacher you decided not to recycle because you own your trash. See how well that clever argument goes.
putting that sorting and disposal on end customers was a business way of offloading their making single use trash problems on somebody else.
point any blame for cans in the trash where it belongs - can makers
why do you need more than a picture of the rembrandt? there's so much good artwork easily available that the rembrandt is only worth money because hoity toity people want to store their money in it and in the future sell it again
If you find this a "breathless emotional take", then I don't even know what to say. Nothing will meet your "evidence" criteria to dislodge your fervent fawning and support for destruction carried out by AI corporations.
If I do something destructive to my copy of Accelerate -- feed it through a scanner, leave it where the munchkin can reach, whatever -- that does nothing to deprive you of your copy. Books are funny like that; there tend to be lots of identical ones.
Surely that's not about AI? Or you mean the small fraction of that that's result of "destructive format shift" process[0], which itself is a consequence of copyright regulation that would otherwise prevent anyone from accessing these works?
> the outages and increased usage bills for once free websites (some resulting in closure)
That's assumed to be AI companies for some reason, even if they have no real incentive to do that, while the usual business underbelly of people scraping web for whatever reasons (which now may include some AI upstart wannabies too, to be fair) is forgotten about. Not helping is people confusing AI agents acting as user-agents and doing one-off fetches with "AI scrappers".
> the increased difficulty to access public information such as Reddit and Twitter
It was never public, and they started locking down before AI, when both platforms (as well as all other social media platforms) run out of VC subsidsies and realized they need to start monetizing; first step they did was to wage a war on third-party clients and API users. AI came later.
> ignore the amount of conversations being taken away from the public in favor of LLMs
You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online? I can understand that, it's usually much better experience (particularly on sites like StackOverflow - they dug their own grave here, and they know it; there have been memes about this way before LLMs were a thing).
You're also not considering the amount of questions answered that would not have been asked otherwise. I for one don't ask many questions on-line, so anything I ask LLMs that they solve for me, is a question that would've remained unanswered for me otherwise. Not everyone is gregarious online, many people are more self-reliant and only answer questions, but solve their own problems without asking for help (it's probably not optimal thing to do, but that's another topic).
--
[0] - Digitizing and retaining physical copy is clear infringement, digitizing but destroying the physical original can be argued to be fair use.
Oh, please, AI is single most hated technology to the level we did not seen before. And that is after staggering propaganda going out of these CEO that hiding behind human agency is beyond lie. It was pushed on us, whether we like it or not, because people with a lot of money are the only ones who matter.
And people with a lot of money have that weird cult belief in emerging AI god and singularity, so they dont care what they destroy in the process.
Regardless, these books were available. Researchers, hobbyists - many people have and will find value.
Justifying the destruction of information for a for-profit company is bizarre.
The emotional reaction we're conditioned to have hearkens back to 1940s Germany, where the goal was to have people burn all the copies of existing works. But this is not that. They are going for "one of each book" so they have a comprehensive corpus of works to train from. They are willing to buy books to obtain this corpus. It sounds like the system working exactly as intended, with even the book sellers selling these books saying they are "dead inventory".
There's a time to be breathless about "destroying all of humanity's knowledge", but let's wait until it's actually happening.
Again, I'm am very anxious for anyone to come forward with credible information about any valuable/rare/irreplaceable book that was destroyed in this process. So far zero people have been able to do this, probably because they are buying the books in bulk and want to obtain them as cheaply as possible. Book sellers will know the value of any actually rare/valuable books, and would charge accordingly, which would be pointless for e.g. Amazon to pay. They don't need some crazy first edition, they just need some copy to digitize. The incentives just don't add up to "Amazon is buying priceless art and burning it!" the way the narrative suggests.
People need to face reality. The old Internet was fragile and couldn't have lasted long. It was dying slowly due to closed social media and LLMs are accelerating that death.
I'd rather it die quickly and be replaced by something more robust than decay slowly. I was sick of how it was before LLMs even came among.
This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.
Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.
They work very hard to protect their own IP so they obviously believe that IP is worth protecting.
So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.
It hasn't been literally stolen but indirectly it has been.
I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.
Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.
AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
There actually was theft. See how some big corporations slurped data from Libgen and Anna's Archive. Where did they pay for this?
See also this article:
https://www.theguardian.com/books/2025/apr/03/meta-has-stole...
And similar articles.
It is theft - there is no denying in that.
All these companies owe us a ton of money. But Trump protects the AI mafia so we won't get compensation for the damage they cause here.
It's literally stolen IP.
One minus epsilon of the corpus did not give informed consent, or get compensated. The fact that it's laundered doesn't make it any less literally stolen.
It is unclear to me just what percentage of tech-companies, are in the business of adding extra steps to an illegal process, of turning signal into noise. For example, the vapes made by Juul are a kind of hack around public health laws. The Amazon marketplace shields merchants who sell counterfeit goods. Uber and Lyft bypassed the local laws that applied to taxi services and the medallion system. Airbnb did something similar with the laws governing hotels, and delineating who owns and who rents.
It is a special kind of disappointment, to be someone who loves technology, and to have to work in the technology industry -- because it appears to be run by people who hate everything that is not money.
I know this is a real harm to individuals, and bringing this up is not some kind of anti-technological thinking (the luddites were OG with that, and I always argue in favor of them - they had a very valid point).
I also realize that in a few years, I very well may be in the same position here. So will most of us.
But that's a different argument than GP was making, different one from what I replied to. That was about losing culture, and losing a business model is not losing culture.
Also I was with you all the way until the last paragraph:
> AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet.
AI is doing exactly what users want it to - what I too use it for - it bypasses the spam and scams that stand between the user and the solution to their problem. Good riddance, in this.
> Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
And that is just bullshit. AI is not "taking everyone's content and selling it back to users", and the vendors are definitely not "keeping all of the profits" - on the contrary, they haven't even begun to figure out how to monetize almost any of the value their inference services provide, because they have zero visibility into how much any stream of tokens ends up being worth for their customers. They cannot tell whether the plumbing advice the LLM generated unclogged my toilet, or prevented a restaurant from having to close for the day; they cannot tell if the code their agent wrote won me a beer from a friend, or unblocked $2 000 000 dollar opportunity for my business. They capture none of the value from this either way.
I know, I worked very hard for 20 years to build up this business. I easily put in tens of thousands of hours to document, record and open source pretty much everything I learned and experienced in tech for free to help build up an audience.
It worked great up until AI, then everything went downhill very fast which is solely related to nothing but less traffic coming to my site. I imagine this will affect everyone, not just me.
> And that is just bullshit
It's not, whether or not they are profitable isn't important. The takeaway is AI creators consumed all of this content then they turned around and are actively selling that information to other people. This is the only part of the pipeline that matters, whether it's profitable on a balance sheet for them isn't important to everyone who had their content scraped, stolen and sold. The damage is already done at that point to the original content creator.
To pour even more salt in the wound, the AI creators actively cause content creators to directly lose money on hosting because every time they scrape your site that costs bandwidth. If your hosting provider is metered, then you're paying directly out of pocket.
No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.
If IP is real, then the AI companies have performed flagrant theft.
If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.
The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.
Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.
Instead we should try to find the most practical solution that has the most benefit - which will probably be more complex than yes/no.
Intellectual property is as real as private property (which is a lot more elaborate and weird than possession and territoriality, which is the most that has any natural basis). It's really foolish to claim one doesn't exist and should be abolished and the other this fundamental sacred thing that should be respected absolutely (as many do).
Intellectual property exists as a concept to foster the creation of more intellectual property. That is not the case for all private property, because I cannot copy your land or your car infinitely. That’s why IP rights expire at some point, or have fair use that doesn’t harm the IP rights holder, something other forms of property don’t have. Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.
I think AI training falls most likely in the fair use category of intellectual property: there is some societal benefit* that requires no actual harm** to the IP holder, therefore it’s a good trade-off for society if we poke a hole in the social construct of property to get that benefit.
*Let’s put aside the question of whether AI is good for society or private ownership of AI models is good for society. Important questions but separate from the theory. IF IT IS GOOD, it follows the above. If it is not good, then of course it does not.
**Also not a fully settled question. Again, an important debate to have and the tradeoffs here matter. If the harms are small enough, the societal benefit could be worth it. Both notes have to be true for this to be worthy of “fair use”.
No. it won't blow anyone's mind. Not a single person.
> but people created music, literature, art long before intellectual property laws existed.
^^^ cherry-picked, incomplete reading of history.
Music and art was created by patronage. Entire periods o where the art that mattered was made by the powers that moved the world.
For my whole life the goal was to move away from that era, not to see it recreated.
Also, this is a debate that English speakers can enjoy on a website for an US based accelerator. Most of humanity doesn’t even have the standing to be heard in this conversation.
The idea of IP as an economic tool to foster creation comes out of the UK and subsequently the US. In mainland Europe, IP comes out of the French Revolution and the idea that copyright is like a "moral right" that you intrinsically deserve for putting in the effort to create something. This viewpoint has de facto won out because as global commerce and global culture has become more and more widespread, everybody has standardized on the longest durations (the standard has long been "life + 50 years" in Europe) so one country doesn't have to worry about freeing up its works for "exploitation" by another country.
Nope. "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
The ends is promoting progress. The means for that is prescribed by a limited monopoly as an incentive to those ends.
At least in the US.
All property is fundamentally about exclusion and control to allow for private exploitation. That common bit of rhetoric about copying misses that point. The reason private property exists is not because land (for instance) can't be copied.
> Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.
You should note that you misread me: I didn't say they were identical in every respect, I said their "reality" is the same. They're both made up social constructs. One isn't more fundamental than the other.
Both things can be true at the same time. Land being economically scarce necessarily means that whatever use is thrusted onto it will produce private exclusion. Today, we give this relationship of private exclusion the name of "property" or "private property", but that doesn't mean you have to formally recognized private ownership in law to suffer its effects - socialists regimes stand as a testament to that. Thus...
> They're both made up social constructs. One isn't more fundamental than the other.
The first sentence is reductive, while the second is innacurate, if not outright wrong. Physical property is primarily predicated on real scarcity, while intellectual property is purely predicated on social control, which, when analyzed more closely, lays claims (even if partially) of ownership over other resources, including peoples' bodies and mind. It's fair to say that one is more fundamental than the other.
Hear me out. Property rights exist to assign stewardship and use rights over rivalrous goods. goods where the use of the good precludes the use of that good for its purpose by another party. If you take my bike, I cannot ride it to work. This concept exists to prevent violent conflict over non-shareables and to prevent the tragedy of the commons (see the highly successful fisheries rights, NOx and SOx emissions markets as propertization schemes)
"Intellectual property" (except for trademark if you want to get pedantic), is not rivalrous, and its primary purpose is to be shared, not hoarded. Therefore intellectual property isn't a thing, creating it as a legal construct was a mistake that has hamstrung society for a long time.
What copyright lets you do[1] is offload that risk onto a publisher[0]. Instead of having to pay to commission every piece of art, a publisher can do that, and then sell the now-monopolized copies of whatever art turns out to actually be valuable.
A lot of hay was made during the Piracy Wars over filesharing tools breaking this bargain. A bunch of data hoarders with an interest in sharing media made it possible to just get the shit for free. This created a social dynamic where artists were annoyed about it, but publishers were Fucking Pissed. You could even measure how publisher-brained an artist got by how angry they were over Napster[2].
AI generated art also breaks this bargain, by making creative labor nearly non-rivalrous. The only cost is electricity and GPUs. This has created nearly the opposite reaction: artists are pissed while publishers don't care, because AI is to publishers like tort reform is to insurance companies. A publisher that gets art for free doesn't care if everyone else has it, because they have the payola dividend: they can push whatever slop they want onto the market and the market will eat it because they're big and powerful.
> This concept exists to prevent violent conflict over non-shareables and to prevent the tragedy of the commons
The copyright maximalists would argue that free reuse of creative works is a tragedy of the commons. I certainly remember hearing that phrase bandied about a lot during the Piracy Wars.
It's also important to note that "tragedy of the commons" is not a natural law, but a specific framing that is used to justify antisocial ends. The communal ownership so decried worked perfectly well in England for hundreds of years until the ruling classes found it inconvenient and abolished it. The kind of ecological collapse the tragedy attempts to invoke did not happen because there were already communal means of preventing overuse of the land. In fact, an emissions market is probably closer to communal management than enclosure.
Also none of this changes the underlying logic that AI companies are trying to enclose the intellectual commons, and that their business model relies on being able to replace human brains with machine intelligence they can rent out by the megatoken.
[0] Individuals who self-publish included.
[1] To be clear, copyright was created as a censorship regime, it just happens to be useful for other things.
[2] In Lars Ulrich's defense, they weren't just angry that Metallica songs were on Napster, they were specifically angry that Napster had their latest album before it was in stores.
That very much falls within the "IP is not real" category. Taking a utilitarian approach here is exactly what the sam altman / effective altruist crowd is doing (or claiming, at least).
Believe it or not, but there are different philosophies. The US Constitution, for example, is written under the framework that all rights are innate, and the government merely endorses, not grants, those which are described. This framework creates a moral basis to rights, such that things like property are not merely social constructs but moral goods. To violate them is itself an immoral act.
AH. Given that framing, then I suppose anything that varies from "these axioms are perfect there can be no others" must fall into the "not real" category.
But isn't there some debate about the moral axioms themselves? A space that is much more complex than "yes/no"?
EDIT: I always wonder, when I get two downvotes in the same moment, if someone is cheating. It happens so often.
EDIT2: And now! 3 upvotes in the same moment. I daresay someone is confessing.
Most people want a clean environment while the current US administration is pulling back environmental regulations to allow for more pollutants. The current US Supreme Court has taken gifts from those that have invested interest in the outcome of their judgements.
Social axioms are mutable while mathematical axioms are immutable.
Human trafficking is bad. How many of the wealthy that partook of Epstein's trafficking have been prosecuted? This shows that human trafficking being bad is mutable based on wealth and power.
Well... no. I'd say that most of the social constructs I'm thinking of predate wealth. Most likely shaped by evolutionary pressures.
> mathematical axioms are immutable
I invite you to Google "the axiom of choice" to see a lively debate.
Money is a social construct for trading value, but is not itself a prerequisite for the accumulation of value (wealth).
> ergo no social construct predates wealth.
What about the social construct of parenting?
But parenting in humans varies from culture to culture. That's what makes it a social construct.
How wealthy is your dog?
If you think the only measure of wealth is money, you have it backwards- money is a polite fiction representing an abstraction of value, and is only good for use as an exchange into another more concrete form of value.
It is this biological imperative that pushes humans to desire ever more, and only higher order thinking- social behavior, religious discipline, whatever floats your boat- that suppresses those instincts into virtue, altruism or civility.
So, when I replied to the following:
> Well... no. I'd say that most of the social constructs I'm thinking of predate wealth. Most likely shaped by evolutionary pressures.
I'm pointing out that the social constructs were shaped as a reaction to our nature, which includes a intrinsic, if subconscious, understanding of the concept of value and wealth.
The number 4 in a set is immutable, 4 == 4 under all stances, while it can be taken out or added to the selected set. Objective reasoning.
Example: The OS function that applies IP port assignments. Application may utilize one or more port numbers. When the application is done, they are moved back into the unused set. The port numbers themselves are immutable.
Social axioms are mutable because subjective reasoning is being applied.
Example:
Ask a person if they support domestic abuse. They most likely will so no. Ask them if they ever funded a performer that committed acts of domestic abuse after the facts were revealed. There is a high chance they did and write it off by saying, "don't look at their character, look at how well they perform. That was a good show, movie, book, ..."
Ever hear someone use the phrase, "that is different" when describing two equal scenarios that apply to two different people?
Kanye West still selling out shows highlights that people are willing to reject objective reasoning in a social environment.
Or to rephrase, social axioms are mutable because they can collapse and no longer exist.
So mathematical axioms are in flux.
Gödel's incompleteness theorems is built around this. By saying axioms change then one could create an universal math theorem. Since axioms are finite in definition and applied in a finite space, functions, there cannot be an universal math.
Example: Fibonacci sequence only works with positive integers and zero. It will work with negative integers wrapped in absolute value. The axioms of Fibonacci did not change while wrapped in absolute value. Rational and irrational numbers can work with Fibonacci too by applying another function bound to absolute value.
Hmmm. I'm afraid we're just going to have to agree to disagree.
You have a nice day.
I'm also a fan of nuance and non-black&white thinking but that is quite literally not how the rule of law generally works
> This framework creates a moral basis to rights
The Constitution (and the thinkers that it was based off of) creates a moral basis to rights. At the time the Constitution was written, this was a pretty radical idea. The most prevalent moral framework at the time was the divine right of kings, the idea that the king was ordained by God to rule and anyone who questioned that right was speaking heresy.
You're right that it was a relatively radical idea at the time, but not without precedent. The Dutch, Venice and Genoa were republics, the Swiss Confederation was a union of cantons and territories, for example, not to mention the influence of the Iroquois Confederacy, John Locke, Montesquieu, and the assorted Greek and early Roman city states.
Note that society is more about justice than it is about philosophy or logic. Whether you believe that IP makes logical sense or not, is not terribly relevant through the lens of justice. To allow the Altmans and Amodeis of the world to escape the kinds of consequences that Swartz could not (despite having purer intentions than either -- despite being a genuine utilitarian, instead of an aspiring oligarch _pretending_ to be one), would be the height of injustice.
It would be a grotesque and cruel insult to all the people who remember Swartz, and share his values and optimism about the internet and computers.
There's a flip side to this as well- President Lincoln famously suspended habeus corpus unilaterally in violation of the Constitution, outright ignoring a Supreme Court ruling in the process. It became moot when Congress retroactively approved the suspension and allowed it to continue for the remainder of the civil war. Lincoln remains celebrated, despite what could rightfully be argued as blatantly unconstitutional power grabs.
The use of the term divine at the time was more of an aftereffect of most people having some kind of belief in god and using that as a casus belli to redistribute power. The actual causal reasons seem more technological and more of England spreading themselves too thin. Lincoln went on to further prove there is no divine power enforcing these rights, it's the distribution of power and information that defines the rights you have. The Constitution is simply a piece of paper, not an agent.
We see this happening again in our modern world. Technology has allowed us to concentrate information like never before. If you believe knowledge (hence processing of information) is power then we've created new power structures that must change society.
Now, we can say things like "well, you can't concentrate information" but the other party can simply disagree and say it's my divine right to. You're at an impasse unless you plan on taking up a path to violence (any and every legal action is implied violence).
This is a straw man that nobody is arguing.
> Now, we can say things like "well, you can't concentrate information"
Again, nobody is arguing that "concentrating information" is bad. They're arguing that the expression of ideas in a specific format is a type of property, and that what AI companies are doing is immoral theft. Further, states and the federal government have the means and impetus to prosecute theft according to the powers given them by the Constitution. The only thing that is stopping them from doing so is that the claim that theft has occurred is contested, and the method in question has not been successfully tried in court in the broad.
Why would I want it to happen to more people? This is, at least to me, an insane point of view.
"Oh my god, the police just shot a black man for no reason. To make this fair they should also shoot more white people for no reason". Do you see how absolutely unhinged that sounds?
Yeah, this double standards thing is profoundly unjust.
>To allow the Altmans and Amodeis of the world to escape the kinds of consequences that Swartz could not ... would be the height of injustice.
Simply put, he's demanding two injustices. Let's break this down with a imaginary law.
"If you do not wear your hat you will be beheaded. This includes the poorest man and the king himself".
In OPs definition, this kind of equality is justice. But I do not consider this justice at all, simply equality.
"If you do not wear your hat you will be fined $1. This includes the poorest man and the king himself".
This is something much more close to justice as the punishment fits the crime.
True, but to put a still finer point on it, still fantastically more harmful to that poorest man than the king.
Intellectual property should not exist, but since it does, one can and should expect it to be applied consistently in all cases. Just look at what they did to Kim Dotcom, he basically got destroyed. SWAT rappelled down helicopters and descended upon this guy for the egregious crime of copying bits. That's exactly what should happen to Anthropic, OpenAI, Google and Meta. Anything less than that is an insult to previous enforcement victims.
Inconsistency robs authorities of all credibility. Selective enforcement exposes them for the arbitrary tyrants they are and delegitimizes their decisions.
And
"Conservatism consists of exactly one proposition, to wit: There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect."
It's a cruel and sad irony that people invoke Aaron Swartz to argue directly opposite worldviews depending on prevailing fashion, making him into some kind of perpetuum mobile in his grave. Literally the same arguments, using his memory to argue position that's opposite to what was argued two years ago, back before AI fully took off.
I personally doubt Aaron would be against AI companies doing what they did on the grounds of intellectual property. There's plenty of things to hate about how AI is transforming the world, but this is not it.
Inches aren't real. Laws aren't real. Language isn't even real.
Have you ever tripped over a line of longitude?
People live such abstract lives now that they're losing track of what is concrete.
On the contrary; I think. Nearly every definition disagrees with the claims you're making. It's not a mere synonym of "concrete" or antonym of "abstract".
https://www.merriam-webster.com/dictionary/real
And regardless, the usage here was clearly in the context of GP's (false) dichotomy between "real" or "false concept", so obviously the definition you're trying to prescribe doesn't apply.
Well, no, the physical laws aren't social. An apple falls whether there is a social group to see it or not.
And yes! Animals are a perfect example. Mother Nature came up with a nice strategy for maximal benefit in a social group, and she encoded it in our instincts and intuitions.
But that doesn't mean it's perfect. I have an instinctual fear of fire - but it's still better to cook my own food.
- amassed then destroyed/denied largest knowledge corpora seen so far
- poisoned knowledge for decades to come, deliberately or not, which would stop new developments by independent teams
- now they are crying that they are being robbed of their precious models, which - at this point in time - are neither truly good not economical
(not to mention actual environmental and social harm they cause)
this is truly first of a kind situation, that was never seen nor heard before
Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.
That's not to say that IP is useless, but treating IP infringement as "stealing" was always a problematic shorthand. Previously, calling IP infingement "stealing" was the domain of corporate interest groups like the RIAA or MPAA, but with AI the winds have turned and supposed anti-corporate lefties are keen to treat IP infringement as theft to attack AI companies. There's very little intellectual consistency on either side of aisle here.
When an AI company takes that type of value, and doesnt provide value back to the person who made it, thats theft.
Information / facts are real and free. But those who helped us get here should choose how to license their work, and an AI company picking it up and re-selling it is clearly unethical.
I can reasonably wrap my head around the idea that an author should be compensated for their work, but where in the social contract does this power to control how that work is used come from?
It's so destructive and tangles up courts and makes contracts complicated and we lose the original versions of the work because e.g., they have to change a background song due to complicated licensing.
I can respect protecting copy rights, but it should never be conditional; if you choose to make a work available to the world, you have a legal right to defend the unauthorized copy of it but you should not have a right to say how it gets used.
It reminds me of manufacturers like John Deer.
If you are freely posting sentences like this was on the internet you are giving away your IP for free. Anyone can read the sentence. If someone can make money off of it then that’s just markets at work. I can’t make money off of what I write here, for example.
But I also don’t think pirating a movie is theft either. You haven’t proven to me you’ve lost money. Maybe wouldn’t have watched it anyway.
Fun topic
That's what contracts and licenses [0] are for. Or perhaps you're arguing for a world where the only possible protection is that of trade secrets... that once protected information is made public by any means anyone can do anything they wish with it? If you are, then this quote [1] seems relevant:
If IP is real, then the AI companies have performed flagrant theft.
If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.
[0] ...which are contracts in disguise...Those contracts and licenses are just a form of protection enforced by the State. In general once knowledge or information is widely available it's also freely available. Whether someone can further protect the distribution or the usage of that knowledge to some ends is up to them.
I disagree with both of those premises.
IP is real, and scraping freely given comments or published material on the Internet and doing something economically useful with it doesn't entitle you to retrospectively go back and decide that you are owed some money. If that were the case you have to prove material damages. How much is my post here worth? one quadrillionth of a cent?
I think it can become different if, for example, a book was scraped or cataloged but that's primarily because it's likely that you can enforce or make a case within a given legal system to enforce your copyright or IP claims. But what if I read a book and then thought the main idea was my own, or just spoke to someone else about it and they spoke to someone else about it and it winds up in an LLM? There's nothing wrong with that or anything you can do.
To expose information is to put that information at risk of being used by others. It's up to you to protect that information or enforce claims on it. If Reddit's public site gets scraped and OpenAI does something economically useful with it, well, that's just life. You can't put information out in the public for free and then demand payment later. Reddit comments are freely accessible by the public, yes? Companies are part of the public.
If you want to argue that it's IP theft then distilling weights or otherwise reverse engineering the models is a violation of IP protections too.
[1] Rhetorical
Yes, that's how property works most anywhere that has stable government. Societies where The State has a monopoly on violence are usually far more stable and pleasant to live in than ones where vigilantism is how correction of injuries is done.
> Reddit comments are freely accessible by the public, yes?
In exactly the same way that the books in my local public library are freely accessibly to the public, yes. I'm sure that you're aware there are so many things you can't legally do with books in your local public library. Reddit isn't substantially different... go read the "User Agreement" contract that governs use of the site when you get a free hour or two.
> To read this post you must pay me $5. Please leave your contact information...
You can totally prevent me from -among other things- making many sorts of commercial uses of your post, but -as I'm sure you know- the made-up system you're gesturing at not how software licenses have worked for roughly as long as software has been a thing that was commonly licensed.
Anyway. It sounds like you want a return to "publish nothing, keep everything of worth secret to anyone not in the guild, and do what needs done to people who the guild suspects has betrayed its secrets". Those were the really bad old days, and the awfulness of that system was why the US adopted the patent system back in the late 1700s. It's also part of why it adopted the copyright system [0] as well as limited term lengths (fourteen years, with an option for fourteen more if the work's author was still alive and explicitly requested an extension).
[0] Though, you had to explicitly register your work with the relevant authorities to get protection... unregistered works received no protection. Things stayed this way for roughly two hundred years, and -IMO- should have stayed that way.
I don't think those user agreements are really enforceable or should be enforceable given how the public utilizes them. That's a digression though.
With respect to your library example, the difference is the author made the big with an intent to sell. There is a market for the book, it has some value as evidenced by the sale of the book. Public Reddit comments have no marketable value on their own. If Reddit doesn't want the public to use them, they must protect them. Potentially via the State, but generally they would have to take some other protective measure.
> You can totally prevent me from -among other things- making many sorts of commercial uses of your post, but -as I'm sure you know- the made-up system you're gesturing at not how software licenses have worked for roughly as long as software has been a thing that was commonly licensed.
How things work now or in the past can change. Maybe you're just slow to adopting a new paradigm? I can't prevent you from doing anything with my post unless you go on to make enough money that it is worthwhile. But if you just use it for your own private coffee cup there's nothing I can do nor is there anything the State is interested in doing about it. Similarly you can deconstruct Reddit comments into individual ones and assess the value, but the value is 0. Aggregating those into something useful is economically useful and Reddit doesn't really own the data if they are publishing it for free, the users do, despite what either party says.
It's similar to how copying != stealing. First you have to prove I'd have bought the movie. Second of all you have to enforce cases where I might accidentally see it I was walking down the street and someone was playing it in their living room. If you can't enforce it, you don't really own it.
Information economics is difficult because it's a mix of different things.
Hell if a car was downloadable why would I not ? Your capability to make endless profit should be capped somewhere, if indeed it is Profit
It’s called Derivative Work and it’s a good feature for IP law: https://en.wikipedia.org/wiki/Derivative_work
You wouldn’t like a world where companies could copyright knowledge and then prevent anyone else from making a derivative of that knowledge.
Imagine Wikipedia being taken down for sharing knowledge that another company wrote about first. It’s a bad idea.
That would be on what the AI model generates, not what it is trained on.
The latter is where the contention is, and it's a valid argument. So much so that some companies are not using stolen information to build their models.
IBM for example indemnifies its models for its customers and has detailed information on where the sources came from to train them.
It has been tried in court several ways already. Remember the lawsuit that forced Anthropic to use physical books? They tried to argue that the books couldn’t be trained on at all. It failed.
Arguably LLM companies could have made large-scale deals with libraries and got the exact same knowledge (much, much more slowly). I wonder if people would have the same issues then? My guess is probably. Goes back to the meme that if libraries were proposed today there's no way they would ever be allowed.
Your analogy is flawed because you are saying the information was free to begin with.
That analogy works for IBM granite models because the information they trained on is free to use.
The major LLMs did not do that.
People tend to forget that IP was not originally about digital distribution at all (copyright), it was about giving inventors exclusivity periods to profit without competition (patents).
It was a misguided attempt to stop sometimes literal theft of designs from rival inventors, by tying the design to the person instead of whoever possessed the schematic.
It's also a regime that in its modern incarnation protects businesses, not artists.
AI scraping for the goal of making a commercial LLM service is different from, say, a commercial file sharing platform.
The first difference is that LLM training is highly transformative. Let's say your LLM ingests the Harry Potter novels during its training. What you get at the other end is not the Harry Potter novels, it is a LLM that can talk to you about Harry Potter, it is not the same thing, and going from one to the other requires a significant amount of work, very expensive work in this case.
Not only that but there is no direct competition. People won't stop buying the Harry Potter novels because a LLM trained on it exists. If you want to read the books, you buy the books, you don't ask a LLM about it. A file sharing service on the other hand competes directly against the official channels, if you want to read the books, you can download it from this service instead of buying it on the official channels.
So, about how free you should be to get these weights from the AI companies. If you just share a 1:1 copy of the weights, that's the "file sharing" situation, not transformative, you took their work, didn't do any of your own. Usually considered unacceptable by IP laws.
Distillation is a more interesting case, you are using a LLM to train your own, it is transformative work, but you may also be competing directly against the LLM you are distilling. So, in a sense it is worse than scraping, but still, despite how much the likes of OpenAI and Anthropic are complaining, it seems to be legal.
So it is somewhat consistent: 1:1 copy and distribution is not allowed, be it source material or LLM weights, and training is, be it source material or another LLM (as in distillation).
This is how IP has always worked. It has never protected the little guy. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.
Nothing more and nothing less, and all of the normal debates about what should remain Capital and what should be The Commons apply.
Nothing short of a global revolution can fix this. Proprietary LLMs should be illegal. Either we achieve post scarcity within this generation or it's literally over.
Considering its become exceptionally difficult to search for things that used to be easy to find, I have to disagree with you there.
The web has been polluted with trash, covering up all the original media with messy imitations.
I doubt it, those are the same group of people anyway.
Orders of magnitude more people are using LLMs to brainstorm or generate content for them, which they are then using to solve their own problems and carry on with their life, increasingly solving many more problems for themselves this way, and not publishing any of it.
The world isn't made of "content creators" flinging crap around in hopes of monetizing it. Most people have actual jobs.
I guarantee it's scraped our work.
So where's our paychecks.
Not the least because I'm already getting many orders of magnitude more value each day from them providing me inference as a service.
It's different from search engines because you're not even getting the chance to monetize it yourself, you lose the ability to even show some ad for a pittance. Hell you don't even get credited most of the time. Many many people didn't give permission for the data to be ingested and used this way.
We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.
All of these lag SOTA by ~6 months on average, so yes, indeed, the thing you could only pay for half a year ago, is now free for entire world forever, and it only keeps getting better with time.
https://en.wikipedia.org/wiki/Cultural_evolution
https://en.wikipedia.org/wiki/Memetics
The past is still there, but our culture, being a live thing, is pretty much being affected by AI. We now have an artificial intelligence in this loop of exchange of ideas. Slopifying everything in its way.
It's a robbery of our future culture, at least.
That is a different argument entirely, though. Not the one GP presented.
On this, I have no concrete views. Sloppifying is real, and AI is short-circuiting "this loop of exchange of ideas", for sure. On the other hand, we had several precedents on this within last 700 years - the printing press, the radio, and the Internet to name the largest three - and culture turned out fine. Different, but fine. AI feels more than this, though, so I'm not putting that much stock in argument from history here.
But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of trying to deny it.
> And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
We don't have a society where "we got back" anything, and if you weren't aware, this "intelligence" is definitely getting metered. If anything "we" got undercut, and "our" assets lost a lot of their value, while others will get very rich from that loss. Come back with communism, and then maybe you'll have a point.
[1] Fuck you, Claude, for making me wince at that.
For your example to be trivially valid, the farmer would need to be entitled to the sun, but the sun is not property of aynone as of yet. There might be an argument to this (especially if more context is given), but, in this general case, I find it unlikely.
Still, even we accept your example as valid, what would be the equivalent to "sunlight"? Labor, maybe? But labor isn't something that is owned nor it's a persistent property of material resources, it's human-derived transformation. It's inconcrete by itself, even more so than light, more akin to work (in the technical physics sense). It just so happens that it "produces" a material outcome.
It isn't you that AI communism is stealing from.
This is good, but will not fly in court. Your honor... I only moved the Picasso...it is still there...
It's there in much diminished market, one that has been saturated overnight on a scale previously unfathomable. Their art is still there but it has been used against them to obliterate the playing field.
> nor are the people involved selling it back in any form.
Aren't all of these AI companies selling subscription models for people to create derivative art?
> And let's not forget what we got back for this
Get back for what, exactly? Didn't you just argue that no theft or sellback has occurred?
It used to be that, no matter what you want to do, there's a video on youtube of some guy who knows what he's doing showing you exactly how to do it. It could be, like, compressing a rear disc brake cylinder on a car or matching fiberglass gel coat colors, or carving a statue out of marble.
Theoretically those videos are still in there somewhere, all 15 years old at this point, but you'll not find them with youtube search instead you'll find a million AI slop videos no matter what you query.
There's a way that many people fear this is true: what you're getting for "free" has some external cost that you aren't taking into account.
Cost 1. The environment. Tech companies were once significantly interested in efficiency, their carbon footprint, etc. With the advent of LLMs, this was thrown out the window. Concerned people are now thinking about water footprint, heat footprint, noise footprint, and probably more. These things are hard to put a dollar figure on in a short comment, but the sustanability of a liveable climate for the billions living on this planet is literally priceless.
Cost 2. Employment. We face a significant cull of employability -- graphic artists, programmers, mathematicians, paralegals, and more are rightly fearful for the end of their career. If we're getting intelligence for "free" in 2026 dollars, but we can't get jobs in 2030, will all that content still be affordable in 2035?
Cost 3. Infinite investor dollars. The AI companies are burning cash at a historically unprecedented rate. This has attracted a lot of attention from traditional investors, like pensions, banks, etc. If all of this ends up in a free product, that doesn't sound like it'll return those investments. And that can crash the economy -- I ask again about real affordability in 2035.
This has another side effect in that it breaks the power that consumers have in the market to have a say in how resources get allocated, by voting with their wallet.
The entire AI buildout is non-consensual. Individuals have zero say. We could all refuse to buy AI subscriptions and it would not matter because businesses will still buy, and, investors have decided that we are moving forward with this no matter what, seemingly whether anyone is actually buying or not.
AI isn't necessarily unique here, but it is one of the biggest new examples of wealth inequality and how society at large are no longer the ones who get to decide how resources are allocated under a capitalist system, especially when the investor dollars behind it amounts to the GDP of a small nation.
And this is fuel for climate change denial. See? Are the wealthy and powerful people acting like they are concerned about the climate? No. They are doubling down on energy use and consumption. It's all a fraud!
1. is mostly bullshit that's used by anti-AI crowd to gather further support. Data center companies did not stop caring about footprint, they're still pursuing that for the same reason they did it before: it aligns with overall lowering of costs. AI added an extra incentive to improve efficiency because the demand far outstrips supply of compute.
For 1., also the general argument stands: yes, these data centers use energy, just like everything else humans do. What matters is the value it provides to people, and in case of AI, it's one of the most objectively useful expenditure of watts on compute (your point 2. notwithstanding).
I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
> Your point about efficiency is particularly obtuse: efficiency doesn't mean burning less fuel, and getting more compute per Joule. Not reducing the overall carbon footprint, but accelerating the burn rate as fast as we can build.
Efficiency does mean getting more compute per Joule. They are doing that. They are also building out as fast as they can, because both of those address the same problem: demand for compute outstripping supply for it. They can't just pursue the build-out strategy, because hardware production is now supply-constrained too - that's where the "RAMpocalypse" came from.
Yes, this is a huge build-out, and little to none of that is 100% carbon-neutral, so ecological footprint adds up. But that's normal and expected and a kind of tradeoff humanity has been making forever: building new stuff, be it hospitals or airports or data centers, has an environmental footprint, and we only hope what we get from it is more important to us on the margin, and that we can offset the environmental costs some other way.
> I'm not even sure I agree that demand is outstripping compute -- nvidia's circular demand-inflating investment/buildout/loan situation certainly muddies the waters on that, see my point 3.
I'm not talking about financial "demand". I'm talking about real demand, for compute. There's nothing to doubt here, it's pretty clear that all major AI players are constantly running at capacity - it's obvious in the very structure and limits they place on even the highest tiers. Unless you believe half the data centers are just spinning busy-loops and burning energy to inflate stock prices and generate a fake reason to build out more compute capacity - if yes, then I don't know what to tell you.
--
[0] - Like the water story, where the reported datacenter usage numbers look big in isolation, but once you relate them to "how much water is there" and "how much is used by other industries", it turns out to be a nothingburger. The only numbers that survive are those showing the story is really about those other industries having a competitor for cheap water now, and having to pay a bit more than they're used to.
Water scarcity is a problem around the world; even my rainy hometown of Vancouver had steep water restrictions this year. The question "how much water is there" is incredibly disingenuous, to the point of feeling like a motte and bailey: yes, the universe, or the earth contains vast oceans of the stuff but datacenters are using drinking water which is quite scarce. Just because "other industries" use a lot of water doesn't make it okay -- I level exactly the same criticism at other industries which waste water at such an egregious degree.
To be clear: if data centers only charged up a pool of water to use as coolant, that's fine. But when the weather is warm, they run millions of gallons of tapwater down the drain just to cool off. And heat spikes are happening with increasing frequency, severity and duration.
We as a species need to conserve water and emit less CO₂, full stop. The tech industry doesn't get a pass just because other industries are dirty.
Again with the weasel words. Real people, including individuals? Pretty clear this is written from a pro-business perspective and can be safely dismissed.
>"robbery of all of our culture to sell it back to us at a mark-up".
Well, it is that. Why is it criminal for me to steal a big-business-movie for personal use, but they can just take all my blog posts and sell their deritives to others?
I respectfully request that you get all the way outta here with this take.
Artists have caught AI generating literally their own work, for free, at scale. There have been tons of articles on this website about the AI hyperscalers slurping up books, copyrighted works, etc through legal and questionable ways. Even TFA says that the NYT suffered absolutely devastating CTR drops.
If you create a blog post about something super esoteric, it will guaranteed end up in the training sets for every big-lab frontier model within 24 hours (probably less) and probably show up in Google's AI Summaries at around the same time. It's also well-documented that their spiders don't honor robots.txt either and blocking them is basically impossible unless you give Cloudflare protection money (since spiders can super easily solve Anubis hashing challenges).
If that doesn't sound like theft-as-a-service to you, I don't know what to tell you. All I know is that hosting content ANYWHERE AI labs can get a hold of it is rapidly becoming a fool's errand, and we will all be at a loss for it.
You'll never agree but I think you underestimate how many people consider all that the AI companies have done to be the wealthy stealing and selling things back.
> nor are the people involved selling it back in any form
I am convinced that the longer one works in AI the less one has any grasp on reality
> too cheap to meter
Its the most expensive buildout in human history - you can't just split half the cost. The spending is the only significant growth in the US economy. Unless you think that all the datacenters and all that capex are for training?
However ordering a car with Uber through your phone just gets people to the destination faster, cheaper, more comfortably and reliably using those same roads.
This is theft I made up the route first >:(
It's actually not there. See how many websites have now closed doors or ceased to operate because of constantly being hammered and bombarded by robotic scrapers. For others, it made websites that turned a small profit into unsustainable money pits.
Sure, their content might now be ingested into an LLM training set sitting somewhere on proprietary servers. But the site itself (the origin of truth) now does not exist.
So no. It actually isn't there. Resting on this falsehood, the rest of your retort makes not much sense.
reified intelligence on a chip, almost too cheap to meter,
“Reified intelligence” is stretching the truth. It is certainly closer to intelligence than anything we’ve come up with before, it can certainly match or beat actual intelligence in some domains, but just as certainly it clearly lacks many features and capabilities of intelligence.“On a chip” is also stretching the truth. The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them. Individuals who have free usage for a project commonly report that, had they paid, it would cost five or six figures. Companies have seen enormous bills, some approaching the size of their payroll. And all of this is true for tokens that are dramatically subsidized, by one of the most intense and largest concentrations of capital in history. It is the polar opposite - “almost too expensive to even do, and absolutely must be carefully metered”.
In summary, let us indeed not forget what we got back for this: extraordinary distortions of reality evenly intermingled with bald-faced lies.
It's not complete or that well-rounded. But it's something that was the domain of speculative science fiction only 5 years ago, and it's rounded enough to be applicable to ~everything to some degree.
> The kinds of models that you might point to in support of “reified intelligence” run on things the size of a desktop computer and cost more than a car, which is way off from the scale that “on a chip” suggests.
By "on a chip" I meant more "in silica" than literally on a single chip" - though this actually is* true, but those chips aren't cheap.
> And “almost too cheap to meter” is aggressively false. Users of this service are known to talk incessantly about being metered, it is a daily fact of life for them.
You are looking at power users that use LLMs in agentic coding sessions. Most people just run off free tier of ChatGPT, which is free for them. There are equivalent open-weight models at this level, and while hardware to run one for yourself is expensive even for most westerners, the marginal inference cost is literally dirt cheap, which is why you can get that for near-free from smaller inference providers - or pony up some money, rent a bunch of compute with friends, and become an inference provider yourself.
(It's only a tough market because the major vendors are giving out better models than you can run for ~same or lower price than you can offer. Which either way is too cheap to meter in terms of solving useful problem for real people. Again, developers are a special case of power users, as usual.)
Not problems, laziness. The way I see students and colleagues use it is to get their work done with less effort. That's its selling point. There are not many real problems LLMs address.
> no one has actually been robbed
In our society, people get paid for work. If you think society is wrong, fine, but you must state so first. Under common assumptions, all that data has been produced through work, and that work represents value. Taking it for free is therefore theft.
I think this is quite the pollyannaish perspective and very much inline with those that think if you can take, then take and only apologize when caught.
We won't get anywhere in these discussion if good chunk of participants cannot admit to the trivially observable facts about the actual objective reality in which they live in.
The kind of "piracy" you're talking about wasn't depriving the authors of anything because you could always make the argument you weren't going to pay for it anyway. If, on the other hand, you were making copies and charging people for them, you could definitely say you were depriving the legitimate authors of that revenue. AI companies are very much doing the latter, not the former.
The other part of it is it's not just copying. Previously, if I decided to make a copy of a work without paying, I'm only copying the work, not the author's whole writing style. Now the AI companies are depriving authors of revenue from works they haven't even made yet.
If they were able to create a model de novo then they could truly claim it hasn't just been lifted from existing culture.
sell it back to us as markdown
Ad hominem.
Exact same pattern - take it as an ad hominem all you want.
The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
> The markup is the millions in training they committed and the connecting the knowledge
Sure, but gated behind a hallucinating idiot.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow
[2] https://sydneyuniversitypress.com/products/9781920898168
[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons
Even if this were an accepted principle, that wouldn't change the principle of free use. In all of your examples, only re-printing all or substantial portions of the books of Andre D. Short would be copyright violations. Just referencing facts from Short's books, or even including small quotes, in your own new work is not a violation.
^ Of course there are other ways to alleviate the concerns too such as universal basic income, government grants, etc for someone who wants to dedicate their life to measuring the dimensions of frogs, or whatever else their interest may be. There would however be some geopolitical/trade issues involved--a population would have to be comfortable doing the heavy lifting only to have another country simply use the work freely and instead dedicate their lives to something less favourable such as building missiles.
No, it wouldn't. "Sweat of the brow" applies to collections of facts whose compilation required effort. "Life's work" is a superset of that. Originality and creativity, which are required to copyright something, are also work.
However, LLMs do sometimes output training data almost 1:1 without sufficient transformation, and these cases may be problematic if they could reduce the market for the original copyright owner. For example, if prompting an LLM with "Translate the first chapter of {book} from American English to British English" reliably did what the user asked, perhaps no one would have a reason to buy the book directly from the author.
[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwzqxzbpw/...
And OP's contention is obtaining the training material and using it in training requires making unauthorized copies. That's the infringement; training, not inference.
Furthermore inference indirectly affects the market for the artist's future work. Don't need the writers and artists the LLM trained on anymore, when it can do similar work for free.
There are ONLY TWO Stories and their details, that we collectively will never see.
1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...
And the other...
2) The Epstein Files.
AI is better at repackaging it back to the end user but ultimately I'm arguing it's the same thing.
(caveat: yes I know there were plenty of people that objected to Google et al indexing everything; famously, Gmail was scary to a lot of people because they read your email to give you ads)
At a mark-up would mean it’s more expensive. The outrage is that they’re taking knowledge that was expensive to access because you had to hire experts or otherwise pay a lot of money for it and making it accessible to anyone who signs up for the ChatGPT free tier.
Calling it “robbery” is also specious as no knowledge was taken away from anyone. The content in the training sets was out there in the world one way or another. It still is!
I’m really perplexed by this sudden swing toward the idea that knowledge is something that we should encourage or incentivize to keep locked away or that other people should be forced to pay for use of knowledge. Roll back the clock a few years and tech sites would be almost unanimous about knowledge being free and unrestricted for the benefit of humanity. I’m keeping knowledge separate from actual direct rote duplication of content.
Now we have this amazing era where I can download models to my computer, run them locally, and have enormous amounts of derived knowledge at my fingertips for the cost of some compute cycles. Except now it’s a “crime against humanity”?
In any case, Microsoft has stolen 25 billion from its employees in 2025, and OpenAI has got 13 billion in revenue from "stolen" content in the same period, so that'd make them about equally bad villains, except OpenAI has mostly stolen from other companies.
The web getting flooded with slop and drowning all original work is one way to go about that.
Another is paywalling and gatekeeping en masse.
A third is shutting down shadow libraries.
"Property" and "IP" discussions are distractions; no amount of it can rationally get us around the utter unfairness of what occurred and the way it will warp our economy at a basic level if not addressed.
It's not even enough to make the weights and models free; access should be free, and everyone who hitched their horse to this wagon should be on the hook for keeping the systems running, on their dollar. They took ownership of a venture that is short one (1) "Humanity's entire cultural corpus", and the only question is if we're going to issue a margin call.
first of all - a lot of people are definitely not forgetting it. perhaps many more are waking up to the fact. when so many people wake up to the fact that a massive theft of intellectual property IS what enabled present day AI, they will inevitably refuse to a) publish that much openly; b) respect any kind of copyright claims imposed by those who perpetuated, facilitated, enabled the theft.
so, really, a lot will be coming out of it, we like it or not.
These are two different things.
Note: am not an AI fanatic.
Honest question. There is a line in the sand somewhere apparently.
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
Probably indicative of America's wider downfall that they've all become so self interested
If it has no soul to save and no body to torture it deserves no rights.
Edit to add: I did buy my modern version of a CD player. It has no tone controls. I just have to accept whatever sound profile the manufacturer thought was normal and proper for everybody, which involves lots of bass and not enough middle. Nevertheless, albums! It's a way of life.
That's most humans.
This idea that corporate personhood is some perverse idea is silly. Corporations are just people acting in groups for economic benefit. Everything they do is done by members of those groups (or people they hire).
Similarly, AI is an inanimate tool, like a keyboard. No comment is posted by “bots” - comments are posted by humans running software.
The fact that they don't is crucial, and probably the very reason why the concept of corporations was invented.
If you can't tell the difference between "I want to share all information freely with my fellow mankind" and "I want to share all information freely, even to billion dollar corporations that are making the human-replacer machines that threaten my fellow mankind" I really just don't know what to tell you
Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.
Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
There should be a carveout for non-profit or government AI.
https://www.reuters.com/world/us/us-appeals-court-rejects-co...
As for your second point I doubt it will ever be economically feasible. AI companies cannot generate profits even while stealing their training data.
The laws at play here are related to Intellectual Property, specifically Copyright. Yes, it is terribly flawed, but it is the product of centuries of case law dealing with very hairy issues, and I believe it is fundamentally sound, and here's why.
As the name implies, it deals with only verbatim copies of works or subsantial portions thereof. It very expressly does not cover abstract things like concepts, ideas, themes, facts, or patterns, and rightfully so, because we really do not want anyone owning something that broad.
But these abstract things are precisely what have been extracted, at unimaginable scale, to build these models! Each pattern in the tokens derived from these works contributed imperceptibly tiny perturbations to randomly initialized weights, interacting in incomprehensible ways into vectors representing concepts and ideas and facts, the cumulative aggregate of which has somehow created a form of intelligence.
There is no copying, only gleaning, and so Copyright Law falls short. But what is the alternative, and do we want it?
To prevent something like this would require some sort of legal protection on the more abstract things. We do have a legal framework for those: Patents! But as is very clear on HN and in many Tech circles, those are an extremely contentious topic (even though they actually protect much narrower ideas than most presume.) I don't think anybody anywhere really wants any protection on broader abstractions, and rightfully so.
So: we as a society expressly decided these abstract things belong to the commons, and those are the exact things these labs harvested. This is probably the only logical culmination of our technological journey, and is within the very reasonable legal frameworks we have evolved over centuries.
As such, it is not productive to dwell on fighting this or bemoaning this. Instead we should focus on ensuring that this technology -- with its immense potential and opportunities and dangers -- benefits everybody as much as possible. That is a better way to compensate everybody's labor, and that is a much richer and fruitful discussion to be had.
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
p. 1
"This case is about, as Microsoft's Director of Applied Science put it, an astonishing theft of unprecedented proportions; SF1437, perhaps the largest theft of labor in human history. SF1652"
p.11
"As Microsoft recognized: millions of people around the world will soon consider large models hoovering up all their work to be an astonishing theft of unprecedented proportions and admitted that almost no one intended for content they created to be used in this fashion, nor are they compensated for its use. SF1437."
p. 74
"As Microsoft's Dr. Glen Weyl put it, compensating creators is in the best interests of my employer, of my country, and of many other groups I belong to. SF1657."
Hyperbolic quotes from Microsoft employees are, IMO, the least interesting elements of this brief
Here is Microsoft's brief. Note how MSFT responds to the "web grounding" claims
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
It seems OpenAI does not want the public to know about (a) OpenAI's data collection and retention practices and (b) the number ChatGPT users have requested deletion of conversations
https://storage.courtlistener.com/recap/gov.uscourts.nysd.61...
"OpenAI seeks to redact specific information about [(a)] the number of users who requested deletion of ChatGPT conversations and [(b)] OpenAI's related data collection and retention practices."
"Disclosure would give OpenAI's competitors insight into OpenAI's confidential business practices and customers and cause competitive harm to OpenAI. Yeats-Rowe Decl. 4."
Perhaps it would causes competitive harm because, upon learning about OpenAI's privacy practices, ChatGPT users might reduce their usage of ChatGPT
Declaration is sealed so we can only guess
LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.
AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.
Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.
Sorry but the point isn't to create artificial scarcity just so intellectual labor is well off, that's a negative side for the consumer that was considered necessary tradeoff. Market economy should be about providing the most value to the consumer.
Disclaimer - I was never a fan of IP laws despite them working in my favor, with AI I can see them finally being abolished.
We could imagine, as an extreme case, a technologically highly advanced society, containing many complex structures, some of them far more intricate and intelligent than anything that exists on the planet today – a society which nevertheless lacks any type of being that is conscious or whose welfare has moral significance. In a sense, this would be an uninhabited society. It would be a society of economic miracles and technological awesomeness, with nobody there to benefit. A Disneyland with no children.IMHO, any LLM output is, and systemically forever will be, inferior to human creative output, if not entirely worthless. We can scale and industrialise the production of slop, sure, but who wants that and to what end? We're not getting the next cultural movement or artistic masterpiece from these systems no matter how often we pull the lever of the regurgitation blender machine. Because human intent is the whole point, and without it, there is no reason to make any of this.
I agree that it should be about creating the most value to the consumer. As a consumer, the value I get out of LLM writing or art is almost always exactly zero, and it makes it that much more difficult for me to find the things I would see value in.
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
Does it?
“We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.”
All these failing companies are trying to bullshit their way out of the decline. Atlassian could have, you know, come up with a usable GitHub competitor. Instead they dream about coding by yapping.
If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.
If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.
Sounds about right for the people and orgs involved.
So training can make it legal as well. Interesting...
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here
He does not see the moral dilemma here?
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.
Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them.
An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible.
> If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it.
Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.
That's why the GenAI companies try to talk their way out by claiming training is "Fair Use". It means they acknowledge that the materials they use are copyrighted and not legal to reproduce without a license.
Where is the loss? Apart from trust?
I can access a public park, but that doesn't necessarily give me the right to also bike on its sidewalks, or walk on the grass, or take some of the plants home with me.
> ... or has been obtained illegally
Similarly, content that is _accessible_ publicly may be illegal to _obtain_, these aren't mutually exclusive.
On the internet, you'll find there are terms of services and licenses. These restrict how you can use even publicly accessible material. Public availability doesn't give you a license to use it however you want.
Of course. You listed complex examples from the real world. A park where walking is allowed but damaging the plants is not, for instance. There it makes sense to distinguish various activities that can be done in there and treat them separately.
But a website offers not much activities that you can do with it. You can read it. And that is about it.
With the advent of trillion dollar corporations selling extremely powerful general purpose imitation as a service, the meaning of “public access” has substantially changed, potentially invalidating the original agreement.
Yes. You can of course always close it to the public if the public usage bothers you. Or require those using it to agree to your terms and conditions where you restrict the speed, the weight, the time of the day, etc. Anything you want.
But if you choose to make it public with no restrictions, you have to be prepared to face the consequences.
This same group of people, now being on the other side table, are screaming an crying that it's not fair.
Grow up and reap what you sow.
Whew, good thing it’s too late to be accountable for that now, huh? Water under the bridge. Mistakes were made. Eggs, omelets.
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Look at who owns that site, no surprise here.
Big tech scrapes ALL OF THE WEBSITES CONSTANTLY to resell to you as knowledge? Shut up and take all of my money.
The 2020s is the wildest timeline indeed.