Microsoft CEO of AI Your online content is 'freeware' fodder for training models
theregister.com
theregister.com
The following rules are agreed upon by pretty much every country that has an interest in copyright: https://www.wipo.int/treaties/en/ip/berne/summary_berne.html
This is called a search engine.
Small extracts or low quality assets are usually permissible, full on scraping is not.
It's really crazy. If I search for any of the things I have posted on SO, for the last dozen years, I get multiple hits.
There are quite a number of sites that just full-on scrape SO, strip out the attributions, then republish, wrapped in their branding.
I suspect that this is done for many, many other content-rich sites.
"anyone can copy it, recreate with it, reproduce with it" are his words. Those "with"s are important!
So if you publish, say, a novel or some expensively researched reporting online, you are now to assume it's legal to have an "AI" paraphrase it and sell it on Kindle? Or MS monetize it on their scummy "Microsoft Start"?
That behavior is a lot more aggressive and unethical than what Google does on its news aggregator (although that has become unusable for me on mobile since the latest redesign, whatever).
sure, Google would probably like to anonymize all "content creators" too, and drive the value of arts and information to zero.
I don't buy the constant equation of linking to external sources (search engine, link aggregator) with the automated mass-production of uncredited copycat content.
That's what I consider an important difference.
It feels weird when link aggregators link to MS Start, which itself seems to be an AI text generator paraphrasing other's content. Not sure because I tend to leave early.
Guess they have some licensed content as well.
When people don't value content, it's cynical to complain about the decline of journalism.
Existing copyright law already covers that. It's the same rules.
> "anyone can copy it, recreate with it, reproduce with it" are his words. Those "with"s are important!
So I was a very strong anti-DRM proposant when younger and I basically still am.
Information is inherently free, yes. And every creation that is understandable to others must be derivative of other works.
But no, I still think that existing copyright law does not cover AI slop generated from copyrighted content.
It covers "remixing" and fair use to a degree (or, more precisely, there is legislative precedence about many nuances of copyright).
But copyright law deals with humans.
It is not designed to cover automatically generated copycat "work" at scale.
Look at actual cases about derivative works and the line being drawn: no way this covers the situation with LLMs, it is not remotely practical to have a complicated court case for each piece of AI slop generated.
To the contrary, LLMs are ideal tools to obfuscate the actual source of any original piece if information.
Yes, the existing copyright of the next HBO series won't be affected.
But ad an indie author or musician, it could become even harder to achieve any kind of meaningful revenue.
I think this tends to degrade all domains of art to a pastime for people without any material needs.
Not that this is principally new, but the kind of automation that AI allows has the potential to ramp this up to eleven.
That's kind of what other human creators collectively are. They're competition and they are the reason indie artists have trouble making money. There's too much supply. But that's a good thing for people in general! If AI acted like that, it would be even better, and perhaps artist would no longer be a job but more like a hobby - like gardening or making model railways.
- publish content (music, book, blog post, whatever).
- content gets automatically plagiarized, so that it's not an exact duplicate anymore
- now you'll have to sue and win in court to claim your rights to the content
So the same scenario as with human plagiarizing, but at a scale that makes it harder for authors to claim their copyright or argue against "fair use".
> There's too much supply. But that's a good thing for people in general
No, I don't want to listened AI-generated slop "in the style of" my favorite artists.
Music is a great example. For small excerpts and some background listening, I might not notice at first. But after a while, I do.
Same with text: I have not ever read a single piece of original content worth reading that was "created by AI", do you?
The "supply" here is not getting bigger because of AI. It is diluted.
What do you prefer to read, a StackOverflow answer from 2012 with some edits afterwards, or a translated sunmarization with subtle errors and no attribution?
To address your question:
> To help me understand you, imagine if AI were as good as it is but somehow not trained on anyone else's material without their explicit permission. Would everything be OK then? It's still undermining an indie author or musician's ability to earn money from their own work.
I have not yet heard any AI output that was interesting music on its own to me. Interesting sounds and loops maybe, compositions: not so much.
I find AI as a tool for creators more interesting.
What I am worried about is that this won't happen. That it will only be used to plagiarize at scale, and replace creativity with a mixture of randomness and interpolation of already present ideas, optimized for commercial exploitation.
In this scenario, everything that you put on the internet will be merged into mass-produced culture products by large companies, before you even have any chance to try and claim any kind of copyright.
In the short term, it's unimportant then if your work would be used as input/context or training data.
> If AI acted like that, it would be even better, and perhaps artist would no longer be a job but more like a hobby - like gardening or making model railways.
I used to argue against copyright in a similar way. Money is a bad motive anyway, etc
But this is not practical. Cultural progress and interesting works are rarely done as a leisure activity of people who have to work a regular job full-time.
The cynicism is mind-blowing. Artificial Stupidities have created exactly zero knowledge so far, which is why companies like Microsoft now roll out people like Suleyman who openly admits that the DMCA is only for the rich with lawyers.
And steal and repackage "content" from altruistic creators.
This one might bite back Microsoft. If any case goes to the Supreme Court, I'm not sure that they'll be amused by this line of logic.
First they violate the law. Then they ignore it. Then ignoring it becomes the new norm. Then the new norm becomes the new law.
We need people who really value creation, democracy, etc and not pure capitalism, profit optimization, etc.
Drowning out creativity on the internet with remixed word salad will only be profitable for a short while. After which, people will start looking for human-created stuff, and avoid ai as much as possible. And slowly, it will become possible. Just like adblockers slowly advance.
"Everyone in the business" - you mean, the AI training business? Paraphrasing a famous call-girl, "They would say that, wouldn't they?"
This isn't even limited to tech. Rich people at worst settle with money to break the law, and at best completely get off scot free doing things that would put much pettier offenders behind bars. I steal some bread from a store and go to jail. Some rich dude commits millions in fraud and cover up and still runs for president.
We post regularly online without any expectation of payment but then we never considered that the output could or would be used for commercial purpose. The value of our collective output is being captured by a very small elite. Not sure what we can do about this, other than support alternative eco-systems at least to ensure that intra-corporate competition might keep prices low
For Hollywood writers, artists, musicians, book authors it is much easier to protest. You could go to Redmond for a couple of weekends, camp on "One True Microsoft Way" and protest.
And the NYT would cover this, since its interests are aligned with the protesters.
We need the public square back, we need it conceptualised as such. The death of twitter was horrifying, apparently you can just BUY it.
Excuse my passionate language, but fuck these predators.
We did. If you are posting here, that means you've explicitly stated that you read and agreed to Y Combinator's TOS. Which include this clause:
> you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed
And people are still fine with this apparently, otherwise they wouldn't use the service anymore.
or press an ear against its hive.
I say drop a mouse into a poem and watch him probe his way out, or walk inside the poem's room and feel the walls for a light switch.
I want them to waterski across the surface of a poem waving at the author's name on the shore.
But all they want to do is tie the poem to a chair with rope and torture a confession out of it.
They begin beating it with a hose to find out what it really means.
It's heavily skewed against us, but you can still argue the 3c's here. consulation is the licence, consent is agreement, compensation is agreed upon as none for the user.
What's not in that agreement is letting non-afiliated companies do any of the above actions. And I believe the privacy policy and GDPR cover other clauses pertaining to the sell of our data.
>Sale or Sharing of Personal Information. We do not sell or share your Personal Information (as those terms are defined under the CCPA).
And that social contract did not include AI. You put content on the internet because you had certain expectations on how it'll be seen or used.
Those things aren't what makes a copy fair use, though.
The EU is probably already acting on it behind closed doors, the usual vitriol against "innovation" peddled by the American-way of doing business went pretty high against the AI Act (which is definitely far from perfect but a step into some direction to regulate it), in the near future I can see more rulings or even new regulation to address the absurdity of AI companies consuming all this data for their own profits with no compensation to the creators of it.
My personal opinion is that leaving "innovation" be the only guidance to what is "good" without any morality imbued is stupid. A lot of us has seen the cycle by now, what was innovative before becomes entrenched, the entrenched companies become behemoths, and obviously start abusing their position of power when consumers have very little options to not be in the system they created. The downfall of tech from what I experienced in the early 2000s to what it has come to be in the 2020s is just sad, it's the new 80s finance yuppie bullshit, instead of coke-addicted greedy as fuck bros we have nerdy-blabbing-about-changing-the-world greedy fucks reaping the profits.
This will get ugly, and companies doing it will deserve the retribution if they get fucked.
I am ambivalent about this overall, but few things are clear. Someone getting sued does not automatically mean they are wrong and whoever is suing is right. We don't know the rules, and hence the lawsuit. I see it being used as an evidence of wrongdoing, and seems plainly wrong. Every thing that becomes big ends up being sued (including the artists with allegations that they copied someone's work). Tells us nothing.
(I think this part is clear) Reproducing content verbatim without permission, and for profit, is plain old plagiarism, whether it's done via AI or human. In some cases, with proper citations, it is allowed, but otherwise it's a no. For summarized content, with or without credit and citations, was always allowed, but never done at this scale, so this "social contract" might need to change.
The social contract has never been sacrosanct, it's enforced through social consequences. These companies don't feel social consequences, individuals making these psychopathic decisions are socially rewarded by their peers and allies, and financially rewarded in turn. Money is a flow of power, not of value. Especially post Bretton-Woods.
People talk about data, but individual data is as useless as they come for many companies. Data is only useful in aggregate and that too after reaching a certain scale. Counter intuitively, individual data (w or w/o PII) is more likely to be used for nefarious purposes compared to aggregated data. Aggregated data can lead to better user experiences, better utilization of our time, and can deliver more value to us (while delivering personalized ads). No one should create a blanket rule which bans aggregation given how it can ban Google Maps and Google search and Google ads in the same breath if interpreted broadly.
With generative AI, the paradigm has changed and it needs a rethink. That was my original point. We don't have a clear answer yet.
If AI follows the digital copyright model, what'll happen is that interest groups with true legal muscle behind them (music, movie industry, etc.) will enforce their interests with draconian laws, and everybody else will be left to fend for themselves.
https://www.microsoft.com/en-us/software-download/windows11/
This is an attitude that has run in Microsoft's veins since the start; they've always been scofflaws, they've always been in the business of stealing other people's inventions. For Microsoft, laws are for protecting their moat.
Frankly, I'm astonished that Suleyman said this on the record.
The copyright issue only comes up if you publish the output of your model. But if the AI is (somehow) clever enough to never reproduce the source material in any way that counts as copying for the purposes of copyright, then there's no copyright problem making it available to the public.
Some artists assume they have more rights than they really do and that other people aren't even allowed to mimic their style.
Just because AI threatens to produce equally valuable material for much lower cost doesn't mean that's wrong. It just means those people got out-competed in the market. That's what the free market is for - to serve the consumers, not the producers.
The reason everyone is irritated at the tech companies is not so much because they made it cheaper to do these things, but because they seem to have such a smarmy, slimy attitude about the whole thing - these comments by the Microsoft CEO of AI as a prime example. They could have just trained their models on old public domain content, Wikipedia, or a million other things, but instead they want to justify raiding content that was shared for free under different expectations. This is only going to result in a locked-down, paywalled web that is pretty much the opposite of 90s net values.
For one, if you serve only the consumers at the detriment of the producers, soon you’ll have no producers left (or, more likely, a handful of extremely powerful ones which will stay stale as they have no reason to innovate).
For another, you appear to be arguing as if the “free market” is an unambiguous positive, when it very much is not. It doesn’t even adequately serve consumers, it serves businesses. The free market is what makes companies knowingly produce and sell harmful substances, like cigarettes or lead paint.
I'm not arguing free market is all positive, just pointing out that the whole reason people see free markets as a good thing is because they benefit people (consumers). They have downsides too but those aren't the reason they're popular. Nobody wants free markets so they can get poisoned more easily.
If I use a reversed Jay-Z sample in my music, I will still get sued. This is just a couple of orders of magnitude higher abstraction.
However, if the model really can't regenerate its training data no matter how hard you try, then that would be fine too. I don't think anyone can really guarantee that now so that might be a problem, but it isn't necessarily a problem.
No, this is not true at all. A lot of things are put on the web with plenty of “restrictions on what for”. That’s why licenses exist. Creative Commons being an example of licenses with well defined limits on what you can and cannot do with the content.
Furthermore, a significant portion of content on the web was put there without permission: think piracy or revenge porn. AI systems have been trained on tons of pirated books.
> Since it's on the web, there really is an implied permission to use it privately for whatever you want.
There really is not. It’s just that if you’re using it privately then by definition no one else knows and thus no one can complain. It doesn’t mean that the creator condones whatever you’re doing.
> Some artists assume they have more rights than they really do and that other people aren't even allowed to mimic their style.
We’re not talking about “other people”. Another person mimicking your style is, to an extent, harmless because they’re also making a time investment to replicate your style. With the AI tools you can reproduce hundreds of knock offs in minutes.
We have to stop with this comparison of “what these AI systems are doing is fine because humans could also do it by hand”. It is not the same thing. Scale matters.
When something has a stated license, that only increases your rights to use it; the default is nothing allowed except fair use.
Copyright literally restricts using the material. If you buy a book, you cannot sell tickets to an event where you read it out loud.
Nope, copyright applies just like it always does.
Make sure to exclude an post with "MVP" in there.
Sheer hypocrisy.
That is exactly what you should be doing to call them out on their bullshit. From the article:
> "I think that with respect to content that is already on the open web, the social contract of that content since the 1990s has been it is fair use," he opined. "Anyone can copy it, recreate with it, reproduce with it. That has been freeware, if you like. That's been the understanding."
Which means that from their logic you can just copy their content and reproduce it.
> "There's a separate category where a website or publisher or news organization had explicitly said, 'do not scrape or crawl me for any other reason than indexing me,' so that other people can find that content," he explained. "But that's the gray area. And I think that's going to work its way through the courts."
Microsoft wants rules for thee, not for me.
Some more discussion: https://news.ycombinator.com/item?id=40826588
This argument cannot possibly hold in any court. This has not been the 'contract'. I cannot reproduce the content of a newspapers online outlet, I cannot reproduce the art of another artist on Instagram, I cannot reproduce someones Youtube video without permission. This same thing sparked the whole fair-use debate some years ago.
The exceptions to these rules have always existed in limbos of regulatory grey areas and are being discussed for decades now.
This guy is still living in the Napster-era apparently and the amount of gaslighting Microsoft, OpenAI, Google etc. perform right now to freeload on data is presumptuous.
What a joke of a person. I hope court will roll over them harshly and explain world doesn’t work like that and just because you can do something it doesn’t make it right.
Read a history book.