Lawsuit claims OpenAI stole 'massive amounts of personal data'
businessinsider.com
businessinsider.com
Next up, all content will be only viewable if you are signed in ala twitters change this morning.
The API changes and the ensuing chaos are at least in part because of OpenAI scraping Reddit’s content.
Purely conjecture obviously, but him being on both sides of that, not hard to see how all this becomes more favorable to his company making the moat wider for others.
It seems rather obvious under his perspective to do everything possible to restrict access to all information that could be used to train an LLM.
I understand Sam’s history with ycombinator gives him leniency around here but we seriously need to wake up to the fact that he’s attempting to monopolize the most powerful tool ever created. He talks a good game but if you pay attention to his actions it’s clear he is openly hostile towards the common good which he supposedly promotes.
I see him in the same regard as I see any oligarch vying for public favor by ‘stealing’ billions, donating a few million, and riding off into the sunset being praised by media for how much of the ill gotten gains he’s given back.
Instead they imposed an extremely abrupt new demand for money with any API call.
More importantly, the laws are representations of human feeling and social negotiation. Copyright law grew out of a complex array of beliefs and interests. E.g.: https://en.wikipedia.org/wiki/Statute_of_Anne
It could be that the various "AI" generative models are legal under current law. But it could also be that this is one of those things, like the printing press, that causes people to say, "Hey, that's not right," and change the law to rule out something that was previously legal.
"This is all 'wild west' to me as a European, but somehow I wonder if there is any guns that work to protect yourself by means of self-justice."
I vaguely remember a Stern magazine cover from around 2004, depicting a giant cowboy boot with "GM" decorated with red, white and blue coming down on a crowd of people standing in the shape of the Opel logo and titled something like "Der Wild-West-Method."
It was clearly not a compliment.
(Correction: GM had owned Opel for decades at that point, but was getting more aggressive with how it was run)
In case anyone wants to read more about it - it’s very relevant today.
If that’s all it takes to skirt copyright restrictions Aaron Schwartz should be rolling in his grave right about now.
And courts never overrule prior decisions...
also, hoovering up all available data on the net does not make them any less private
I certainly never said Google could crawl my website, and I would be surprised if the presence of a robots.txt file is written into law as a requirement to prove you didn't want them on your web properties.
I feel like this must have been hashed out in lawsuits already, probably against search crawlers.
If you want more info about the case you should look for the "affaire Bluetouff*.
IANAL and so on.
Publicly accessible =/= You having the right to do whatever you want with something.
The internet is closer to leaving your own journal open at the library or a magazine store, and people going through it. You made it available for everyone to peek at in a public space. Maybe you didn't intent to leave it there, maybe you didn't intend to make it visible, but you did.
IP is still IP even when it’s publicly available.
But this is a nice setup for a straw man argument to distract people from the actual issue.
Why aren't they suing the hosts that allowed the data to be stolen? Where's the notice of security breach?
You sometimes can use things that are available. It’s not an absolute. See: https://www.theregister.com/2022/04/19/scraping_public_data_...
All of the stuff I've put on the web, I put there for actual people to use. Crawlers and scrapers came around, and while I certainly didn't like it or approve of it, it was something I put up with. The only defense against it was to stop putting things on the web entirely, which seemed like an overreaction.
Now, however, the use of that data to train AI (something that I consider actively harmful to society in general and don't want to support in any way), is a degree too far in the pot that's been slowly coming to a boil.
While I do want to give back and contribute to the larger body of work available to people, I want it to be available to actual people. I don't want it to be available for training AI.
I don't see why that's such a ridiculous stance. However, it's not legally possible to both make a work available to the public and prevent that work from being available for other uses. That's a real shame, and that the only protection available to me is to no longer have the works available publicly, it's loss to everybody.
"I don't want it to be available for training AI" is a perfectly reasonable personal preference. We can discuss whether it's selfish or not, whether others agree or not, etc.
But this is a lawsuit. What makes it illegal for OpenAI to use your content? Like, is there some license you've put up on your content that disallows it? Is there anything that's relevant to the case at hand?
As an EU citizen, some random company can't use PII related to me if I don't give consent or revoke my consent. They'll have to remove it or face a huge fine. The law doesn't care about the cost for your company to comply.
> The Children’s Online Privacy Protection Act (COPPA) gives parents control over what information websites can collect from their kids. The COPPA Rule puts additional protections in place and streamlines other procedures that companies covered by the rule need to follow. The COPPA FAQs can help keep your company COPPA compliant. Learn about the COPPA Safe Harbor Program and about organizations the FTC has approved to implement safe harbor programs. You can also get information about ways to get verifiable parental consent– including new methods the Commission has approved – and the process for seeking approval for new methods.
There's one. There are similar laws in Europe.
The AI scraped, utilized, kept, keeps... data on children. Parents did not consent. Children /can't/ consent.
The AI use of such data is not only illegal but could end up with people being jailed over it.
So my question would be "what license entitles OpenAi to use my content"? It could be that as part of the user agreement of the website you gave some rights to the website and they sold it on to OpenAi, could be fair use, could be "does not count as 'more'",... .
Bad analogy: For money laundering we have similar rules. If you suddenly appear with a large amount of money/data, then the obligation is on you to show that it is clean, not on others to show that it is dirty. You can disagree with that (innocent until proven guilty etc.), but it is not clear cut which way around things should be.
As to the legality, that remains unknown until a court makes a ruling. I suspect that this lawsuit will go nowhere, but I'm just speculating along with everyone else.
On the other hand, you can’t prevent people from using your data if you put it out in public without condition (in the form of some acceptable licenses).
As an analogy, if you put a picture of your living room in the sidewalk, you can’t prevent me from looking at it, study it, when I walk by. I may even benefit from it by copying your style or decoration. You may however cover it, with a warning. I’d clearly violate your terms if I still look at it against your will, although I may not violate any laws.
That’s bullshit. Copyright still applies to property available to the public. Maybe you are confused between “available to the public” (access) and “in the public domain” (copyright).
Listening to a song on the radio doesn’t give you the right to copy it and sell it.
You can say the same for other forms of IP. Trademarks, logos, etc are all “available” to the public, but they are still protected IP.
When you post to a public website, you are giving access to that content , and (usually) legally giving rights to the website to publish your content. Those rights don’t transfer to a 3rd party.
PS - these comments are going to be used to train the next GPT aren't they?
When a person goes to Amazon and looks at what is splashed on the page, there are any number of chances that they will click "Buy!" on any one of them, and winning that lottery is Amazon's business; when a crawler does, it does more "looking" with no chance of purchase, and the data is then used to reduce the value of Amazon's lottery game.
I'm not arguing whether or not crawling is or should be legal, simply saying the "it's not theft because it's copying" argument is inadequate to the task.
Amazon is inherently anti-consumer, and uses every single thing it can to advertise to and/or profit off of you. Their 1984-inspired “security” and home automation systems, their app/website, their policies are all meant to take your money. Which is why web scrapers like CamelCamelCamel are good, because with all of this anti-consumer garbage Amazon shoves at you, you have the power to turn the tables and pick up what you need at the price you want.
The only time I’ve heard scrapers/indexers being a problem is Bing was terrorizing someone’s website, so they just banned every Google/Bing/Yahoo IP. A problem that $1.3T companies don’t have.
You can claim that “its not theft because it’s copying” is inadequate, but I would say the same about your own argument. Because there is no good argument against scraping data. It’s the only way to have a free Internet.
the rest of what you wrote is a combination of marxism--why can't society cooperate to meet my needs!?--and laissez-faire capitalism--bastards think they can use technology against me, I'll use it against them--neither of which either is a good way to run an economy.
don't worry, we don't "gotcha" here, we got you, brother!
It can change the availability of future data. I, for one, am altering my posting habits knowing that my data can be scraped into LLMs.
But I digress. This is more about selling an app with locations to soup kitchens. The kitchens may not have explicitly given permission to be used in the app, but their business location is public knowledge and not expected to be hidden.
If you post something to the public internet, you lose privacy ... that's how internet works.
For this we have robots.txt and authentication ... if a site allows you to browse their content, it's free to take, whatever the purpose.
- Indexing vs. Displaying: Search engines primarily index web pages to create a searchable database of information. They do not generally host or display full copyrighted content directly. Instead, search results usually provide brief snippets, page titles, and links that direct users to the original source. This approach aims to respect copyright by driving traffic to the copyright holders' websites.
- Fair Use Considerations: In some cases, search engines may display limited portions of copyrighted content under the fair use doctrine, which allows for the limited use of copyrighted material for purposes such as commentary, criticism, news reporting, or educational purposes. The application of fair use can be subjective and depends on the specific circumstances of each case.
Replace "search engine" with "LLMs", it's (practically) the same.
I'm not okay with things I create (written, photos, etc) being used by these companies in datasets.
You mean like all those credit card and other databases that keep getting left on public endpoints on AWS?
Taking and allowing something to be taken are two separate issues, morally, ethically and legally.
AI can be one of the most powerful forces in the world but it needs to pay content creators. If people stop creating new content for AI to train on then it will get stale.
AI can be the product of our dreams and our passions but we need to make sure those who choose to create and share content to train it are treated fairly and compensated when applicable.
This lawsuit is alleging that OpenAI trained GPT-3 and -4 on inadvertently published information. Web crawlers are very good at finding things you wouldn't expect to be public; there's techniques you can use to, say, abuse Google to search for such things.
Does anyone know if this includes enterprise level versions out of the box. Specifically for Teams as i would assume, maybe incorrectly though that the other would require a plugin of some kind. But Microsoft being mS i feel would be more comfortable putting chat gpt into their base products.
If that IS the case then there will be a big backlash against them.
As a result, GPT-5 creates a brief that, while superficially appears incredibly compelling to OpenAI and its lawyers, is actually specifically tailored to rub the judge the wrong way so he or she rules against OpenAI... all thanks to GPT-5 having had access to enough personal data about the judge to know how to piss them off.
Let’s grab some popcorn and watch…