You sometimes can use things that are available. It’s not an absolute. See: https://www.theregister.com/2022/04/19/scraping_public_data_...
If you post something to the public internet, you lose privacy ... that's how internet works.
For this we have robots.txt and authentication ... if a site allows you to browse their content, it's free to take, whatever the purpose.
- Indexing vs. Displaying: Search engines primarily index web pages to create a searchable database of information. They do not generally host or display full copyrighted content directly. Instead, search results usually provide brief snippets, page titles, and links that direct users to the original source. This approach aims to respect copyright by driving traffic to the copyright holders' websites.
- Fair Use Considerations: In some cases, search engines may display limited portions of copyrighted content under the fair use doctrine, which allows for the limited use of copyrighted material for purposes such as commentary, criticism, news reporting, or educational purposes. The application of fair use can be subjective and depends on the specific circumstances of each case.
Replace "search engine" with "LLMs", it's (practically) the same.
PS - these comments are going to be used to train the next GPT aren't they?
When a person goes to Amazon and looks at what is splashed on the page, there are any number of chances that they will click "Buy!" on any one of them, and winning that lottery is Amazon's business; when a crawler does, it does more "looking" with no chance of purchase, and the data is then used to reduce the value of Amazon's lottery game.
I'm not arguing whether or not crawling is or should be legal, simply saying the "it's not theft because it's copying" argument is inadequate to the task.
Amazon is inherently anti-consumer, and uses every single thing it can to advertise to and/or profit off of you. Their 1984-inspired “security” and home automation systems, their app/website, their policies are all meant to take your money. Which is why web scrapers like CamelCamelCamel are good, because with all of this anti-consumer garbage Amazon shoves at you, you have the power to turn the tables and pick up what you need at the price you want.
The only time I’ve heard scrapers/indexers being a problem is Bing was terrorizing someone’s website, so they just banned every Google/Bing/Yahoo IP. A problem that $1.3T companies don’t have.
You can claim that “its not theft because it’s copying” is inadequate, but I would say the same about your own argument. Because there is no good argument against scraping data. It’s the only way to have a free Internet.
the rest of what you wrote is a combination of marxism--why can't society cooperate to meet my needs!?--and laissez-faire capitalism--bastards think they can use technology against me, I'll use it against them--neither of which either is a good way to run an economy.
don't worry, we don't "gotcha" here, we got you, brother!
It can change the availability of future data. I, for one, am altering my posting habits knowing that my data can be scraped into LLMs.
But I digress. This is more about selling an app with locations to soup kitchens. The kitchens may not have explicitly given permission to be used in the app, but their business location is public knowledge and not expected to be hidden.
All of the stuff I've put on the web, I put there for actual people to use. Crawlers and scrapers came around, and while I certainly didn't like it or approve of it, it was something I put up with. The only defense against it was to stop putting things on the web entirely, which seemed like an overreaction.
Now, however, the use of that data to train AI (something that I consider actively harmful to society in general and don't want to support in any way), is a degree too far in the pot that's been slowly coming to a boil.
While I do want to give back and contribute to the larger body of work available to people, I want it to be available to actual people. I don't want it to be available for training AI.
I don't see why that's such a ridiculous stance. However, it's not legally possible to both make a work available to the public and prevent that work from being available for other uses. That's a real shame, and that the only protection available to me is to no longer have the works available publicly, it's loss to everybody.
"I don't want it to be available for training AI" is a perfectly reasonable personal preference. We can discuss whether it's selfish or not, whether others agree or not, etc.
But this is a lawsuit. What makes it illegal for OpenAI to use your content? Like, is there some license you've put up on your content that disallows it? Is there anything that's relevant to the case at hand?
> The Children’s Online Privacy Protection Act (COPPA) gives parents control over what information websites can collect from their kids. The COPPA Rule puts additional protections in place and streamlines other procedures that companies covered by the rule need to follow. The COPPA FAQs can help keep your company COPPA compliant. Learn about the COPPA Safe Harbor Program and about organizations the FTC has approved to implement safe harbor programs. You can also get information about ways to get verifiable parental consent– including new methods the Commission has approved – and the process for seeking approval for new methods.
There's one. There are similar laws in Europe.
The AI scraped, utilized, kept, keeps... data on children. Parents did not consent. Children /can't/ consent.
The AI use of such data is not only illegal but could end up with people being jailed over it.
As to the legality, that remains unknown until a court makes a ruling. I suspect that this lawsuit will go nowhere, but I'm just speculating along with everyone else.
On the other hand, you can’t prevent people from using your data if you put it out in public without condition (in the form of some acceptable licenses).
As an analogy, if you put a picture of your living room in the sidewalk, you can’t prevent me from looking at it, study it, when I walk by. I may even benefit from it by copying your style or decoration. You may however cover it, with a warning. I’d clearly violate your terms if I still look at it against your will, although I may not violate any laws.
As an EU citizen, some random company can't use PII related to me if I don't give consent or revoke my consent. They'll have to remove it or face a huge fine. The law doesn't care about the cost for your company to comply.
So my question would be "what license entitles OpenAi to use my content"? It could be that as part of the user agreement of the website you gave some rights to the website and they sold it on to OpenAi, could be fair use, could be "does not count as 'more'",... .
Bad analogy: For money laundering we have similar rules. If you suddenly appear with a large amount of money/data, then the obligation is on you to show that it is clean, not on others to show that it is dirty. You can disagree with that (innocent until proven guilty etc.), but it is not clear cut which way around things should be.
That’s bullshit. Copyright still applies to property available to the public. Maybe you are confused between “available to the public” (access) and “in the public domain” (copyright).
Listening to a song on the radio doesn’t give you the right to copy it and sell it.
You can say the same for other forms of IP. Trademarks, logos, etc are all “available” to the public, but they are still protected IP.
When you post to a public website, you are giving access to that content , and (usually) legally giving rights to the website to publish your content. Those rights don’t transfer to a 3rd party.
You mean like all those credit card and other databases that keep getting left on public endpoints on AWS?
I'm not okay with things I create (written, photos, etc) being used by these companies in datasets.