Google says it'll scrape everything online for AI
gizmodo.com
gizmodo.com
Time to make a webring of Markov generated garbage to feed the bots...
https://developers.google.com/search/docs/crawling-indexing/...
99% of the time people's complaints about "permission" and "information accidentally made public" are the same arguments they have been using about search engines for years. Most of them have court cases even.
If it is legal to do copyright laundering with AI, then copyright is dead. What's the point of writing a blog if people read it somewhere else, through an AI?
As if scraping was a new and unusual thing. Google and other have always been doing a lot more than just creating an index.
The opponents should attack the new stuff, not the other stuff that people have been doing for years.
Source? That makes no sense to me. The problem is what being done with the copyrighted material, not the fact that an algorithm loaded it.
Are you sure they complain about the scraping itself, or do you want them to complain about that, so that you can say that they make no sense?
"The defendants are alleged to have conducted widespread web-scraping campaigns, violating various platforms' terms of service as well as state and federal privacy laws..."
https://www.iotworldtoday.com/security/openai-faces-3b-lawsu...
But reasonably, I would assume that the humans who see a problem with AI are not bothered by the fact that some text appears in some memory somewhere. They are bothered by the use of that text, which they believe goes against their copyright.
2. Put it behind a click-wrap agreement that provides that it is provided to the user in exchange for an agreement not to use it to train AI.
3. For images, using on of the AI-poisoning techniques and publicizing that fact may encourage people to exclude your content from training sets; OTOH, it may backfire, especially against people looking specifically to defeat the AI poisoning technique in use.
Otherwise, those operating under the theory that training AI is Fair Use will have no reason not you use the content for training AI.
Such a thing is only as good as your ability to enforce it, though. Are you in a financial situation that allows you to afford to sue Google? I'm not.
Having a price list doesn't change the fact that if someone ignores your demands, your only recourse is to sue them. If they have a bigger warchest than you, and are sufficiently motivated, they can drag the whole thing out until you run out of money and can't pursue the lawsuit anymore.
This is why such an approach is of limited value in terms of protecting yourself against companies that are sitting on a ton of money.
In memory of Don Lancaster, he made a point about this sort of thing: many companies would rather spend $100,000 on legal fees than stoop to paying you a $10,000 fee or royalty.
If your threat model is “those only concerned with what they can be forcibly compelled not to do, after evaluating the capacities of each content supplier to resist”, things are different.
Which category Google falls into on this is a debate I’m not interested in engaging in.
So, for now, all of my public websites are no longer publicly available until/unless I can find a solution.
Ultimately the decision to check they're not utilizing copyright material rests with the human commercializing it.
Now my opinion is that humans are not machines, and therefore it makes sense to not apply the same laws. To me, it seems natural to think that a human will never remember every single word of a book they read. Whereas a machine can totally do it.
> It's how the human uses it's output is what is important
I would take it from the other direction: the AI should be able to prove that it "read" the words like a human (i.e. with a "bad" memory and everything that goes with it).
These are probabilistic models,it's basically a large algorithm. They're not copy pasting a database verbatim. They don't have access to it, only a list of parameters (sort of the algorithm weightings) and the list of tokens (translated words).
With OpenAI's API you can even see the certainty/probability associated with the next word it quotes. We know for definite whether a model is doing this by looking at it's source code. Humans do also come up with the same ideas from time to time, without any external input.
> A computer can remember to do that only if it's programmed to do so.
Not sure what your point is there. So your belief is that if the programmer did not explicitly put a specific intent into the code, then the computer cannot do anything consistent with that intent? Like nothing unexpected can ever happen with computers, because they are strictly limited to what the programmer explicitly wanted them to do?
Given these models have billions (hundreds of billions) of parameters, the odds of it exactly copying your code is less than 100%.
Academics when publishing run through plagiarism checkers as it happens even if they're 100% sure they haven't consulted other thesis'. (Usually 20%-30% is acceptable).
Some coders run their code through checkers to make sure it's not copyrighted. The onus similarly would be on whoever is using the LLM output to check it before monetizing it. It's just a tool after all.
What I am concerned is big conglomerates munching on it and reselling it at a high price.
A distinction without a difference.
> More akin to a search engine database.
The point of a search engine is to direct you towards the content. The point of LLMs is to keep you within the walls of the LLM.
You mean like a search engine scraping content and putting ads by results?
1. https://developers.google.com/search/docs/crawling-indexing/...
If AI could honor those licences (and all the others), I wouldn't see a problem. But they can't.
There are also more sophisticated attacks like even single pixel attacks that one can place in his content.
I expect if I put enough counter AI attacks in my websites media/text they will notice that and put me on their exclusion list.
Well they are first to pull asshole move so won’t feel guilty.
Have different service levels. Authenticated users get the real thing (rate limited), anyone else gets a dynamic Markov-generated version with subtle errors deliberately introduced.
Basically, a company called HiQ was scraping public LinkedIn data to sell a service where they would inform companies if their employees were looking for a job on LinkedIn -- because LinkedIn will hide info from your employer for obvious reasons. LinkedIn sent cease-and-desist, then later sued by lost.
https://www.fbm.com/publications/what-recent-rulings-in-hiq-...
Are you writing open source code? Would you still make it open source if people never saw it as "your" code, but instead paid an AI company to access your knowledge?
Are you getting those strange internet points by answering StackOverflow questions? Will you still do it if people stop seeing the answers you redacted (and voting for them)?
If it goes there, I will stop blogging, stackoverflow and open sourcing. And everything else that I do for free, just for the satisfaction of getting a "thank you" from a human from time to time. If AI wants to make me disappear, I will disappear.
Don't you see a difference between 1) a human reading your blog post, understanding it, making their own opinion, and then sharing it under their name and 2) a machine taking your blog post as an input and distributing a modified version of it?
Let's take two examples:
First example: I create a blog that copy-pastes your posts to the word, but I don't give any kind of attribution, and I get money from it. Would you be fine with that, or would you consider that I should not be able to do it, because I am basically selling your copyrighted work without authorization?
Second example: I do exactly the same as above, but I have a small script that will replace some words with synonyms. Is it now better than the previous example?
Where do you put the limit? If the script is so good that you can't recognize your post anymore, then you find it okay? But that's still an automated script that just duplicates your blog and benefits from it, right?
https://www.fbm.com/publications/what-recent-rulings-in-hiq-...
For example, we may collect information that’s publicly available online or from other public sources to help train Google’s languageAI models and build *products* and features like Google Translate, Bard, and Cloud AI capabilities. Or, if your business’s information appears on a website, we may index and display it on Google services.
https://policies.google.com/privacy/archive/20221215-2023070...
For situations like this precision matters and guesses don't count.
And if they can use leaked proprietary code, why couldn't they use basically everything they have access to, like your emails?
In my opinion, it should all be forbidden, but you will find plenty of voices online that think AI should be able to (ab)use open source code, so...
We're fully in agreement there.
If you do not want your “intellectual property” taken and used by others, don’t give it away for free.