Pulling my site from Google over AI training
tracydurnell.com
tracydurnell.com
It is reading and learning. A person would read and learn.
This has no bearing on plagiarism or copyright. A work can be considered plagiarized or to breach copyright if the author hasn't even seen or come across the copyrighted/published work.
This is no different. I can write some code and use it, subconsciously referencing a work.
If I don't check my written work and put it out there, someone might have a claim against me. If I don't check the machine generated work and put it out there, someone might have a claim against me.
OpenAI,Meta,et al are providing the model, basically a regression model or tool. I'm adding the variables or secret sauce that makes it output that set of data in that specific order not them. It'd be like suing Parker for making the pen.
Existing conventions around “learning” are built on assumptions of human scale, and the expected consequences thereof.
I can’t understand why one would expect people to go “oh it’s technically ‘learning’ I guess I’ll ignore all the consequences that weren’t present when it was just humans”.
It's also important to note these laws were always intended to strike a fair balance between the copyright owner and the good of society as a whole. Copyright is not an end in itself.
Although everyone probably wishes otherwise, there isn't really a hard objective line for determining if something is infringing or not. And that's actually a good thing! But it means that, in the end, it's up to a judge to weigh a bunch of different, rather fuzzy, factors.
We could very well distinguish between machine learning and human learning although they're no different from each other in principle.
And although, we can't say that ML is plagiarism, at scale it breaches our other moral principles like privacy and individual identity.
Let's say in 10 years from now Google'll train ML on (close to) all the data available in the world, be it text, visual or audio. Would they be able to prompt this AI: "You are George Wilkinson from Colorado ** st. 23". How close will it be from real George on whose data it was trained on?
We can't tune our human learning to this level of precision, but it is only a matter of time for the machine.
We absolutely do, and expectations of what's possible are completely different.
It's a good thing, but at the moment language copyrights is wielded by corporations and ignored by corporations. Either we make copyrights like maths and discard them, or we hold corporation's to the same standards that they hold us to.
Do you really think that of I train my model on Mickey Mouse and then go ahead and generate "not Mickey" that Disney won't try and sue the crap out of me? Almost all these complaints boil down to different standards for them versus us.
A search engine that exclusively indexes noindex sites (you can use other sites while spidering) and builds an LLM model with the results.
I'm not going to be apologetic about it either, since the same people who think they're invulnerable also tend to espouse sadistic glee over the impending immiseration of millions due to these developments... just practicing what they preach, after all.
then be competing with it for the rest of their lives as it slowly reduces their labour potential to zero
Just today I googled (and duck duck go'd?) alternatives for Discord (because reasons). Entire search results page was "X top alternatives to Discord." It was all blog posty kind of stuff with an "author".
And like 90% of it was written by indian and african sounding names. These were clearly "content farms" with low paid labour and bad grammar, or just authors with nothing better to do than write Yet Another Blog Post about Top Discord Alternatives. Sure, they weren't generated, but the fact that a human was involved in creating something crappy doesn't make it better or unique.
What I was actually looking for was unique content. Either an actual curated list of alternatives (NOT a blog post they update every year). Or an extract from a book where someone posted fiction about a fictional Discord user that meets aliens. Or comments in a forum, or a link to a song-lyrics website for a Weird Al parody song about discord, a website dedicated to expounding the virtues cutting the discord cord, a link to a PDF where someone saved a IRC chat server's logs about a person switching from discord to IRC, or an "IRC-MF do you speak it" crass website, or something. Anything but a damn content blog post by some third-world content creator or hipster-blog-poster from the 1st world.
What I got was garbage. Human-level garbage. Garbage that across hundreds of thousands of websites basically took a piece of content and expanded it with every known combination of words, sentences, and mini-stories and pasted it on a stupid blog post with an author.
And this garbage is what this AI is training on so we can have content farms make more copies of itself with more variations and in different languages now, all so we can pay Google et al attention-coins to magically sift through all that garbage and present us with something a little less garbage-y for us to consume.
Side note, I have friends that crawled a massive amount of the internet over several months for their own purposes.. at this point it's probably impossible to exclude your site since tons of other people probably link to your site if it's at all of value.
Instead, I'd just keep an eye on my access logs and block obvious crawlers when I saw them.
Is incredible.
I'm curious, this is the first time I've heard of a webring, I'd like to learn more about these alternate discovery routes. Anyone have any concrete experience or recommendations to share?
After that, sites also stop linking to each other.
Sure search engines got better, and then they got worse..
Many of the things 'web rings' promoted were 'similar sites'.
Yet it was when google become big that it destroyed web rings and blog rolls (which were similar to web rings, but often included a variety of sites not just similars) -
It became known that google penalized people for linking to other sites, linking to 'bad neighborhoods' - they would sometimes call out sites publicly for giving 'link juice' to others for profit..
Web rings disappeared and blog rolls as well, because google.
Not because search engines got better, but because google threatened to penalize you for linking out.
It feels like at some point in the last 5 years we crossed some threshold where search engines are so optimized to the commercial space and selling junk that obscure searches yield no useful information. Any keyword which is so unfortunate to overlap in the commercial product space will just dominate your results and make them mostly useless. Even my advanced google fu to tack on certain phrases and other bits of language to narrow things down and formerly gave very focused results but now feels like google is atrophying in that domain.
It doesn't help that our culture has become so word overload friendly where instead of creating new words and phonetic combinations existing words are used instead and now require disambiguation to say not the product kind of this word.
This must be illegal, but how are all the little bloggers going to oppose it?
But even that requires the permission of the copyright holder. Nobody is required to accept an infringing use of their work in exchange for royalties.
Which doesn't really mean anything.
At this point, the only defense I can think of is to not make the content publicly available. Which is what I've done.
My point being: it's exceptionally hard to create laws that deny precisely what you don't want and allow precisely what you want, without quickly getting into details that bring the entire law's assumptions into question. Here being "because an ML training is not a person, it has no right to scan the web".
Also on a whole different scale and instead of supplementing the web content it’s devaluing it to a degree.
Basically you're assuming the agenda of the operator, saying "that's bad an shouldn't be allowed". But I see the web- except for things specifically labelled with standard copyright disclaimers- as effectively a large corpus of publicly available data, "in the market square for all to see".
I don't think training on copyrighted stuff will be ever banned, but we need to figure out how much they can be allowed to generate based on that. Eventually new models will just pop up with more carefully curated data anyway.
I don't see why AI developers should be expected to think otherwise and worsen their training data over this.
The point is that Google has a certain market position that makes it very different when they "recite the vague plot of a novel or a fact they learned". The point of competition law is to "distort" free market capitalism for the betterment of society. This is one of those cases where practical considerations trump information idealism. The quality of information on the internet will go down if we stop rewarding original publishers.
It could be illegal if the AI reproduces vast portions of it. If you could ask the LLM over a course of prompts to generate a significant portion of the content (as the copyright law defines it), then yes.
As long as the AI isn't reproducing it, then I am not sure if it would count.
From a US copyright law point of view, this is most likely correct. Copyright law doesn't prevent you from ingesting copyrighted works, it prevents you from distributing them.
There is also a great deal of existing case law about how different a work has to be before it's not infringing from another work anymore. There are existing rules of thumb judges go by when trying to determine if infringement occurred. They include things like the amount of difference in expression, the quantity, whether or not it's incidental, etc.
And that's not even getting into the question of fair use -- which is a whole other kettle of fish.
I suspect that the courts will deal with these issues the way that they've always dealt with these issues: on a case-by-case basis.
P.S. Permission is not granted to downvote my comment!
It's not clear to me that it should be illegal.
[0] https://tracydurnell.com/2023/07/07/the-next-big-theft/
[1] https://www.tumblr.com/nedroidcomics/41879001445/the-interne...
Got it.
> In addition, I created a robots.txt file to tell “law abiding” bots what they’re not allowed to look at. I ought to have done this before but kind of assumed it came with my WordPress install (Nope.)
> I specifically want to deter my website being used for training LLMs, so I blocked Common Crawl.
Instead of blocking, it would be neater to present and alternative version to the crawlers (like many paywalled sites already to for SEO) that's full of dynamically generated LLM-generated garbage. That'll help the LLMs poison themselves.
My websites have been closed to the public since shortly after the release of ChatGPT, but I've been considering opening them up again, sort of. The not-logged-in experience being full of dynamically generated LLM poison as you suggest -- for everybody rather than trying to single out crawlers -- and you have to log in to get to the real contents of the site.
You have a legal right to do this (assuming the LLM bullshit isn't libel) but I don't see how it could be considered a moral act.
Yes, although I've been blocking googlebot for years, so nobody will get to my sites through google anyway.
> I don't see how it could be considered a moral act.
I'm curious about this -- why do you think this is in any way an immoral act?
If a naïve human comes across the site, they'll quickly realize that it's not useful and move on. No harm done. How does morality enter into it?
Would your moral objections be eased if the first line on the page is something like "this page is full of machine-generated nonsense. Please ignore it"?
I do not share your confidence.
>Would your moral objections be eased if the first line on the page is something like "this page is full of machine-generated nonsense. Please ignore it"?
That would certainly help.
That seems reasonable enough. If I do this, I'll include such a disclaimer.
IIUC, if a site presents different (view of) content to the crawler than users, the site can get de-indexed.
Assuming there's a difference, anyway. I suspect that with Google and Bing, there isn't.