AI Has Created a Battle over Web Crawling
spectrum.ieee.org
spectrum.ieee.org
Get the data, make embeddings, categorize and tag the data, then do search and have LLM summarize findings.
I have never done anything serious, but have always tried to keep my hit rate fairly modest.
*Except YouTube, I'll yt-dl an entire playlist, no probs*
I am not looking for ways of circumventing DDOS, more what do people consider a reasonable rate without putting undue burden on a host?
Can I pat myself on the back for only ever having 1 browser tab open at a time? Doesn't my minimalist take on browsing the web "keep my hit rate low" - I think I deserve a tax credit for helping the environment too, wdyt?
Would it make you feel better if I said the pages I'm scraping are also hosted on AWS? That is - I'm effectively paying for the data by paying for lambdas. Or is it the poor hardware itself you worry for?
I'm also concerned that Free and open APIs might become a thing of the past as more "AI-driven" web scrapers/crawlers begin to overwhelm them.
Or a list of IP addresses or user agents or other identifying info known to be used by AI crawlers?
Ultimately pointless but just curious.
I wonder if there are crowd-sourced iptables instead. Similar to how ad filters are maintained.
There are probably tons of those crawlers and with intent to later launder that data, probably using LLMs intended to change the content just enough to be outside of copyright zone.
I welcome all useragents and have been delighted to see openai, anthropic, microsoft, and other AI related crawlers mirroring my public website(s). Just the same as when I see googlebot or bingbot or firefox using humans. For information that I don't want to be public I simply don't put it on public websites.
And there is no middle ground?
I want others to read it but not to sell it?
The first couple decades of the web were build on an implicit promise: publishers could put their content out on freely accessible websites, and search engines would direct traffic to those sites. It was a mutually beneficial arrangement: publishers got the benefit of traffic that they could monetize in different ways if they saw fit (even hosting ads from the search engines a la AdSense), and Google etc. could earn mountains in AdWords revenue. I'm not saying it was always a fair tradeoff on both sides, but both sides had an incentive to share content openly.
AI breaks that model. Publishers create all the content, but then the big search engines and AI companies can answer user questions without giving any reference at all to the original sites (companies like Google are providing source citations, but I guarantee the click-through rates go WAY down). This breakdown in the web's economic model has been happening for a while even before AI (e.g. with Google hosting more and more "informational content" directly on SERP pages - see https://news.ycombinator.com/item?id=24105465), but with AI the breakdown is really complete.
So no, I don't fault people at all who don't want all the toils of their labor to be sucked up by trillion dollar megacorps.
Alternative systems, like paywalls with prominent terms and conditions, exist to protect copyrighted content. They intentionally avoid them. So, they’re using a system designed for unrestricted distribution to publish content they hope people will use in restricted ways. I think the default, reasonable expectation should be that anyone can scrape public data for at least personal use.
HTTP is not designed for unrestricted distribution - if that were the case we wouldn't have the HTTP status codes 401 (Unauthorized), 402 (Payment required) and 403 (Forbidden).
But there is currently a lot of content that is published openly solely because the pre-AI economic model of the web made it viable. That economic model is now going away, so you'll see a lot more content put behind paywalls (or just not published at all) because AI means that it no longer makes sense for some publishers to spend time and money to produce content they can't get a return on.
Considering my blog's niche focus on A/V production with Linux, I can only imagine how much more frequent their crawling would be on more popular websites.
Screw that. Research is about ethical consent. Also this is very much "well everyone should let me (a researcher) access whatever I like"
Obviously most of the 'AI researchers' right now are not altruistic, but it is possible to take the position that advancing AI will be sufficiently valuable to society that it overrides corporate preferences against bulk scraping.
If you don’t want people using your images, consider data poisoning techniques to punish systems that train on them.
https://www.lakera.ai/blog/training-data-poisoning https://www.technologyreview.com/2023/10/23/1082189/data-poi...
I would apply those adjectives to the companies (or researchers out for accolades and citations) who attempt to build their systems based on the non-consensual data extraction in the first place.
To not only deprive systems ultimately benefitting humanity ...
By and large, these systems are built to finance the lush early retirements of the founders, investors and high-level ICs of these companies -- not to benefit humanity. Whether they even benefit humanity as a side effect is very much open to question.
If these companies won't pay, or even condescend to ask permission for access - fuck 'em.
This is like saying if a publishing company takes excerpts from your work, and builds products from it -- without your permission, and of course without paying you royalties of any kind -- they have benefitted without taking anything away from you.
The tech companies call this "learning" of course, but that's just subterfuge.
This is clearly not selfishness. Selfishness is believing you're entitled to other people's work for free, which is the implicit assumption made by people claiming that they should be able to train their models on your work.
Selfishness is trying to take other people's work and turn it onto an AI model without compensating them on their turns. This demonstrates extreme entitlement, lack of a consistent moral system, and a fundamental ignorance of basic economics.
> no differently from a human artist
Yes, humans are objectively different from computers, including neural networks. Humans are sentient, and existing AI is not. It's not even accurate to say that humans have "neural nets" in their brains, because human cognition not understood.
The AI poisoning technique is very moral - if someone doesn't explicitly contact me and ask me for permission to train on my data, then they absolutely deserve a corrupt model. I think I'll do this.
> Entitlement is expecting a royalty from someone who saw your publicly posted art and learned anatomy from it.
As I already pointed out, it's extremely clear that AI are not humans nor are remotely comparable to them (objectively from both a physical and a functional perspective, and morally from the perspective of the vast majority of people holding on to wildly different moral systems, both inconsistent and consistent), so this is completely irrelevant, and additionally not what I claimed (also known as a "strawman fallacy").
Additionally, it is objectively not entitlement to expect payment from someone who consumes your content. If someone posts a learning resource as a paid course on a platform like Udemy, then paying them is morally required - you have to obey the license that the creator makes the content accessible under. Furthermore, the vast majority of the content publicly posted on the internet is not done so under the assumption that it will be used for AI training data - if someone publishes their work for other humans to view for free, that does not extend to a license for AI developers to use it to train for free, which is exactly what the article is about.
> Selfishness is attempting to sabotage it so the learner is mislead into learning wrong anatomy.
See above.
> Lack of a consistent moral system is believing these are variably wrong or right to do when it's a meat neural net or a digital neural net in question.
Objectively incorrect. I can describe a moral system that's non-arbitrary and consistent and perfectly delineates the boundaries between humans and artificial intelligences, and the rights ascribed to both. You cannot. (if you think you can, you're welcome to put it here, and I'll show you why it's inconsistent and/or arbitrary)
Further demonstrating the issues with reading comprehension, you missed/ignored the statement that I made that "It's not even accurate to say that humans have "neural nets" in their brains, because human cognition [is] not understood." which neuters your claims.
> And finally, thinking any such loom-stomping tantrum
The only person making emotionally manipulative fallacies is you. I've articulated my points logically - you've made multiple fallacies, logical mistakes, and your entire first comment was emotional pleading without a shred of logic or reason.
> will meaningfully halt progress and let you keep art creation in exactly the same state as the past
Yet another strawman argument and/or reading comprehension failure. I never claimed that it was possible, necessary, or desirable to keep art creation in the same state. You really didn't read my comment before replying.
> is a fundamental ignorance of economics
...and so this isn't valid. However, I can point out the mistake that you initially made - you believe it's economically viable for those training AI to steal the work of artists and use it to replace them. It isn't.
Your entire comment reads like an AI response that attempted to mirror the structure of mine without any of the understanding, or the ability to make coherent arguments.
However, any moves will increase the cost of training on un-licensed internet data. This will shift the balance slightly closer towards AI companies licensing data as opposed to using scraped data. Emphasis on slightly.
I wonder if “affirming the consequent” is still ok (not that you’ve done so, your post just brought it to mind).
If publicly available content can be used in training sets, any one model willing to use them will have an advantage. Is it reasonable to assume then that all publicly available content is in the most popular training sets, or is that falling into the affirming the consequent fallacy?
>In February 2022, Sama and OpenAI’s relationship briefly deepened, only to falter. That month, Sama began pilot work for a separate project for OpenAI: collecting sexual and violent images—some of them illegal under U.S. law—to deliver to OpenAI. The work of labeling images appears to be unrelated to ChatGPT. In a statement, an OpenAI spokesperson did not specify the purpose of the images the company sought from Sama, but said labeling harmful images was “a necessary step” in making its AI tools safer. (OpenAI also builds image-generation technology.) In February, according to one billing document reviewed by TIME, Sama delivered OpenAI a sample batch of 1,400 images. Some of those images were categorized as “C4”—OpenAI’s internal label denoting child sexual abuse—according to the document. Also included in the batch were “C3” images (including bestiality, rape, and sexual slavery,) and “V3” images depicting graphic detail of death, violence or serious physical injury, according to the billing document. OpenAI paid Sama a total of $787.50 for collecting the images, the document shows.
>Within weeks, Sama had canceled all its work for OpenAI—eight months earlier than agreed in the contracts. The outsourcing company said in a statement that its agreement to collect images for OpenAI did not include any reference to illegal content, and it was only after the work had begun that OpenAI sent “additional instructions” referring to “some illegal categories.” “The East Africa team raised concerns to our executives right away. Sama immediately ended the image classification pilot and gave notice that we would cancel all remaining [projects] with OpenAI,” a Sama spokesperson said.