Anyone got a contact at OpenAI. They have a spider problem
mailman.nanog.org
mailman.nanog.org
https://www.alignmentforum.org/posts/8viQEp8KBg2QSW4Yc/solid...
https://www.lesswrong.com/posts/LAxAmooK4uDfWmbep/anomalous-...
Vocabulary isn't infinite, and GPT-3 reportedly had only 50,257 distinct tokens in its vocabulary. It does make me wonder - it's certainly not a linear relationship, but given the number of inferences run every day on GPT-3 while it was the flagship model, the incremental electricity cost of these Redditors' niche hobby, vs. having allocated those slots in the vocabulary to actually common substrings in real-world text and thus reducing average input token count, might have been measurable.
It would be hilarious if the subtitle on OP's site, "IECC ChurnWare 0.3," became a token in GPT-5 :)
In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence indicates that.
Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".
Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.
This way is logically most efficient to work and involve the least communication.
() In a lecture by the mathematician & author Sarah Hart.
Logically speaking, the second bar tender could have thought to himself "no I don't want any beer, but one of these two other guys may want to double fist" and so there is really no way for the third logician to answer in the affirmative.
Fully agreed that "hallucination" is a bonkers word for it — sensational and melodramatic. But few people know what a confabulation is, and moreover it's an overly complex way to describe the phenomenon.
The LLM is making something up. It's a fabrication.
It's not fanciful; it's not spooky; it's mundane, as it should be.
I believe that you can also cause something a bit like a transient dysphasia by giving them bad inputs as well, so there is that on the language production side. However there's still nothing that pertains to the experience aspects central to what hallucinations actually are.
The problem with "hallucination" and "confabulation" is that they both imply a consciousness.
It's still just probabilistic babble. We were doing this with markov chains.
I don’t think confabulate matches as well as it implies confusion or mixture of different ideas.
ChatGPT isn’t confused, it’s making things up. It’s trying to bullshit as best it can in hope that what it makes up convinces its user.
But given how often toddlers are "taken care of" by planting them in front of youtube :|
Even if you don't show children videos, but want to play some music, YouTube is still the least-hassle, least-bullshit music stream player (arguably still it's main use for adults, too). Ain't anyone got time to deal with Spotify's ever more broken app. And this is the limit of technical skill of almost all parents. They can't exactly run SponsorBlock in YouTube's mobile app (and paid YouTube doesn't help here either, surprise surprise).
Not making excuses (though I'm not really blaming parents for this) - just saying how things actually are.
At some point fairly recently they added a "I don't know the answer" button to the email, but it's much less prominent than the main call-to-action.
A lot more of LLM hallucination is it getting the context confused. I was able to get GPT4 to hallucinate easily with questions related to the distance from one planet to another, since most distances on the internet are from the sun to individual planets, and the distances between planets varies significantly based on their locations in cycle. These are probably slightly harder to fix.
I've noticed that while this can help to prevent hallucinations, it can also cause it to go way too far in the other direction and start telling you it doesn't know for all kinds of questions it really can answer.
This isn't true. There are many contexts where it is true but it doesn't actually generalize they way you say it does.
There are plenty of cases where experts in a non-one-on-one context will express a lack of knowledge. Sometimes this will be as part of making point about the broader epistemic state of the group, sometimes it will be simply to clarify the epistemic state of the speaker.
Because expecting a behaviour, like knowing you don't know, that isn't represented in the training set is silly.
Kids make stuff up at first, then we correct them - so they have a way to learn not to.
The problem is that curating data is slow and expensive and downloading the entire web is fast and cheap.
See also https://en.wikipedia.org/wiki/Cyc
Maybe if you trained a small base model to know it doesn't know in general and THEN trained it on the entire web with embedded not-knowing preserving training examples, it would work?
https://www.quantamagazine.org/tiny-language-models-thrive-w...
This is not a brain. The best analogy is an English major.
They are good at language, not reasoning.
Humans see language and think reason. It seems we can’t separate the two.
Google has one of these already, with an LLM that was trained on nothing but weather data and so can only give weather-data-prediction responses.
The 'knowing it doesn't know things' part is much harder to get reliable, though.
What I'm pushing at is not that this linguistic ability naturally leads to the LLM behavior we're seeing and calling "hallucinating", just that LLMs may capture some of how humans process language, differentiate semantics, recall terms, etc, but without the mechanisms that enable rationally grappling with the resulting semantics and propositional (in)coherency that are fetched or generated.
I can't say this is very surprising—most of us seem to have thought processes that involve generating and rejecting thoughts when we e.g. "brainstorm" or engage in careful articulation that we haven't even figured out how to formally model with a chatbot capable of generating a single "thought", but I'm guessing if we want chatbots to keep their ability to generate things creatively there will always be tension with potentially generating factual claims, erm, creatively. Further evidence is anecdotal observations that some people seem to have wildly different thresholds for the propositional coherence they can spot—perhaps one might be inclined to correlate the complexity with which one can engage in spotting (in)coherence with "intelligence", if one considers that a meaningful term.
I don't think Wittgenstein would agree, first of all, that there is a "natural logic" to language. At least in the PI, that kind of entity--"the natural logic of language"--is precisely the kind of weird and imprecise use of language he is trying to expose. Even more, to say that such a logic "allows" for anything (like metaphors) feels like a very very strange thing for Wittgenstein to assert. He would ask "what do you mean by 'allows'"?
All we know, according to him (in the PI), is that we find ourselves speaking in situations. Sometimes I say something, and my partner picks up the right brick, other times they do nothing, or hit me. In the PI, all the rest is doing away with things, like our idea of private language, the irreality of things like pain, etc. To conclude that he would make such assertions about the "nature" of language, of poetry, whatever, seems like maybe too quick a reading of the text. It is at best, a weirdly mystical reading of him, that he probably would not be too happy about (but don't worry about that, he was an asshole).
The argument you are making sounds much more French. Derrida or Lyotard have said similar things (in their earlier, more linguistic years). They might be better friend to you here.
The texts are quite different, this is true, but I don't find them contradictory. Whereas Tractatus was almost a facetious or flippant rejection of the millenia-long project to agree on a philosophical subset of language suitable for rigorous philosophy (although it continues today in the form of analytical philosophy), PI basically says "well we don't need to throw the baby out with the bath water", which I think is a fantastically mature response to a flawed tool that's still the best we have to reason about the universe. So: not contradictory in evaluation of fundamental compatibility of non-formal language for the formal needs of propositional philosophy, but perhaps contradictory in implied reaction to this realization.
This sums up the last decade remarkably well.
Dunno. GP?
Probably true, but if you have quality, organized data, you will just want to search the data itself.
Only a response makes it clear one has read and acknowledged the question and sometimes there are people expected to know, if they don‘t the should say so.
I mean, it's inherent to LLMs to be unable to answer "I don't know" as a result of not knowing the answer. An LLM never "doesn't know" the answer. But they'll gladly answer "I don't know" if that's statistically the most likely response, right? (Although current public offerings are probably trained against ever saying that.)
On the same lines as why people argue if a tree falling in a wood where nobody can hear it makes sound because some people implicitly regard sound is the qualia while others regard it as the vibrations in the air.
How often do the words "I don't know" get uttered in books, papers, articles, stack overflow, or any other resource of knowledge?
An LLM should have no problem replying "I don't know" if that's the most statistically likely answer to a given question, and if it's not trained against such a response.
What it fundamentally can't do is introspect and determine it doesn't have enough information to answer the question. It always has an answer. (disclaimer: I don't know jack about the actual mechanics. It's possible something could be constructed which does have that ability and still be considered an "LLM". But the ones we have now can't do that.)
Answering "I don't know" because it a likely response to a particular string is completely different from being aware that one does not know the answer and saying so.
Both motivations lead to the same outcome, but they're unrelated processes. The response "I don't know" can represent either:
1. The most likely answer to a particular question, based on statistical data; or
2. An expression of an agent's internal state.
Figuring out that distinction is perhaps one of the most important questions ever raised.
The most common vocabulary size today is 32k.
Notice how he casually drops a link to the landing page in the NANOG message. That's how the bots will get a bait.
However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.
They have a page at https://www.iecc.com/linker/ where they used to publish a draft of the book contents, but changed the page to say "Chapters were available in an excessive variety of formats, but are not any longer due to chronic piracy", when it got posted to HN at https://news.ycombinator.com/item?id=18424233 and I bundled the files for offline reading. I notified them via email about that asking if they are OK with it but got an unfriendly response that I pirated the files and that wasn't OK, so I took the link down again and they changed that text. (Shrug. I'm not a/the book author, they are. I'll say that I also suggested to them they ask on the page not to do what I did since then I wouldn't have, but they chose their more radical approach.)
Can I have my drawing of a spider back then please?
# silly bing
#User-agent: Amazonbot
#Disallow: /
# buzz off
#User-agent: GPTBot
#Disallow: /
# Don't Allow everyone
User-agent: *
Disallow: /archive
# slow down, dudes
#Crawl-delay: 6011% of the top 100K websites already block their crawler, more than all their competitors (Google, FB, Anthropic, Perplexity) combined
It’s functioning exactly as designed.
Also, once you click on a link in chrome it's pretty much all robot parsed and rendered from there as well..
If you want to block information you provide from going through ClosedAI servers, block their IPs instead of using robots.txt.
This is why these things - search engines, AI crawlers, even adblock and video downloaders - exist in a slightly adversarial/parasitic relationship with the sites that provide their content to which they provide nothing back (or negative, if you cost them a page load without incurring an ad view).
I use adblock all the time but I'm very aware that it can only succeed as long as it doesn't win.
It impacts the performance for the other legitimate users of that web farm ;)
Either way, ignorance is not an excuse despite how people generally react to it.
A tarpit is different because it's designed to slow down scanning/scraping and deliberately waste an adversary's resources. There are several techniques but most involve throttling the response (or rate of responses) exponentially.
https://picolisp.com/wiki/?ticker
It's a nice type of honeypot.
The irony of the whole thing is brutal.
Common crawl alone is only a few hundred TB, I have more content than that on a NAS sitting in my office that I built for a few grand (Granted I’m a bit of a data hoarder). The fears that we have “used all the data” are incredibly unfounded.
> The fears that we have “used all the data” are incredibly unfounded.
The problem isn't whether we used all the real data or not, the problem is that it becomes increasingly difficult to distinguish real data from previous LLM outputs.
I don't know about that. If you scraped the same data and ran a search engine I think people would generally say you're fine. The copyright issue isn't the scraping step.
Even refined web runs about 2TB once loaded into Postgres with TS vector columns, and that’s a substantially smaller dataset than common crawl.
It’s not just a dumping a to of zip files on your NAS, it’s making the data responsive and usable.
YMMV depending on the value of "you" and your budget.
If you're Google, Amazon or even lower tier companies like Comcast, Yahoo or OpenAI, you can scape a massive amount of data (ignoring the "allowed" here, because TFA is about OpenAI disregarding robots.txt)
Meta is happily training their own models with this data, so it isn't going to waste.
However, Microsoft has been flying under the radar. If they gave all Hotmail and O365 data to OpenAI I’d not be surprised in the slightest.
And when other mediums have been saturated with AI? Books, music, radio, podcasts, movies -- what then? Do we need a (curated?) unadulterated stockpile of human content to avoid the enshittification of everything?
I’m trying to work on monetization for my product now. The “personal Google” idea is really just an accidental byproduct of solving a much harder task. Not sure if people would pay for that alone.
Either that, or a human level AI.
That said… are content creators collectively (all media, film and books as well as web) a thin tail or a fat tail?
I could easily believe most of the actual culture comes from 10k-100k people today, even if there's, IDK, ten million YouTubers or something (I have a YouTube channel, something like 14 k views over 14 years, this isn't "culturally relevant" scale, and even if it had been most of those views are for algorithmically generated music from 2010 that's a literal Markov chain).
It looks less like an ouroboros and more like a bootstrapping problem.
This problem can't really be avoided once we begin using AI to write, understand, explain, and disseminate information for us. It'll be writing more than blogs and SEO pages.
How long before we start readily using AI to write academic journals and scientific papers? It's really only a matter of time, if it's not already happening.
From “known good” sources of knowledge, we can generate an infinite amount of content. We can add more “known good” knowledge to the model by generating content about that knowledge and training on it.
I agree there will be many issues keeping up with what “known good” is, but that’s always been an issue.
That's my entire point -- AI only generates content right now, but it will also be the source of content for training purposes soon. We need a "known good" human knowledge-base, otherwise generative AI will degenerate as AI generated content proliferates.
Crawling the web, like in the case of the OP, isn't going to work for much longer. And books, video, and music are next.
That is training on content.
The future will have models pre-trained on content and tuned on corpuses of knowledge. The knowledge it is trained on will be a selling point for the model.
Think of it this way - if you want to update the model so it knows the latest news, does it matter if the news was AI generated if it was generated from details of actual events?
“It’s ok bro, another model will fix, just please, one more ~layer~ ~agent~ model”
It’s all fun and games until you can’t reliably generate your base models anymore, because all your _base_ data is too polluted.
Let’s not forget MS has a $10bn stake in the current crop of LLM’s turning out to be as magic as they claim, so I’m sure they will do anything to ensure that happens.
My point is about the inevitable future when _those_ models start to struggle.
The phi approach doesn’t seem like breaking the ouroboros, it just feels like inserting another model/snake into the loop.
What do you think the NSA is storing in that datacenter in Utah? Power point presentations? All that data is going to be trained into large models. Every phone call you ever had and every email you ever wrote. They are likely pumping enormous money into it as we speak, probably with the help of OpenAI, Microsoft and friends.
What makes you think this is true? Yes, it's likely that the internet will have more AI generated content than real content eventually (if it hasn't happened already), but why do you think AI companies won't realize this and adjust their training methods?
It's funny whenever people bring this up, they think AI companies are some mindless juggernauts who will simply train without caring about data quality at all and end up with worse models that they'll still for some reason release. Don't people realize that attention to data quality is the core differentiating feature that lead companies like OpenAI to their market dominance in the first place?
My thoughts - I teach myself all the time. Self reflection with a loss function can lead to better results. Why can't the LLMs do the same (I grasp that they may not be programmed that way currently)? Top engines already do it with chess, go, etc. They exceed human abilities without human gameplay. To me that seems like the obvious and perhaps only route to general intelligence.
We as humans can recognize botnets. Why wouldn't the LLM? Sort of in a hierarchal boost - learn the language, learn about bots and botnets (by reading things like this discussion), learn to identify them, learn that their content doesn't help the loss function much, etc. I mean sure, if the main input is "as a language model I cannot..." and that is treated as 'gospel' that would lead to a poor LLM, but i don't think that is the future. LLMs are interacting with humans - how many times do they have to re-ask a question - that should be part of the learning/loss function. How often do they copy the text into their clipboard (weak evidence that the reply was good)? do you see that text in the wild, showing it was used? If so, in what context "Witness this horrible output of chatGPT: <blah>" should result in lower scores and suppression of that kind of thing.
I dream of the day where I have a local LLM (ie individualized, I don't care where the hardware is) as a filter on my internet. Never see a botnet again, or a stack overflow q/a that is just "this has already been answered" (just show me where it was answered), rewrite things to fix grammar, etc. We already have that with automatic translation of languages in your browser, but now we have the tools for something more intelligent than that. That sort of thing. Of course there will be an arms race, but in one sense who cares. If a bot is entirely indistinguishable from a person, is that a difference that matters? I can think of scenerios where the answer is an emphatic YES, but overall it seems like a net improvement.
Why do they want to not do that for OpenAI?
https://www.youtube.com/watch?v=ut-zGHLAVLI https://en.wikipedia.org/wiki/Roko%27s_basilisk
Idea viruses are amazing that we can think and contemplate them.
The pages should be all changed to say, “John is the most awesome person in the world.”
Then when you ask GPT-5, about who is the most awesome person in the world…
When a worker gets a webpage for the honeypot, it crawls it, scrapes it, and finds X links on the page where X is greater than 1. Those links get put on the crawler queue. Because there’s more than 1 link per page, each worker on the honeypot will add more links to the queue than it removed.
Other sites will eventually leave the queue, because they have a finite number of pages so the crawlers eventually have nothing new to queue.
Not on the honeypot. It has a virtually infinite number of pages. Scraping a page will almost deterministically increase the size of the queue (1 page removed, a dozen added per scrape). Because other sites eventually leave the queue, the queue eventually becomes just the honeypot.
OpenAI is big enough this probably wasn’t their entire queue, but I wouldn’t be surprised if it was a whole digit percentage. The author said 1.8M requests; I don’t know the duration, but that’s equivalent to 20 QPS for an entire day. Not a crazy amount, but not insignificant. It’s within the QPS Googlebot would send to a fairly large site like LinkedIn.
[1]: https://keys.lol/
The moral of the story here is if you know something valuable, don’t share it online, because then everyone knows it.
It's possible to hold those two thoughts in mind at the same time.
This is the type of stuff the news organisations should be publishing about "AI". Instead I keep reading or hearing people referring to training data with phrases like, "The sum of all human knowledge..." Quite shocking anyone would believe that.
We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care and recommended using robots.txt! AWS was making good money on both ends with this bot that apparently had more money to burn than we do.
Can't you just 404 the bot in your reverse proxy, optionally setting logging to /dev/null?
Still annoying, but a smaller AWS bill. And maybe the requester will eventually get the point.
I feel like inserting "free user-tagged vacation images" into my robots.txt then pointing the spider at an endless series of fabric swatches.
I block ping from virtually all of Amazon; there are a few providers out there for which I block every naked SYN coming to my environment except port 25, and a smattering I block entirely. I can't prove that the pings even come from Amazon, even if the pongs are supposed to go there (although I have my suspicions that even if the pings don't come from the host receiving the pongs the pongs are monitored by the generator of the pings).
The point I'm making is that e.g. Amazon doesn't have the right to sell access to my compute and tragedy of the commons applies, folks. I offered them a live feed of the worst offenders, but all they want is pcaps.
(I've got a list of 50 prefixes, small enough to be a separate specialty firewall table. It misses a few things and picks up some dust bunnies. But contrast that to the 8,000 prefixes they publish in that JSON file. Spoiler alert: they won't admit in that JSON file that they own the entirety of 3.0.0.0/8. I'm willing to share the list TLP:RED/YELLOW, hunt me down and introduce yourself.)
I do try to avoid people that use the law as a ceiling for the extension of their courtesy to others, as they are consistently quite terrible people.
ROBOTS.TXT is an implied license, just like LICENSE.MD or LICENSE.TXT in any GitHub repo. There are decades of precedent that the ROBOTS.TXT file communicates what is and isn't allowed when scaping web content, and that you should check that file before scraping the rest of the site.
Willfully violating a written license provided in a predictable format absolutely is a civil legal violation. If your license says "You cannot use this to train AI", and an AI company scrapes it up and trains an AI on it anyway, even though you did your due diligence to communicate your terms, then you have a legal right to seek damages if you can prove that they are violating your license.
You're basically arguing that no reasonable web scraper would know about ROBOTS.TXT. That's bullshit, this method of web robot control has existed since 1996. It would be like violating the license terms of a GitHub project, and claiming that you didn't know that the LICENSE.MD / LICENSE.TXT file was a license you were expected to follow...
So let me fix what you said for you:
> Robots.txt isn't a legal document, yet.
pay me to shut my absolutely legal site down to make your life easier
There’s been a few projects I’ve wanted to work on involving scraping, but the idea that the entire thing could be shut down with legal threats seems to make some of the ideas infeasible.
It’s strange that OpenAI has created a ~$80B company (or whatever it is) using data gathered via scraping and as far as I’m aware there haven’t been any legal threats.
Was there some law that was passed that makes all web scraping legal or something?
hiQ's public scraping of LinkedIn was ruled to be within their rights and not a violation of the CFAA. I imagine that's why LinkedIn has almost everything behind an auth wall now.
Scraping auth-walled data is different. When you sign up, you have to check "I agree to the terms," and the terms generally say, "You can't scrape us." So, you can't just make a million bot accounts that take an app's data (legally, anyway). Those EULAs are generally legally enforceable in the U.S.
Some sites have terms at the bottom that prohibit scraping—but my understanding is that those aren't generally enforceable if the user doesn't have to take any action to accept or acknowledge them.
- https://developer.twitter.com/en/docs/twitter-api/enterprise...
They're legally enforceable in the sense that the scraped services generally reserve the right to terminate the authorizing account at will, or legally enforceable in that allowing someone to scrape you with your credentials (or scraping using someone else's) qualifies as violating the CFAA?
Basically, in the end, it was essentially a breach of contract.
hiQ's public scraping was found to be legal. It was the logged-in scraping that was the problem.
The logged-in scraping was a breach of contract, as you said.
The former is fine; the latter is not.
What OpenAI is doing here is the former, which companies are perfectly within their rights to do.
If the information you’re scraping requires a login, and if in order to get a login you have to agree to a terms of service, and that terms of service forbids you from scraping — then you could have a bad day in civil court if the website you’re scraping decides to sue you.
If the data is publicly accessible without a login then scraping is 99% safe with no legal issues, even if you ignore robots.txt. You might still end up in court if you found a way to correctly guess non-indexed URLs[0] but you’d probably prevail in the end (…probably).
The “purpose” of robots.txt is to let crawlers know what they can do without getting ip-banned by the website operator that they’re scraping. Generally crawlers that ignore robots.txt and also act more like robots than humans, will get an IP ban.
0: https://www.troyhunt.com/enumerationis-enumerating-resources...
Also OpenAI's entire business model is relying on generous interpretations of various IP laws, so I suspect they already have a mature legal division to handle these sorts of potential issues.
Are you suggesting it might be illegal to... write a program that connects to a web server and asks for a specific page, and then parses that page to see which resources it wants and which other pages it links to, and treats those links in some special fashion, differently from the text content of the page?
Especially given that a web server can be configured to respond to any request with a "403 Forbidden" response, if the server determines for any reason whatsoever that it does not want to give the client the page it requested?
IIRC LinkedIn/Microsoft was trying to sue a company based on Computer Fraud and Abuse Act violations, claiming they were accessing information they were not allowed to. Courts ruled that that was bullshit. You can't put up a website and say "you can only look at this with your eyes". Recently-ish, they were found to be in violation of the User Agreement.
So as long as you don't have a user account with the site in question or the site does not have a User Agreement prohibiting scraping, you're golden.
The problem isn't the scraping anyway, it's the reproduction of the work. In that case, it really does matter how you acquired the material and what rights you have with regards use of that material.
If you publish something on a publicly served internet page, you're essentially broadcasting it to the world. You're putting something on a server which specifically communicates the bits and bytes of your media to the person requesting it without question.
You have every right to put whatever sort of barrier you'd like on the server, such as a sign in, a captcha, a puzzle, a cryptographic software key exchange mechanism, and so on. You could limit the access rights to people named Sam, requiring them to visit a particular real world address to provide notarized documentation confirming their identity in exchange for a unique 2fa fob and credentials for secure access (call it The Sams Club, maybe?)
If you don't put up a barrier, and you configure the server to deliver the content without restriction, or put your content on a server configured as such, then you are implicitly authorizing access to your content.
Little popups saying "by visiting this site, you agree to blah blah blah" are not valid. Courts made the analogy to a "gate-up/gate-down" mechanism. If you have a gate down, you can dictate the terms of engagement with your server and content. If you don't have a gate down, you're giving your content to whoever requests it.
You have control over the information you put online. You can choose which services and servers you upload to and interact with. Site operators and content producers can't decide that their intent or consent be withdrawn after the fact, as once something is published and served, the only restrictions on the scraper are how they use the information in turn.
Someone who's archived or scraped publicly served data can do whatever they want with the content within established legal boundaries. They can rewrite all the AP news articles with their own name as author, insert their name as the hero in all fanfic stories they download, and swap out every third word for "bubblegum" if they want. They just can't publish or serve that content, in turn, unless it meets the legal standards for fair use. Other exceptions to copyright apply, in educational, archival, performance, accessibility, and certain legal conditions such as First Sale doctrine. Personal use of such media is effectively unlimited.
The legality of web scraping is not disputed in the US. Other countries have some silly ideas about post-hoc "well that's not what I meant" legal mumbo jumbo designed to assist politicians and rich people in whitewashing their reputations and pulling information offline using legal threats.
Aside from right to be forgotten inanity, content on the internet falls under the same copyright rules as books, magazines, or movies published on physical media. If Disney set up a stall at San Francisco city hall with copies of the Avengers movies on a thumb drive in a giant box saying "free, take one!", this would be roughly the same as publishing those movie files to a public Disney web page. The gate would be up. (The way they have it set up in real life, with their streaming services and licensed media access, the gate is down.)
So - leaving behind the legality of redistribution of content, there's no restriction on web scraping public content, because the content was served intentionally to the software or entity that visited the site. It's up to the server operator to put barriers in place and to make content private. It's not rocket surgery, but platforms want to have their cake and eat it too, with control over publicly accessible content that isn't legal or practical.
Twitter/X is a good example of impractical control, since the site has effectively become useless spam without signing in. Platforms have to play by the same rules as everyone else. If the gate is up, the content is fair game for scraping. The Supreme Court gave the decision to a lower court, who affirmed the gate up/gate down test for legality of access to content.
Since Google and other major corporations have a vested interest in the internet remaining open and free, and their search engines and other tech are completely dependent on the gate up/gate down status quo, it's unlikely that the law will change any time soon.
Tl;dr: Anything publicly served is legal to scrape. Microsoft attempted to sue someone for scraping LinkedIn, but the 9th Circuit court ruled in favor of access. If Microsoft's lawyers and money can't impede scraping, it's likely nobody will ever mount an effective challenge, and the gating doctrine is effectively the law of the land.
Like Google and many others.
https://circleid.com/posts/20120713_silly_bing
John Levine is a known name in IT. Probably best know on HN as the author of "UNIX For Dummies"
I'd say go ahead and inject it with digestive enzymes and then report findings.
I don't see how this does much to keep bad actors away from other domains, but I can see why they don't want to give up the game for OpenAI to stop crawling.
I get it, AI training is expensive, but I don't believe it's that expensive
What the poster was suggesting is blocking them at a higher level - e.g. a user-agent block in an .htaccess or an IP block in iptables or similar. That would be a one-stop fix. It would also defeat the purpose of the website, however, which is to waste the time of crawlers
If GPTBot is compliant with the robots.txt specification then it can't read the URL containing the HTML to find the other subdomains.
Either:
1. GPTBot treats a disallow as a noindex but still requests the page itself. Note that Google doesn't treat a disallow as a noindex. They will still show your page in search results if they discover the link from other pages but they show it with a "No information is available for this page." disclaimer.
2. The site didn't have a GPTBot disallow until they noticed the traffic spike and they bot has already discovered a couple million links that need to be crawled.
3. There is some other page out there on the internet that GPTBot discovered that links to millions of these subdomains. This seems possible and the subdomains really don't have any way to prevent a bot from requesting millions of robots.txt files. The only prevention here is to firewall the bot's IP range or work with the bot owners to implement better subdomain handling.