With the rise of AI, web crawlers are suddenly controversial
theverge.com
theverge.com
The basic social contract of the web fell apart long ago when almost everyone decided that Google was the only search engine worth serving and started aggressively blocking other crawlers.
You mean like this public API
https://hacker-news.firebaseio.com/v0/item/39421253.json
Or this RSS feed
And if you use your cookies they temp-block your account after x (like x every 30 minutes) times.
And I didn't even care that there was no content on the RSS feed, I just wanted the title / notification so I don't have to use ... (I can't put anything in here as there is no alternative that allows me to track everything).
curl -si40A "" https://old.reddit.com/new.rss > 1.json
sed 10q 1.json
HTTP/1.1 200 OK
content-length: 40582
content-type: application/atom+xml; charset=UTF-8
x-ua-compatible: IE=edge
x-frame-options: SAMEORIGIN
x-ratelimit-remaining: 93
x-ratelimit-used: 3
x-ratelimit-reset: 96
x-reddit-pod-ip: 10.104.158.170:80
x-reddit-internal-ratelimit-rls-type: ip-standard
In theory, APIs to retrieve public information suck because they are designed to be rate-limited or subject to quotas. Whereeas IME public-facing websites are much less likely to set rate limits and enforce quotas.I prefer to use the public-facing website instead of "Web APIs" and convert HTML to SQLite which I prefer over non-LD JSON. If the public-facing website changes their HTML, I change the code in the filter. This is rare. It is very simple for me to do.
How did perplexity.ai crawl the internet for their AI?
Many revenue-based websites tried to have it both ways with web crawlers wherein they wanted to block automated access or repeat viewers while letting first time viewers get a free taste. Others have noted that basically Google gets a free pass for all the traffic it brings in but everyone else has to respect robots declarations.
It seems like a no brainer - if your web server is configured to reply to GET requests with a 200 status and some content then they get to do pretty much whatever they want with it.
Don't want to give access to everyone? Stop sending your content for free and get them to agree to some contract and authorize/license their access to your stuff.
I'm sure the sites would rather that google paid them in licencing too, but without changes to our laws that's not going to happen
I'm not sure what the solution is, but paywalled articles in search results are bad. If they want to be indexed they should have to offer that same indexed content to anyone browsing the index.
Google does advertise that they index based on the same content that's available to anyone viewing the page, and has policies against presenting a different version of the page to their crawler versus what you're showing to visitors.
Should Google also stop indexing Facebook, since Facebook puts a login wall for people to access their content? Should YouTube (ie Google) ban movie trailers, since it's just a tease for paywalled movies? The iTunes store let people listen to 30 seconds of a song before purchasing at the paywall. Was that wrong?
Yes - I thought they already did? (I know LinkedIn edges around this by putting up a login wall only if you have a cookie showing that you'd logged in previously).
> Should YouTube (ie Google) ban movie trailers, since it's just a tease for paywalled movies? The iTunes store let people listen to 30 seconds of a song before purchasing at the paywall. Was that wrong?
A free sample of a paid thing is fine if everyone knows that's what it is. It's when you bait-and-switch by offering something that seems like it's free to start with that it's a problem. Like imagine showing a movie in the town square and then 10 minutes in you pause it and tell everyone they need to buy a ticket or leave.
Just checked, Google still indexes Facebook and puts relevant results on top. If you're not logged in you can't continue.
Should Google Maps remove businesses that charge for their products and services from their search results as well?
If by clicking on the thing it does not have the content I searched for (how am I even certain I get it when I pay you?) I would call that result bad.
If you want to charge for stuff that's great, I recommend it, and if you want to give out a free sample or an index that's great, but it should be the same to all comers.
If I went to a public water fountain and found that someone had turned it into a coca cola dispensing machine, I wouldn't be happy and wouldn't pay to use it.
"Journalists" creating pay walls, using SEO tactics to push their articles into my search results, and then trying to extract rent don't deserve money.
If the search results are full of paywalled articles that claim to have text relevant to my query, but won't show me the article because publishers are trying to extract money from me, the publishers of those articles have made my task harder and shouldn't be rewarded. This is a form of spam.
And sometimes the answer is behind a paywall. It's not spam at all. On the contrary, spam is always free.
When this happens, I will continue to pretend to be google to access the content they are pushing. If publishers want to change this behavior they could try not letting Google index it, so I don't need to see it in my search results.
For other users, they see value in having paywalled results if they are the best results, because they do not have a block against paying for content.
If you for example search for a movie on Google, they'll show you paid options to watch it on streaming services or rent it from streaming services. That's good and what should be expected from a search engine.
Paying for stuff is how the world works. If a restaurant boasts about having the nicest steaks, you're not going to get a free steak just to be sure that it's good.
But I really think it is time for a better way to pay for content and articles instead of having to subscribe to each source.
It's like if a friend of yours takes you to a nice Mexican place. Why would you expect to get a burrito al Pastor to eat for free, just because you eat for free when you visit relatives? Nobody said it would be free.
For Google, they have made a product decision about how to treat paywalled content. They don’t care. It hurts the user experience but the days when Google cared about improving their search experience are long gone.
getting all journalism for free isn't sustainable though
I guess it might be a catch twenty two. Low quality -> low income -> low quality, but the sadness of such a dynamic does not make me want to pay for an inferior service.
Is ok to serve a full article to Google and put visitors behind wall?
Giving the search engine the full content so it will rank higher is lying, because that information isn't actually there unless you pay, which most users aren't going to do.
In theory you could solve this with a requirement for the site to disclose that it's doing this, so the search engine can have a box that says "exclude paywalled content", but then everybody would check the box.
I agree the option would be nice, but I think you are wrong that everyone wants paywalled articles excluded
Which is what they do by penalizing sites for showing something different to Googlebot than the user.
> I agree the option would be nice, but I think you are wrong that everyone wants paywalled articles excluded
What they want is for non-paywalled links that include the relevant information to be listed first, which is de facto equivalent to excluding paywalled links (because they'll be on page 75) in the vast majority of search results.
A search based on the results being something otherwise than the contents of the search is misleading, a spam tactic, and bad.
Agreed, we need to wrap the public web with DRM immediately. We can't expect companies like OpenAI to waste time worrying about pesky legalities like copyright and content licensing.
If you're trying to prevent anyone from getting a copy of what's on your site, you're probably screwed, because technical solutions are hard and legal solutions are only going to be violated by The Internet since it isn't all in your jurisdiction.
If you're just trying to keep them from putting an excessive load on your servers, technical solutions are easy. Just give them an API or some other efficient way to receive the data.
I don't see how arguing that it should be harder for journalists to get paid helps anything though
(Except I'm pretty sure that doesn't work either).
We are sure that OpenAI is building its business on copyrighted content, under the defense that to do otherwise is "impossible."
The courts have not decided whether this is "fair use". However, the four factors make it pretty clear.
If it is a use that falls under fair use, I'm thinking about the Google Books case, where an act that is far less transformative (digital archiving and search) was found to be fair use (https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....)
(It's going to be "who's got more money?", isn't it? Fun fun fun.)
This is a very strange way to look at it. If I encode Oppenheimer to MPEG-4, do I now own that IP? In both cases, I'm looking at the source and encoding it to a compressed representation.
A couple notes on "transformative": (1) It's just one of the four factors. (2) It means that you added new expression or meaning to it. By definition, an LLM can't create anything it wasn't trained on — it's pattern recognition, interpolation, and recombination.
https://chat.openai.com/share/2b9a8245-2063-4722-81e8-20d10c...
It's funny because you say it's "pretty clear" without saying which way you think it goes, which then makes it unclear what you think the result is.
It never governed anything, web crawlers were never under any obligation to follow robots.txt.
This article seems like they took an existing controversy, rebranded it as something new, then blamed in on AI.
Basically describes most of what passes as journalism these days.
other than the fact that you could get successful sued if you do not follow them i.e.: eBay v. Bidder's Edge in 2000
This really hit home for me in Hard Fork’s interview with the CEO of Perplexity AI [1]. The crux of the problem is, if Perplexity does what it strives to do, provide answers rather than links, no one will be clicking on links and generating revenue for newspapers anymore. I encourage you to watch him struggle with this question because this is really an irreconcilable problem.
I'm not sure I'm ready to concede the fundemental value judgement being made here. At least I refuse to accept it as a given rather then the core issue to be decided.
Data accessible. Data free. Period.
Or did it ever?