A robots.txt Problem
avodonosov.blogspot.com
avodonosov.blogspot.com
When would you ever want to serve gzipped content to a client who cannot accept it?
From what OP linked [0], it clearly is a problem Google is aware of and are trying to communicate a workaround for.
I suppose the answer is that it is hard to change a default behavior once it has been released, even if it is wrong.
[0] https://cloud.google.com/appengine/docs/legacy/standard/java...
It is relatively rare, for sure, but it is nothing new.
In these rare cases, it is impossible to disable compression from the client side. The Accept-Encoding header is ignored. I frequently use clients that are not popular browsers nor even popular programs such as curl. I suspect that people using popular browsers would never even notice that varying the Accept-Encoding header had no effect, and that is why the CDNs believe that this is OK.
I feel like this is a trend for pretty much everything on the internet these days, the 1% of anything doesn't matter to any service, not using google chrome? "please use a supported browser - aka google chrome", using a VPN? CDN level block or "here's 1000 captchas because Google hates you", You want en-gb? not from that IP enjoy your crash course in the native language, You aren't in the USA? no content for you.
Excluding the 1% might make business sense as an optimisation abstractly, but when that 1% could affect different people each day it can affect everyone eventually and you will actually piss off all users.
Now it's really common for stuff just to straight up not work or you are actively blocked and considered collateral in protection against spam and botnets, the latter is an entirely new phenomenon.
I laughed at the "five month", then realized it is actually impressive that OP got any response at all. What a time to be alive.
Also, why not just ban the offender?
Surely the offender here is Google AppEngine.
I was thinking about that, but didn't find an easy way ban a crawler. Google App Engine has a firewall, but it works based on IP addresses. Banning based on User-Agent would need to be done in the app code, and that essentially handling a request, even if in a cheaper way. I didn't want to touch the application at all, hoping to resolve this on the crawler side, whom I suspected being an unintentional "offender".
Speaking about the five months - that's fine. We were not communicating every day of course. And indeed impressive that I had my case handled at all.
I knew for years that unwanted crawling happens by various crawlers, and was reminded of that in metrics from time to time. One day I was in the mood to study deeper, found two crawlers in the access logs, studied their web sites and emailed them.
One didn't respond at all. The moz.com created a ticket, four days later a support engineer replied, a week later I replied. We had some back and forth. I supposed they don't recognize `User-agent: *` and need `User-agent: Dotbot`. David - the support engineer - expressed several other hypotheses. There was a period of silence, then I raised my issue again, David had it reviewed by some other people at moz.com and they pointed to the gzipped response.
BTW, what I learned, is that "If no Accept-Encoding field is present in a request, the server MAY assume that the client will accept any content coding." (https://www.rfc-editor.org/rfc/rfc9110.html#name-accept-enco...).
So if we make an HTTP request, unless we explicitly specify `Accept-Encoding: identity` we'd better be prepared to inspect the Content-Encoding in the response and decompress data if necessary.
But since Google App Engine returns gzipped content even for requests with `Accept-Encoding: identity`, I accepted that the the failure is on my side and went on with the config changes. Still, left a recommendation for moz.com to support gzip on their end.
Don't put it online if you don't want it crawled. Welcome to the internet.
So does Wikipedia. https://en.wikipedia.org/robots.txt . That points out:
# enwiki:
# Folks get annoyed when VfD discussions end up the number 1 google hit for
# their name. See T6776
Disallow: /wiki/Wikipedia:Articles_for_deletion/
As does GitHub - https://github.com/robots.txtEven the Internet Archive, which doesn't honor directives in the robots.txt files, has one - http://archive.org/robots.txt .
Welcome to the internet.
Also, cargo culting has never been a good reason to do anything.
Also, sloths can hold their breath underwater for up to 40 minutes.
Also, some poorly rate-limited crawlers actually abide by robots.txt, so it's useful to prevent unnecessary load.
Robots.txt isn't for hiding/suppressing information.
Often times you can have whole URL structures that are redundant with other ones, mainly database-generated pages with all sorts of possible query parameters often disguised as paths. Robots.txt is extremely useful in ensuring crawlers can make life easier for themselves by limiting to the "real" content, as opposed to the redundant stuff. Crawling the 5,000 real pages, not the 500,000 additional URL's that return the same content.
Also for ignoring "interactive" pages like login pages that make zero sense to be crawled.
People "give a crap" about robots.txt because it's useful for that.
Also - Google is perfectly good enough at turning down services on their own, we don't need to give them any ideas!
Yeah, AppEngine isn't what it could be, but deprecating it would be a step backwards.
It's been a while since I've done web development but IIRC some web servers (I seem to recall Apache doing this) do this implicitly, i.e. you don't need to add a Vary header for Accept-Encoding since the web server is smart enough to know that this is what you mean. And sending a "Vary: Accept-Encoding" response header likewise seems silly since it's extra data in every request that ought to be implied. Nevertheless I think that by a strict reading of RFC 2616 the behavior of Google here is allowed, and that to be pedantic you should send a Vary: Accept-Encoding header in this case.
> Fixed the Gooble App Engine behaviour by adding an explicit configuration to the appengine-web.xml
> If an Accept-Encoding field is present in a request, and if the server cannot send a response which is acceptable according to the Accept-Encoding header, then the server SHOULD send an error response with the 406 (Not Acceptable) status code.
I agree with your strict reading of the spec, but can’t relate to your attitude towards that reading, nor your attitude towards placing the burden on users to work around it. The spec’s allowance here is very probably intended to produce graceful successes in likely usage scenarios when the client might be able to handle a response it didn’t request and when the server logic is insufficiently robust.
It’s (my speculation here) a transfer of the robustness principle to the client on the basis that clients are fewer and better resourced to be robust than servers in this case. That reasoning applied to Google as compared to its own customers is untenable. There would very little burden placed on Google, or anyone else for that matter, by expecting them to honor the intent of the spec and the intent of requests. Even the most naive solution would somewhat less than double their cache size (which for smaller orgs would be a real burden, but for Google that’s laughable) and at worst would degrade to uncached performance for an initial cache miss. Deferring to Google’s documentation, however clear, relieves them of probably a single engineer’s sprint time, some budgeting consideration… and costs M/N engineering hours of frustration and contribution to burnout while people think they’re convenienced by offloading work to Google as a reliable vendor. Then after shaving so many yaks, if they have the temerity and energy to post about their frustrating experience, they’re criticized on HN for not reading the documentation which they referenced after struggling to find it.
Google is, from my understanding of the post, fully standards-compliant and your reading is correct. But Technically Correct isn’t the best kind of correct when a giant megacorporation gets megabucks to provide a service which doesn’t do what even the spec says is not ideal, and their “solution” is that every single one of their customers must discover this fact individually and cater to them.
To your point about savings, it seems like it would be most cost effective, over time, to simply send an empty 406 response, saving the machine power to read and return the cache and the network traffic of sending a file that can't be used anyway.
And "Gooble App Engine" near the end of the article.
I'm imagining some kind of knock-off brand cloud provider.
There was a "Telsa" misprint that always gave me a chuckle. It made me imagine "Telsa, by Eron Muks".
User-agent: * Disallow: /
And then no one can find them via search and they go broke.
We used it earlier on in a startup we were working on and had so many issues that I would never recommend it to anyone.
You'll be much better off using GKE or some other kubernetes variant
The current way is akin to saying ... "You do not have a lock on your door. So we are welcome to come in ..."
Robots.txt is a way of saying "these are not the pages you are looking for".