Google’s robots.txt parser is now open source
opensource.googleblog.com
opensource.googleblog.com
Having written the robots.txt parser at Blekko, I can tell you what standards there are incomplete and inconsistent.
Robots.txt files are usually written by hand using random text editors ("/n" vs "/r/n" vs a mix of both!) by people who have no idea what a programming language grammar is. Let alone follow BNF from the RFC. There are situations where adding a newline completely negates all your rules. Specifically, newlines between useragent lines nor between useragent lines and rules.
My first inclination was to build an RFC compliant parser and point to the standard if anyone complained. However, if you start looking at a cross section of robots.txt files, you see that very few are well formed.
With the addition of sitemaps, crawl-delay, and other non-standard syntax adopted by Google, Bing, and Yahoo (RIP). Clearly the RFC is just a starting point and what ends up on website can be broken and hard to interpret the author's meaning. For example, the Google parser allows for five possible spellings of DISALLOW, including DISALLAW.
If you read a few webmaster boards, you see that many website owners don't want a lesson in Backus–Naur form and are quick to get the torches and pitchforks if they feel some crawler is wasting their precious CPU cycles or cluttering up their log files. Having a robots.txt parser that "does what the webmaster intends" is critical. Sometimes, I couldn't figure out what some particular webmaster intended, let alone write a program that could. The only solution was to draft off of Google's de facto standard.
(To the webmaster with the broken robots.txt and links on every product page with a CGI arg with "&action=DELETE" in it, we're so sorry! but... why???)
Here's the Perl for the Blekko robots.txt parser. https://github.com/randomstring/ParseRobotsTXT
Lots of us really didn’t know what we were doing and we’d made all the action buttons in the listing screens regular links. As you can imagine, pandemonium ensued.
Hey, at least we’d figured out that sql injection was a thing.
It was a simpler time.
Microsoft did this via an explicit web manifest; the web page author needed to list all of the resources they wanted to use in offline or pre-cache mode.
Netscape tried to do this by urging web authors to Be Very Careful with the links on a page, which usually required a specially-crafted offline-crawler-only version of the site. Predicably, hilarity ensued.
The term of art at the time was "push technology" or "web push", the irony of which was not lost upon those tasked with making it work.
1997. Good times.
https://web.archive.org/web/20050505061702/http://webacceler...
The issue you described seems to match this: https://signalvnoise.com/archives2/google_web_accelerator_he...
[1] https://webmasters.googleblog.com/2019/05/the-new-evergreen-...
https://qph.fs.quoracdn.net/main-qimg-002d1f819e1bfbbd14fa2d... shows some of the interface, but I seem to recall getting notified in webmaster tools when I messed up the robots.txt on a particular site.
I haven't read through all of the code but it assuming this is actually what's running on Google's scrapers this section [2] seems to be pretty conclusive evidence to me that this Noindex thing is bullshit.
[0] https://www.deepcrawl.com/blog/best-practice/robots-txt-noin...
[1]https://www.stonetemple.com/does-google-respect-robots-txt-n...
[2] https://github.com/google/robotstxt/blob/59f3643d3a3ac88f613...
:D
Ah, the southern version. :)
If you want to make sure a URL is not in their index then you have to 'allow' them to crawl the page in robots.txt and use a noindex meta tag on the page to stop indexing. Simply disallowing the page from being crawled in robots.txt will not keep it out of the index.
In fact, I've seen plenty of pages still rank well despite the page being disallowed in robots.txt. A great example of this is the keyword "backpack" in Google. You'll see the site doesn't want it indexed (it's disallowed in robots.txt) but the site still ranks well for a popular keyword).
URLs blocked in robots.txt can get discovered through other links and they will get displayed in the search results.
However, you will not see any information like the meta description on these blocked URLs.
There's a good explanation about this here, including a video from former Googler, Matt Cutts: https://yoast.com/prevent-site-being-indexed/
True, but that's not the only thing. If it ever was in the index, it takes forever to be removed, if it gets removed at all. Send 404 or 410, Disallow it or set it to noindex - you may get lucky or you may not. You can of course "hide it from search results", but that only works for 90 days (iirc, may be 120, something in that range). Those leftovers will typically lose rankings, but they often stay indexed, easy to spot with a site: query.
Percolator white paper: https://ai.google/research/pubs/pub36726
Granted, these are edge cases, in most circumstances, 410 + 90 day hiding means they are hidden instantly and don't resurface. These edge cases do make me take Google's official statements on how to deal with things with a grain of salt though: bugs exist, and unless you happen to know somebody at Google there's no way to report them.
https://www.searchenginejournal.com/google-404-status/254429 "How Google Handles 404/410 Status Codes" -- "If we see a 410, they immediately convert that 410 into an error rather than protecting it for 24 hours"
Which site? [Edit: I have now found https://www.gcsbackpack.com/ on page 6 of the results, and this was presumably the intended site.]
It’s bad enough I started using DDG for search because the results are now more relevant. Google’s advertising algorithms are designed to subtly nudge sites into paying for placement — which means there’s a “non-content” element to the search results that makes it into the user experience. I feel like there was a tipping point a year or two ago where the results just stopped being useful — The best analogy I can find is how search engines used to be in the days before AltaVista. Then AltaVista came out and the results were far more relevant (if not perfect). Google -> DDG feels like that in 2019.
That “non-content” element will only grow over time as Google seeks revenue growth — growth across all of Google’s non-advertising revenue streams combined are not enough to move the needle compared to the scale their ad business has — of which search ads are by far the most profitable. So they will further try to monetize search; it’s their cash cow but I think a small player like DDG could easily overtake them as the quality of Google’s search results (to the end user) continue to decline.
Basically Google finds the link in other places -> oh that must be interesting, I'm indexing it, without even reading it. So they don't have the actual content, and just use the texts from the sites that link to it.
Google is also slow to honour 404 and drop pages which can hang around for ages, Bing is much faster to remove 404 pages.
[0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/410 [1] https://support.google.com/webmasters/answer/6065812?hl=en
It makes more sense when you realize that the SEO people (with a few exceptions) are usually pretty shady as well. You rarely hear them recommending that you write better content to get better results, it's always nonsense like "put nofollow on everything so your score doesn't leak".
But I understand, there's a lot of snake-oil and "one weird trick to rank first" that brings a bad name to the SEO world.
I've seen people go on Fiverr and expect to find top-notch SEOs there.
There's more to SEO than just writing good content. There's a lot of technical stuff that can bite you and your awesome content will never rank.
Stuff like improving site structure, canonicals, learning to deal with multi-language versions of your content, implementing proper redirects, etc,etc is something that a good SEO should be able to fix and improve.
- Putting important text inside of images
- Duplicate content out the wazoo
- Not making use of canonicals
- No sitemaps, html or xml
- Page performance issues
- Broken mobile support
And of course, poor content. You can't rank if you don't have content.
> Putting important text inside of images I'm sure the reason for this is that it's hard to parse text from images, and while Google could use their AI to figure it out, they don't bother. But it also prevents blind people from being able to read the text, so it does worsen the experience. > Duplicate content This makes the site harder to navigate for users as well. > Page performace issues Quite obviously makes the experience worse. > Broken mobile support. -..-
No site's output is 100% because of the tech team - content writers can put in weird code, marketers can add all sorts of stuff to say Tag manager, the robots.txt is likely from 2008. And a site built with code as the primary goal is likely lacking in some marketing oomph somewhere.
Someone who's job it is to find the right balance, and aim to maximise the returns from the single largest source of traffic, is pretty valuable.
Google has now clarified that they're removing the code behind the undocumented items, with noindex called out explicitly.
https://webmasters.googleblog.com/2019/07/a-note-on-unsuppor...
It wasn't officially supported / the recommended way - but it worked (in many cases.)
For those who (like me) don't know a lot about this, which side of the argument is bullshit? Have you just been proved right or wrong?
As far as I can tell the inception of this idea was that it was briefly mentioned by some Google employee in an interview. Maybe it was supported in the past or maybe he just misspoke, but I bet even now we'll see people still using this tag.
That seems like a can of worms not really worth opening.
If you'd like a specific example of why people might seek this courtesy, someone might have a page or group of pages on their site that works fine when used by the humans who would normally use it, but which would keel over if bots started crawling it, because bot usage patterns don't look like normal human patterns.
By analogy: humans drive cars and cars can respond to human problems at human time-scales, and so humans (e.g. pedestrians) expect cars to react to them the way humans would. But there are other things on, and crossing, the road, besides cars. Everyone knows that a train won't stop for you. It's your job to get out of the way of the train, because the train is a dumb machine with a lot of momentum behind it, no matter whether its operator pulls the emergency brake or not.
There are dumb machines on the Internet with a lot of momentum behind them, but, unlike trains, they don't follow known paths. They just go wherever. There's no way to predict where they'll go; no rule to follow to avoid them. So, essentially, you have to build websites so that they can survive being hit by a train at any time. And, for some websites, you have to build them to survive being hit by trains once per day or more.
Sure, on a political level, it's the fault of whoever built these machines to be so stupid, and you can and should go after them. But on a technical, operational level—they're there. You can't pre-emptively catch every one of them. The Internet is not a civilized place where "a bolt from the blue" is a freak accident no one could have predicted, and everyone will forgive your web service if it has to go to the hospital from one; instead, the Internet is a (cyber-)war-zone where stray bullets are just flying constantly through the air in every direction. Customers of a web service are about the same as shareholders in a private security contractor—they'd just think you irresponsible if you deployed to this war-zone without properly equipping yourself with layers and layers of armor.
In my younger years the only time I ever dealt with robots.txt was to find stuff I wasn't supposed to crawl.
For instance it explicitly says "To exclude all files except one: This is currently a bit awkward, as there is no "Allow" field."
And the behavior is so different between different parsers and website implementations that, for instance, the default parser in Python can't even successfully parse twitter.com's robots.txt file because of the newlines.
Most search engines obey it as a matter of principle but not all crawlers or archivers [1] do.
It's a good example of missing standards in the wild.
[0] https://www.robotstxt.org/robotstxt.html
[1] https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
That is changing, and was announced today: https://news.ycombinator.com/item?id=20326067
Disallow
The D got mangled and that disallow directive got ignored.
When I read "This library has been around for 20 years and it contains pieces of code that were written in the 90's" my first thought was "that commit history must be FASCINATING".
This may[0] be because it is exported from there monorepo
Whilst I am sure there are good reasons for the omission, it would have been interesting to see the entirety of the commit history for this library.
From Google's perspective it's probably too much work. I would assume this was a part of the cralwer code and extracted over time into a library, while part of the monorepo, so changesets probably didn't only touch this code, but also other parts and this code probably depended on internal libraries (now it depends on Google's public abseil library) publishing all that needs lots of review (also considering names and other personal information in commit logs, TODO comments and their like)
https://github.com/google/robotstxt/blob/master/robots_test....
// A user-agent line is expected to contain only [a-zA-Z_-] characters and must
// not be empty. See REP I-D section "The user-agent line".
// https://tools.ietf.org/html/draft-rep-wg-topic#section-2.2.1
So you may need to adjust your bot’s UA for proper matching.(Disclosure, I work at Google, though not on anything related to this.)
Of course, in practice robots.txt tend to look less like [1] and more like [2].
[0]: https://tools.ietf.org/html/draft-rep-wg-topic#section-2.2.1
What do huge robots.txt files like that contain? I tried a couple domains just now and the longest one I could find was GitHub's - https://github.com/robots.txt - which is only about 30 kilobytes.
Or they have a ton of auto generated pages they don’t want crawled and call them out individually because they don’t realize robots.txt supports globing.
http://www.antipope.org/charlie/blog-static/2009/06/how_i_go...
(reminds me how Y Combinator's co-founder Robert Morris has a bit of youthful notoriety from a less innocent program)
[1] and former code monkey from the dot-com era
1. https://github.com/google/robotstxt/blob/master/robots.cc#L6...
But I'm sure someone out there will fuzz it...
https://opensource.googleblog.com/2019/02/open-sourcing-clus...
Fascinating.
Blocking in robots.txt will stop Googlebot downloading that page and looking at the contents, but the page may still make it into the index on the basis of links to that page making it seem relevant (it will appear in the search results without a description snippet and will include a note about why).
To have a page not appear in the index you need to use a 'noindex' directive [1] either in the file itself or in the HTTP headers. However, if the file is blocked in robots.txt then note Google cannot read that noindex directive.
Also, in the StackOverflow response you linked to that the user agent is listed just as 'Google', but it should be 'Googlebot' as per the 'User agent token (product token)' table column listed in [2].
Good luck! :)
[1] https://support.google.com/webmasters/answer/93710?hl=en [2] https://support.google.com/webmasters/answer/1061943
Can you imagine how many billions of time this code has been executed? I love software like this.
Honestly, excessive cleverness does not generally pass code review @ Google. Especially something that would get this many eyes.
But it being old and critical, I'd also be wary of major changes.
I imagine these days it’s been incredibly hardened and is additionally sandboxed. But back in the day?
And yet the Google style guide literally says: "Assume the person reading the code knows Python better than you do."
https://github.com/google/styleguide/blob/gh-pages/pyguide.m...
If you're going to have to explain it at the next code review, you should comment it now. Complicated operations get a few lines of comments before the operations commence. Non-obvious ones get comments at the end of the line.
The section you're quoting says:
On the other hand, never describe the code. Assume the person reading the code knows Python (though not what you're trying to do) better than you do.
By "clever code" we're talking about weird unidiomatic tricks and hacks that maybe writes things in a slightly shorter or in a fractionally more (unnecessarily) optimised way, and makes you feel clever, but makes it harder and more time consuming for anyone else to understand what your code is doing, or verify that it's actually doing what it's supposed to.
What are the chances that Google is releasing this as a preemptive response to the likely impending antitrust action against them? It would allow the to respond to those allegations with something like, "all the technology we used to build a good search engine is out there. We can't help it if we're the most popular." (And they could say the same about most of their services: gmail, drive, etc.)
There's already https://github.com/temoto/robotstxt
I mean, I get it; it feels that way to me intuitively too. But I'd still recommend against trying it, because I've learned the hard way the intuition here is, if not wrong, at the very least very badly underestimating the cost, especially in the "unknown unknown" department.
Adding a cgo dependency is generally something that isn't done lightly by teams. Having a port to go instead of a wrapper around go would be much more likely to see widespread adoption.
That's why I'm saying there's no point trying to re-implement this. If you were going to re-implement this, there's probably already a library that will work well enough for you. The value here is solely in being exactly what Google uses; anything that is a "re-implementation" of this code but isn't exactly what Google uses is missing the point.
If they formalize it into a spec, others may then implement the spec, but they can and should do that by implementing the spec, not porting this code.
This seems like a weird assertion. The specification isn't particularly complex (ignoring the implicit complexities of unicode). There are ~5 keywords and like 3 control characters. Why would you expect to need all that much?
To be more direct: what are all of these assumptions you assume google's parser is mishandling?
I had thought most of the systems code inside Google would be golang by now. is that not the case ? the code doesnt look too big - I dont think porting is the big issue.
Rewriting decades of core business logic would be a tremendous effort and amount of risk.
Java and C++ still run most of Google.
Why do it in the first place? Just because you can? The code works and it's written in a popular language which plenty of people know. What's the upside?
What would this mean for C++, if not an `extern C` interface?
Google has gazillions of lines of system code already built. Why rewrite everything in go? There is so much other stuff to do. All rewriting achieves is add additional risk because the new code isn't battle tested.
Depends the context, but in general, yes. C++ is very close to C on this aspect, trading memory safety for performances.
Concerning google, as far as I know the codebase is mostly C++, Java, and python. Go will surely eat a bit of the Java and Python projects but it’s unlikely to see C++ being replaced any time soon.
I don't believe this is the case. Most optimized, natively compiled languages all perform similarly. Go, C, CPP, Rust, Nim, etc. I'm sure there are edge-cases where this isn't the case, but they all perform roughly the same.
The performance rift only starts when you introduce some form of a VM, and/or use an interpreted language. Even then, under certain workloads their optimizations can put them close to their native counter parts, but otherwise are generally slower.
The real reason Google didn't re-write this in Go is likely because the library is already finished, it works, a re-write would require more extensive testing, etc. Why spend precious man-hours on a needless re-write?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Like REALLY behind, this surprises me a lot actually. Thanks for showing me that awesome benchmark page :D
How many seconds do you suppose startup time is for those tiny programs?
How many seconds do you suppose JIT is for those tiny programs?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
This argument to me is usually comes from people who have not done projects of significant scale or that required high performance, which is fine not everyone works on that level of a project. But the small difference of 10ms per operation when having to do a million operations is nearly 2.7 hours of extra time. Even 1ms is an extra 0.25/hr in time. These things start adding up when you are talking about doing millions of operations. And there is nothing wrong with Go or Rust or Python, just they aren't always the right tool in the toolbox when you need raw performance. Neither is C/C++ the right tool if you don't need that level of control/performance.
When doing distributed systems or embedded work you generally learn these rules quickly as one "ok" performing system can wreck a really well planned system, or start costing a ton of money to spin up 10x the number of instances just because of one software component isn't performant.
It's still somewhat early but I do already see software being written in Rust with best in class performance (take ripgrep for a prominent example), so lumping it in with Go and Python is really a category error in my opinion.
Personally, I'm still writing C++ for the platform support, etc. but not pretending to like it.
Totally agree C++ definitely has pain points still, but I do love the fact C++ is getting pretty regular updates so it is getting better and less painful generally. Rust is something I want to use in production but haven't seen the right opportunity to do it where the risk to reward ratio was right, yet.
You're certainly right, there IS a performance difference, and in high-computing workloads, such as the one this parser is used for.
From a "regular" web developer perspective (ie. where you only a few servers/VPS's MAX) a lot of newcomers often worry about performance, and usually for most web development the answer with performance is "Yes language [here] is faster then Python/Javascript/Ruby/etc. But those languages/frameworks allow us to develop our application far faster, and ~10ms isn't an issue." Only after performance bottlenecks are discovered would we consider breaking out pieces into a lower level language.
You're completely right though, in HPC it is totally worth worrying about every millisecond, I took the wrong perspective with the implications of the performance differences.
Most of the time, and to your point, that level of performance isn't necessary so using a language that is less likely to let you take your foot off is generally the best & most correct choice. I only resort back to C/C++ when I need the pure raw performance like this parser would, or when doing embedded work. Otherwise I reach for other tools in the tool-bag that are less likely to let me maim myself unintentionally.
Please show why you don't believe this is the case.
The rest of the comment makes some claims about performance, but does not show why we should believe those claims.
For example?
Otherwise we just have: yes it is! no it isn't!
The amount of arrogance in this sentence is insane.
Because Google way is the only one true way?
No really. Microsoft? BSD TCP/IP stack for win95 maybe saved them but there was trumpet winsock and probably would have survived to writing their own on the next release.
Google doesn't get off the ground and has literally no products and no services without the GPL code that they fork, provide remote access to a process running their fork and contribute nothing back. Good end run around the spirit of the GPL there and that has made them a fortune (they have many fortunes, that's just one of them).
New projects from google? They're only open source if google really need them to be, like Go which would get nowhere if it wasn't and be very expensive for google to have to train all their engineers rather than pushing that cost back on their employees.
At least they don't go in for software patents, right? Oh, wait...
At least they have a motto of "Don't be evil" Which we pretty much all have personally but it's great a corporation backs it. Corporate restructurings happen, sure, oh wait, the motto is now gone. "Do the right thing" Well this is fine and google do it, for all values of right that equal "profitable to google and career enhancing for senior execs".
But this is great a robots.txt parser that's open source. Someone other than google could do something useful for the web with that like writing a validator, because google won't. Seemingly because it's not their definition of "do the right thing."
"Better than facebook, better than facebook, any criticism of google is by people who don't like google so invalid." Only with more words. Or none just one button. Go.
The only reason Linux is as mainstream as it is today, is exactly because of this freedom to leverage the code. You even point out that the cause for Golang's success is for precisely the same reason. Overall opensource isn't about making money, it has never been about making money. Its been about making an impact, and bettering the world around us all by giving a piece of technology to be freely used by everyone. There are a variety of opensource licenses that can/will protect your code from any/all closed source uses, for example AGPL explicitly states if your application so much as interacts with the code over a TCP connection or furthermore a single UDP packet it must be opensource as well. However you will rarely see libraries/applications using this license. Why you might ask? The answer is simple, it reduces the impact that code can have.
Really at the end of the day, it comes down to a choice of the developer(s), do you want to make money? i.e. go the Microsoft/Apple route? or do you want to make an impact? i.e. go the Linux/BSD route?
Let me ask one final question, which of the above operating systems do you think are more widely used, or have changed the world in a more dramatic manner?
Google is built on an end run around the spirit and intent of the GPL. "Don't distribute software, distribute thin client access to it! No GPL! Hurrah! Money!"
Decide for yourself what you think of that but it happened. Without it, no google.
But hey, list anyone you think derived more value and contributed less back. It's a reasonable thing to do. Doesn't affect criticism of google.
Oh and your link? That's a propaganda site heavy on aesthetic design and basically devoid of fact.
Apart from that it's a really strong response. Do you love google? Work for them? Reflexively stick up for big business?