Google search indexes itself
google.com
google.com
However, if you search site:http://www.google.com/search and show omitted search results, you get a bunch of results (all 404s).
If you do this there are some strange results on the last couple pages.
For example: Obama won't salute the flag | Phallectomy | horse+mating+video | feral+horses+induced+abortion | Lactating+dog+images | animal+mating+video | mating+mpg+-beastiality+-...
So, Half Life 3 confirmed.
Note that switching //search to /search eliminates the phenomenon.
Note too that all the results on page 1 and page 10 are related to hostgator and coupon codes. I expect that there is some site which contains some text or links that cause these results.
Note also that the `site:` search operator isn't supposed to include anything but a domain or subdomain: no http:// nor /search should be included.
Finally, note that the results are actually google search pages, though! So I do think this is some kind of bug.
But NOT an instance of Google indexing its result pages. Please change the title to 'This one weird google bug will make you scratch your head!' :)
Edit: andybalholm suggests (on this page) that the double slash is in fact causing the googlebot to visit those search results page and indeed index them. Hm, sounds true.
Has anybody visited the spamfodder pages and found instances of malformed yet operative links to google search? (I don't feel like visiting those sites on this machine on this network.)
Correction: this works as you (or a muggle) might expect: https://www.google.com/search?q=site:https:%2F%2Fgithub.com%...
...Though logically the operator should be named `page:` now. :)
But that's clickbait! :)
That's what it looks like to me. Could you explain the difference?
In hindsight, your comment alone would have changed my tune: nope, I can't explain the difference between a page appearing in search results and a page being indexed. Thanks for the illumination. :)
google recommends the site:example.com/path shortcut itself https://support.google.com/webmasters/answer/35256?hl=en
and it's ok to use, as site:example.com inurl:path could mean example.com/hudriwudri/path, too
Traditionally, consecutive slashes in a path name are treated as equivalent to a single slash, presumably to simplify apps that need to join two path fragments -- they can safely just concatenate rather than call a library function like path.join().
Unfortunately, this makes it much harder to write code that blacklists certain paths, as robots.txt is designed to do. Clearly, Google's implementation of robots.txt filtering does not canonicalize double-slashes, and so it thinks //search is different from /search and only /search is blacklisted.
My wacky opinion: Path strings are an abomination. We should be passing around string lists, e.g. ["foo", "bar", "baz"] instead of "foo/bar/baz". You can use something like slashes as an easy way for users to input a path, but the parsing should happen at point of input, and then all code beyond that should be dealing with lists of strings. Then a lot of these bugs kind of go away, and a lot of path manipulation code becomes much easier to write.
But that doesn't in and of itself solve the problem, because "foo/bar//baz" would map to ["foo" "bar" "" "baz"/] without any additional convention.
This is actually not that unusual. this site does not treat two consecutive slashes as a single slash. There are likely others implementation differences.
Certainly in posix consecutive slashes count as one for file paths, but URL query strings are not file paths.
No, I think it'd be more like proto://host/thing?foo&bar&baz (put an =1 on each of those if you like).
Yeah, I'm employing a convention, but so to is the concept of list of strings that the commenter invoked.
https://www.google.ca/search?q=site%3Ahttp%3A%2F%2Fwww.googl...
and no, it's not clickbait and i'm not affiliated with hostgator or any of that other crap.
a few strange points i would like to point out:
the indexed result pages are http:// not https:// - but to my knowledge google forces https:// everywhere.
the double slash issue is probably the reason why googlebot does indeed index this. robots.txt is a shitty protocol, i once tried to understand it in detail and coded https://www.npmjs.org/package/robotstxt and yes, there are a shitload of cases you just can't cover with a sane robots.txt file.
as there are no https://www.google.com/search (with "s" like secure) URLs indexed google(bot) probably has some failsafes to not index itself, but the old http:// URLs somehow slipped through.
but now lets go meta: consider the implications! the day google indexes itself is the day google becomes self aware. google is a big machine trying to understand the internet. now it's indexing itself, trying to understand itself - and it will succeed.the "build more data center algorithms" will kick in as google - which basically indexed the whole internet - is now indexing itself recursively! the "hire more engineers to figure how to deal with all this data" algorithm will kick in (yeah, recursively every developer will become a google dev, free vegan food!), too.
i think it's awesome.
by the way, a few years ago somebody wrote a similar story http://www.wattpad.com/3697657-google-ai-what-if-google-beca... fun enough the date for self awareness is "December 7, 2014, at 05:47 a.m" [update: ups, sorry seems to be the wrong story, but i'm sure the "google indexes itself becomes self aware" short story is out there, but i just can't find it right now ... strange coincident?]
Google only forces HTTPS for certain User-Agent strings. I just tried fetching http://www.google.com with the Googlebot User-Agent string and Google did not redirect to HTTPS.
I reckon this will be fixed in a matter of days, judging by how quickly the latin lorem ipsum google translate thing was sorted out.
Just use http://www.google.com/custom I use either DuckDuckGo or this site all the time, I'd probably switch to DuckDuckGo completely if this search would go down.
Seriously, why don't you let people do this?
Which is nothing wrong on it's own, as long it's protected by good password and doesn't fail to likes of thc-hydra.
They also had some ancient snapshots from 192.xxx range
search?q=site%3Ahttp%3A%2F%2Fwww.google.com
%2F%2Fsearch%3Fq%3Dproranktracker.com%2B%2B
%2BHostgator%2BCoupon%2BCode%3ACOUPON333&pws=0&
hl=en#pws=0&hl=en&q=site:http:%2F%2Fwww.google.com
%2F%2Fsearchhttps://www.google.com/webhp?gws_rd=ssl#safe=off&q=site:goog...
This will help construct a proof of Göogdel's Incompleteness Theorem.
Without being able to find anything in Google, including Google searches, and including that search for Google searches itself, Google is not a completely powerful search engine; however, it cannot be complete and consistent at the same time. There are searches which cannot be shown to be conclusively either in the index, or not in the index.
But not http://www.google.com////search because that's just crazy, come on.
https://www.google.com/search?q=site:http://www.google.com/s...
I got some searches like:
www.google.com/search@q=tetris+sorry+henk
https://www.google.com/search=pupuk+cair+alami
www.google.com/search&q=strobe+trigger+schematic
www.google.com/search@q=transvestites+used+in+rituals (!!!!)
Edit: roland-s found it first :) , and yes, the last pages of results are pretty weird.
site:http://www.google.com//search
but not with[1] site:http://www.google.com/search
[0] https://www.google.com/search?q=site%3Ahttp%3A%2F%2Fwww.goog...[1] https://www.google.com/search?q=site%3Ahttp%3A%2F%2Fwww.goog...
Any ideas why they're doing this?
User-agent: *
Disallow: /search
but maybe //search slipped thorough? site:http://www.google.com/search
, but all the results are considered duplicates and omitted. Hit the button.https://news.ycombinator.com//item?id=8297241
Where in all other cases tested it won't
Is this a server specific stuff? Or it's configurable
http://url.spec.whatwg.org//#concept-url-path http://www.nytimes.com///pages//politics//index.html http://www.bing.com////search?q=site%3Abing.com%2Fsearch%3Fq... https://www.cloudflare.com///index
Furthermore, it's pretty common to rewrite URLs, doing things like adding/removing trailing slashes, whatever. So it wouldn't be too difficult to have it condense multiple slashes into just one.
For example, this link worksfine: google.com//////////////////////////////////search?q=foobar
Google search tries to cover a lot of typos or be pretty user-friendly for people who don't understand tech. I wouldn't be surprised if there's a grandma out there who thinks http://google.com//search is the correct method.