Google's robots.txt
google.com
google.com
Allow: /maps?hq=http://maps.google.com/help/maps/directions/biking/mapleft.kml&ie=UTF8&ll=37.687624,-122.319717&spn=0.346132,0.727158&z=11&lci=bike&dirflg=b&f=dTurns up three results in google. Very weird indeed.
Apparently some security scanner tests some php vulnerability based on fopen trying to load that.
% curl www.google.com/humans.txt
Google is built by a large team of engineers, designers, researchers,
robots, and others in many different sites across the globe. It is
updated continuously, and built with more tools and technologies than
we can shake a stick at. If you'd like to help us out, see
google.com/jobs.Not tested. :)
THIS is evil. You could use this argument for banning any new search engine.
This is the reason they're blocked, not because they're new or non-English.
This site coolshell.cn is boycotting baidu, so it tells Baiduspider that it doesn't want to be indexed by baidu.
Baidu respects this and doesn't index anything from this website.
http://help.yandex.com/webmaster/controlling-robot/robots-tx...
Yandex's online tool to check specific URL on allowness in robots.txt:
I added that to the last.fm/robots.txt many years ago (http://www.wired.com/business/2010/08/robot-laws/all/), and have just been made aware that it appears to have spread across the internet:
https://www.google.co.uk/search?safe=off&q="Disallow%3A+%2Fi...
Favourite sites I've found with it so far: php.net, princeton.edu, nest.com, songkick.com.
https://www.google.com/maps?hq=http://maps.google.com/help/m...
What a weird little product. It's like Yahoo Answers, but somehow with even less sorting or categorization.
They have an SMS companion service that only works in Ghana..
Baraza never really took off.
More about it here: http://whiteafrican.com/2010/10/05/google-baraza-qa-for-afri...
- why is computer an idiot machine
- i casted a love spell on my ex,should i tell her now that she is back
- What is the colour of the black box which is using in planes ?
In practice, I'm not entirely sure but it looks like it's quite old since those pages don't seem to exist anymore, even as folders which don't exist as pages in themselves, and they aren't 301-redirected to the current relevant pages.
In fact, they're all 404's. So perhaps they used to be pages, were deleted, and kept being crawled which made their site look crap (because of the 404's). Now they could use 301's, but I assume that the reason they didn't is because they might want to restructure the site in the future and re-use those pages. They don't use 302's because 302's are unreliable and freaky.
Does that sound right to everyone else?
Is this a common mistake?
Content-Type: text/html;charset=UTF-8
It also tries to set no less than four cookies.Wow, really? Who put up this sign?
> permission. See: http://www.facebook.com/apps/site_scraping_tos_terms.php
I feel like once a company allows public access by posting stuff on the web, they can specify terms, but not include/exclude groups specifically. (In a legal sense; I understand blocking systems that hammer servers but will respect robots.txt. IME bing is the worst offender -- they hammer my sites, send no traffic, but will stop if I specify in robots.txt.)
Does anyone have an opinion about "once public, I can crawl"?
Suppose Facebook is getting paid by Bing, and won't offer crawling to those that aren't paying it? Suppose Facebook considers Baidu's crawler to be evil and chooses to prohibit it for that reason? Suppose Facebook just kind of likes the guys at Bing and decides to allow them special access? If you agree in the first place that Facebook should have the right to put ANY sort of restrictions on who can crawl their side, then why should ANY of these be prohibited? This is not a "common carrier" kind of situation.
I someone writes a curl/wget script wrapper & points it to the top 10 websites, they don't enter into any kind of written contract or agreement.
They're telling the public that it does not have permission to crawl the site which try have the right to do. What is the problem with that?
They run their PHP on HHVM for one:
https://en.m.wikipedia.org/wiki/HipHop_for_PHP
https://github.com/facebook/hhvm
...and yeah, it executes PHP code, for sure. But right there, things are already different, and the reality is that they've written a substantial code base in C/C++.
And, two, I'm sure they retain some serious business proprietary trade secrets about their server infrastructure, meaning that while the web front-end might render out HTML like a souped-up CDN, behind the scenes, there is a shit ton of other stuff going down.
Honestly, I think they just leave the file name extensions in the URL for the sake of nostalgia.
Also their frontend is the only thing written in PHP. You hit the site, you're hitting PHP pages.. Not just an extension for nostalgic reasons.
A description for this result is not available because of this site's robots.txt – learn more.
Which is odd.
> # Folks get annoyed when XfD discussions end up the number 1 google hit for
> # their name.
For the smaller guys, sure it makes sense to have some kind of simple robots.txt policy.