What one may find in robots.txt
xn--thibaud-dya.fr
xn--thibaud-dya.fr
So at least 11 years ago it was blocked.
(I didn't know about spoofing a user agent back then, so it might not have been as easy as that to get around it.)
Google's Advanced search used to be a great tool, until around 2007/08. For some reason it never received an upgrade and several things are broken or don't work any more or were removed (e.g. '+' which is now a keyword for Google+, the '"' does mean the same; e.g some filetypes are blocked, some show only a few results).
But then, so has the www that the original Google worked so well for.
I suspect original Google would be horrible on today's web.
I miss 1998 and I mourn for what could have been.
Certainly feels like content is migrating to the walled gardens and there are fewer and fewer personal websites injecting edges into the open graph.
“This is going to bite them big time in the end, because Google got large by indexing the Geocities-style web, where everybody did have their own web page on a very distributed set of web hosts. What Google is doing is only contributing to the centralization of the Web, the conversion of the Web into Facebook, which will, in turn, kill Google, since they then will have nothing to index.
They sort of saw this coming, but their idea of a fix was Google+ – trying to make sure that they were the ones on top. I think they are still hoping for this, which is why they won’t allow a decentralized web by using SRV records in HTTP/2.”
("term1" and "term2") not ("term3" or "term4")
and whatever tiny "power user" features that google had, like "allinsite:term term term" or '+' don't seem to work at all now.
Google is not optimized for finding things. Google is optimized for ad views and clicks.
I don't see how altavista is superior here
I think I know what they were talking about. A lot of times it appears that adding advanced terms to a query will change the estimated number of results yet all the top hits will be exactly the same. Also, punctuation seems to be largely ignored, e.g. searching "etc apt sources list" and "/etc/apt/sources.list" both give me the exact same results. Putting the filename in quotes also gives the same results as before.
Searching for specific error messages with more than a few key words or a filename is usually a nightmare.
Care to say more? I'm not asking about specifics (but won't mind reading about them either), I'm just curious about the type of information. Internal documents of corporations? .mil? Something else?
It's not rocket science - go to Google.com, type "[some keyword] filetype:doc" (or pds, or xls, whatever), skip to page 10 or so, then start looking. I had to click just twice before I found (5 minutes ago) meeting minutes from some British council meeting, marked 'confidential and not for distribution', complete with names and dates/times. Go to your 'private files' folder on your machine (you know, where you store your job applications, that invitation you made once for your sister's birthday party, that sort of junk), look at the documents you'd least want somebody else to see, identify some keywords that are in those documents but not usually in others, use those as keywords in Google.
Here's a fun one: "site:gallery.mailchimp.com filetype:pdf coupon". Nothing shocking, but I still don't think Mailchimp's customers expect their email list attachments to show up like this... (disclaimer: pure speculation, maybe these are meant to be publicly available, or maybe they're not even real customers and just test data, I read it for the articles, etc.)
For me - at no point was I shown the domain decoded to ASCII (either on HN or in the browser). I recognized the pattern and decoded it manually. For users who are not technical - this is a failed experience because the domain looks suspicious and at no point was it decoded.
I wonder when punycode decoding will begin to get attention from developers. Last year's Google IO had a great talk about how Google realized the inconsistency of their domain handling with regard to I18N:
https://www.google.com/events/io/schedule/session/22ce27dc-7...
Allow: /
# A robot may not injure a human being or through inaction allow a human being to come to harm.
# A robot must obey the orders given it by human beings, except where such orders would conflict with the First Law
# A robot must protect its own existence, as long as such protection does not conflict with the First or Second Laws.
The search result will consist only of the URL, and the snippet will say "A description for this result is not available because of this site's robots.txt"
Use the noindex tag, folks. Also, Google Webmaster Tools allows you to remove URLs from Google's index.
Proceeds to announce the name of a stalking victim. Classy.
Not that I agree it should be further divulged, mind you.
With a fake name there is no proof that the information in question ever existed.
I would've probably masked a name, but institution url should stay there, so anyone could check that the point is valid.
A relatively famous French blogger was convicted once in trial for having publicly reported that a governemental agency was letting GoogleBot index confidential/restricted documents : http://bluetouff.com/2013/04/25/la-non-affaire-bluetouff-vs-... [FR] http://arstechnica.com/tech-policy/2014/02/french-journalist... [EN].
Also, showing the link but not the content is a smart move because it doesn't prove that the author looked for the content of these documents, when the robots.txt is obviously a file you should be able to consult.
I doubt that this is a valid interpretation of the law. The robots.txt file simply mentions parts of the site that shouldn't be indexed, not which parts of a site shouldn't be viewed by the general public. You use a robots.txt file to prevent a search engine from following links that would enumerate all the possible dynamically-generated content on your site to conserve resources and to prevent junk results appearing for a search user.
Tangentially relevant: when you do have something you want indexed, it should probably be a static page that lives at a permanent URL. But never attempt to use a robots.txt file to "hide" sensitive data.
edit: I think you edited to clarify since I commented. Thanks.
There's plenty of others with the same name as her (58k according to Google), that file is now 404, and archive.org doesn't have it (due to its mention in the robots.txt), so whatever information in it that could've been useful for a stalker is long gone.
Her name shows up in a news letter, inc middle name, published and indexed by the school in 2012. It lists her occupation (a very public one) and even where you can find her works. This alone would be enough to get in touch with her.
It was pure morbid curiosity that lead me to search for it, but its totally still relevant if you were looking for her. Unfortunately, given her occupation, its unlikely that anything short of the stalker giving up or getting arrested will grant her any sort of reprieve.
That said, he only reveals that she was stalked which I presume the stalker and her are quite aware of, but no more than that. A better way to report the finding would have been to keep her name and the name of the institution out of it.
At the moment, for me, your post is the only hit on Google for the query you suggest.
Some mobile networks have every person on the network originating traffic from the same IP. Some large institutions, universities, government departments, large companies have all their traffic coming from one IP.
This person has effectively created a feature that will perform a denial of service attack on their own website.
The destination directory doesn't even need to exist. Worst case scenario you could handly the hash via your 404 handler or via .htaccess file if all of your hashes are prefixed. Those are only examples though - there's a multidude of ways you could handle the incoming request.
In short, it's just too much pain for little gain.
Maybe a login form can be served from that URL, and any attempts to login would then get the visitor banned via a session cookie / browser fingerprint combo (Easy to get around but at least then you're not blocking IP addresses).
So add a salt. Just like you would when hashing passwords. You then make it more time consuming to crack the hash than it would be to perform a more typical denial of service attack
> you need to make sure that they can't obtain robots.txt via GET requests from a visitors browser (No Access-Control-Allow headers).
If an attacker already has control over the victims browser (to pull the robots.txt file) then they really don't need to bother with this attack.
> Also don't forget a single visit from a network, sharing an IP would still ban all the network.
Multiple users of the same site behind the same NATing really isn't that common unless you're Google / Facebook / etc. And when you're talking about those kinds of volumes then you'd have intrusion detection systems and possibly other, more sophisticated, honeypots in place to capture this kind of stuff. Also some busier sites will have to comply with PCI data security standards (and similar such as the Gambling Commision audits) which will require regular vulnerability scans (and possibly pen tests as well - depending on the strictness of the standards / audit) which will hopefully highlight weaknesses without the need to blanket ban via entrapment. And in the extremely rare instances where someone is innocently caught out, it's only a temporary ban anyway.
> Maybe a login form can be served from that URL, and any attempts to login would then get the visitor banned via a session cookie / browser fingerprint combo (Easy to get around but at least then you're not blocking IP addresses).
You can do this same method of banning with the honeypot you're arguing against!
While I don't disagree with any of your points per se, I do think you're being a little over dramatic. :)
The login-form method would be a bit less silly, I thought, because it can be a POST. But, well...
The referrer header can be subject to all sorts of subtle edge cases such as switching between secure and unsecure content (or is it the other way around, I can't recall off hand?) which many broswers will then refuse to send a referrer header. So while checking the referrer might work most of the time, it's really not robust enough to be considered trustworthy for anything security related.
Edit: Thanks for the links! :)
Wikipedia: https://en.wikipedia.org/wiki/Uniform_Resource_Locator#Inter...
Edit: typo.
[...]
User-agent: nsa
Disallow: /
From slack.com/robots.txtIt looks like they used robots.txt to do that.
https://web.archive.org/web/20130413152316/http://www.state....
Each line is missing `/documents` in the snippet of the `robots.txt`
Looks like HN needs to learn how to decode Punycode…
I guess using DNS, or by query some engine, like google, or archive.org
is there a service somewhere?
For more try https://commoncrawl.org/
They have some pretty good zone files for major TLDs.