Nike.com's robots.txt
nike.com
nike.com
Webmasters also put highly specific URLs they don't want the world to see in robots.txt. Also not a great idea.
edit: s/unlimited/unmetered/
True, when there is no deliberate cap stated, it's unmetered, not unlimited, meaning you can use what however much fits through the pipes (usually best effort instead of guaranteed bandwidth), and that's the natural limit. A 1 Gbit/s link for example will get you close to 300 TiB in each direction in a 30 day month if you are able to saturate that link 24/7, while a 100MBit/s link will get you a tenth so "only" close to 30TiB.
There is often a “fair use policy”, and even without anything like that you need to consider the potential effect on the stuff you are actually running the service for (while a search engine is pulling random data as fast as they can, that transfer is competing against other network load). I have a couple of inexpensive hosted servers with genuinely unlimited bandwidth, but only at 100 and 250mbps respectively, better rates cost a fair chunk more without some sort of cap (usually of the form “after some-TB-or-some-tens-of-TB we'll throttle you to 100mpbs until your next billing month”.
Unlimited throughput on gbit/more lines is commonly available, but never both cheap and reliable.
It’s been a long time since you could call OVH “unreliable”, Hetzner chugs along just fine too.
I think you might just be stuck in 2004, reliable unmetered gbit+ has been cheap for years.
For example Deutsche Telekom: They don't accept free direct peerings. Sending data to them indirectly frequently runs into congestion issues. For a while Hetzner offered a paid option for better connectivity to DTAG. It looks like Hetzner added a direct peering in 2020.
https://web.archive.org/web/20200205014018/https://wiki.hetz...
I suppose that depends on what you call cheap, and whether you include “guaranteed”/“dedicated” in the definition of reliable for bandwidth.
I'm looking at things mainly from a personal projects & packups PoV, but demanding at least two drives for RAID1+ rather than the cheapest units. If you have money-making projects then the lines between cheap/inexpensive/reasonable might move.
> OVH / Hetzner
I have Hetzner down as “inexpensive” rather than cheap, assuming you count the older hardware in their auction lines (with their new-kit offerings being “reasonable”).
OHV is cheap if you are considering their lower-spec offerings, but those are at limited rates (100mbit usually in the case of Kimsufi branded services, 250mbit for SYS) and IIRC that rate is not guaranteed (while not a massively oversold resource like you'll see in many VPS/shared/similar hosting arrangements, there are enough machines sharing a larger resource that if many try saturate their allocation at the same time they'll hit congestion even within the DC rather than just when your traffic touches external peering).
(Don't take the above as a criticism of Kimsufi/SYS - they are honest & open about the limits of the services they sell (older hardware, limited max network throughput) - I get what I pay for, which is no less than promised, and I've found them to be reliable over the years I've had services with them)
Then there are no cheap options. Even frivolously expensive “cloud” providers don’t give you this.
The reality is that there are very few customers that actually have a real need for “guaranteed”/“dedicated” bandwidth.
(a) may not be that much, in reality, depending on how their network is set up, and how much other traffic there is at any given moment. It also depends where you're routing to, whether e.g. (i) your neighbour or (ii) a DC across the world (neither is clearly better, it all depends). You have to think of it through the lens of physics and information theory, rather than law and business models.
If he were only talking about meters or caps, then that would therefore make his point a bit of a non sequitur in the context of an argument about whether tarpitting/blackholing a connection is risky for the serving party.
(FWIW, at past companies I've not-infrequently encountered network saturation even without an artificial cap from whomever we were peering with.)
I’ve heard this false notion so many times on HN and every time I check it out, the provider says clearly that they don’t allow for unlimited bandwidth.
So please tell me the provider. I’ll pay for a month of service for 10 Gbps unlimited service and then I’ll saturate that line up and down stream for an entire month.
And we’ll just see what happens, ok?
Unmetered rather than unlimited of course, but OVH, Online.net, Hetzner. These are rather well-known providers.
I haven't checked all of them for the exact language but Hetzner explicitly says "All dedicated root servers have unlimited traffic" for example.
Of course, they do have a limit for servers with 10 Gbps connections probably because of people like you who like to push the limits beyond what is reasonable.
I have extensive experience maxing out unmetered 10gbit lines with these providers 24/7. It’s been years and years since a host last tried to FUP me, bandwidth is just really cheap now.
A search engine pulling data from /dev/urandom as in the post that started this discussion is likely to impact other uses of the 100mbit link from Kimsufi. Worse if bad luck means multiple hits of that sort at the same time.
Check out https://discord.com/invite/7Gv8tdM
There are really good deals available, all of these still come with a decent margin for the reseller.
Core i3-2310 8GB RAM 2x2TB HDD 1Gbps Unmetered €16
Core i5-2300 16GB RAM 2x2TB HDD 1Gbps Unmetered €24
E3-1225 16GB RAM 2x2TB HDD 1Gbps Unmetered €35 (iGPU enabled)
W3520 16GB RAM 2x2TB HDD 1Gbps Unmetered €40I was hitting it plaintext first, so a simple redirect to some subdomain instead of a bare redirect to https would probably work fine.
I don't trust letsencrypt and I don't want to give them or anyone else a list of which subdomains I use.
The certificate is only for "*.marginalia.nu", which simply doesn't cover "marginalia.nu". It should give the same error on any platform and browser, unless their SSL implementation is broken.
Some browsers try to be smart and insert www automatically though.
https://marginalia.nu = https://crt.sh/?id=6046506678 (includes wildcard but NOT apex)
https://search.marginalia.nu = https://crt.sh/?id=6125359537 (includes both wildcard and apex)
Don't really understand their motive either. Maybe they thought I was cloud hosted or used some expensive API to do searches and were attempting to rack up big bills or something.
I wish I had a good solution to this. Cloudflare to mitigate DDOS attacks has somewhat of a “baby with the bathwater” vibe (considering how much of a pain it is if you can't pass the CAPTCHA, or if you're on Tor), but I can't think of an alternative.
I've considered having a naked endpoint with a rate limiter that, when it hits some ceiling (dunno, sustained load of 2 RPS or so), offers the alternatives of going through an unlimited cloudflared domain, or waiting until the bots give up. But that might be annoying too.
Oof, I'm not sure if I would do that. Since "" stands for all sub domains.
When I visit https://search.marginalia.nu I'm served with this different cert, which does include both wildcard and apex: https://crt.sh/?id=6125359537
I serve my robots.txt files with Transfer-Encoding: gzip and Content-Encoding: gzip, chunked and then just serve a two layered 10G gzip bomb (240kB file size) instead.
Can recommend. Works pretty often.
¹anything .php,.asp and some other tech we never use, also phpmyadmin, cgi-bin, and some more, actually.
Drupal, PHPBB, etc all have such generic config defaults online. For frameworks such as Rails, Django etc its a bit harder because the logging and url-schemes are not standard, but defined by the person developing in that framework.
nobody own their IP, and at some point a legitimate person will be assigned that IP
and most people who do bad things on the internet use some proxies, usually residential ones
it always makes me sad when people put such rules
This means that commonly only about 30 IPs are on the banlist. The chance that anyone using the same IP as a bot who is visiting any of the apps and sites on that server is not 0%, but tiny.
The only goal is to stop bots and worms from hammering my server. It hardly improves security; it shouldn't anyway.
Is there a way to make it so it doesn't outright make processes crash or lock up, but instead really slow them down / hang them up for a long time?
Also lookup "slowloris attack" -It's pretty efficient if the attacker has more network bandwidth than you have, and can even out the advantage.
Won't be as effective against them having more networked machines though, so you have to use some other tricks against shodan, websnort, burp suite etc.
There are some other tricks though, and you can also abuse HTTP smuggled headers (in the response) to specifically crash their proxies, for example :)
Most common vulnerability scanners use some standard http libraries, so they are vulnerable to slowloris + chunked encodings, or they can't recover from a broken http trailing header, or from an http2/3 upgrade/downgrade loop that breaks their UDP packet parser.
What also works is if you spoof their SOCKS5 proxy with UDP bind requests until they can't use it anymore. So for a lot of DDoS attack scenarios there are active defense options to mitigate the damage which the redteam can cause.
BTW, how's your stealth browser project coming along? (I'd go look but I'm in a weird situation right now where I can't view most of the web. It's browser related so maybe this would amuse you: I'm running Firefox 91.5.1esr on FreeBSD from the official pkg. The vast majority of HTTPS sites give a "Your connection is not secure" error (Error code: NS_ERROR_NET_INADEQUATE_SECURITY) with no recourse (no button to let you make an exception or anything like that. It literally says, "The website administrator will need to fix the server first before you can visit the site.") The fix is to disable SPDY for HTTP2 connections. That works, but then FF crashes (with a core dump, hello FreeBSD) within ~10 seconds no matter what I do. There are a few clues, but I'm just not interested in shaving this yak. Frankly, I'm shopping for a new browser as I find the situation ridiculous. Your project seems way more interesting than what FF has become.)
I think I'd also try to run a UDP tunnel or try to use an amplification redirect on a UDP relaying proxy to get an open port that you can send requests through. Sometimes public DNS resolvers can help to punch a hole through the NAT because their IPs are allowlisted to unbreak secure browsing modes.
> re: stealth
I forked webkit into retrokit [1] these days and am gradually removing all APIs that could be abused for tracking purposes ... with the idea to create a minimalistic webview that is embeddable in a browser and still is able to parse/lex/render modern HTML5 and CSS4.
This way I can later bundle nodejs with the retrokit webview that just displays the Browser UI that's served locally.
Retrokit looks great. :)
I can confirm; I tried to helpfully put a bunch of addresses which would be pointless for a web crawler to crawl in there (links to the login page, etc), but as soon as you say "don't" somebody's gotta go find out why.
I like the idea of putting gzip bombs at commonly-scanned URLs.
# comments
// comments
<!-- comments -->
/* comments */
plain text lines as comments
TeX-markup
script tags
HTML
PHP code
Whatever this is: https://www.costablancaworld.com/robots.txt
search-engine specific directives (Yandex has a few strange ones)
That's an RTF file that was saved as a TXT (without removing formatting)...
Somebody screwed up and made their robots.txt in a rich-text editor like wordpad.
Even stranger things can be found in a `humans.txt`[0] file!
Ah the never ending joy I get by putting a 1GB file as the robots.txt or even having a 1GB `favicon.ico` file.
That’s not really how it works now. Google pretty much ignores robots.txt now and crawls your site anyway.
They can and will crawl even if you say no, they've officially said so as well in their docs.
> A robots.txt file tells search engine crawlers which URLs the crawler can access on your site. This is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.
https://developers.google.com/search/docs/advanced/robots/in...
They won't crawl it, but it they already did crawl it or if there's a link to it, it will still be part of their index (thus search results).
They even go further on their noindex page, by saying that it need to be crawlable (thus not part of the robots.txt) so that they can see the no index directive on the page...
> Important: For the noindex directive to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can't access the page, the crawler will never see the noindex directive, and the page can still appear in search results, for example if other pages link to it.
https://developers.google.com/search/docs/advanced/crawling/...
Honeypot or snarky?
Disallow: /harming/humans
Disallow: /ignoring/human/orders
Disallow: /harm/to/self
> First Law. A robot may not injure a human being or, through inaction, allow a human being to come to harm.> Second Law. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
> Third Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
https://en.wikipedia.org/wiki/Three_Laws_of_Robotics
If you ever get the chance to read a bit of Asimov, I highly recommend it. I'm particularly fond of his Foundation series.
1) don't give a robot a gun
2) don't teach it how to put more bullets into gun after breaking #1
3) don't teach it how to make bullets after breaking #2
3a) don't give it your amazon API keys or other credit card info
(edited for formatting)Wired even wrote an awful article about it at the time: https://www.wired.com/2010/08/robot-laws/
Absolutely delighted people remember it 12 years later.
Everyone Loves Rodrigo
># unless they're feeding search engines.
>[disallows UbiCrawler, DOC and zao]
Anyone understand this bit? They are obedient spiders yet are disallowed because of purpose? Are these obedient but still overly intensive spiders?
Google, facebook and bbc's are surprisingly boring :)
Edit: I've now noticed the disallow directive for other countries specifically targeted for Chinese spiders.
I see the Chinese spiders are only allowed on the /cn/ directory, while disallowed from the root (/).
Another thing I find mildly strange is that the help for certain countries is restricted. This is not what robots.txt should be used for, and it creates incentive for bots to disregard robots.txt.
It seems that they did exist, but moved to specific TLDs (like for Brazil, https://www.nike.com.br). Lazy that they didn't redirect all help files to the correct location though.
It makes sense considering they have this: https://purpose.nike.com/statement-on-xinjiang
https://www.checkbot.io/robots.txt
Also, it's a really common misunderstanding that robots.txt can be used to keep pages out of search results. You need to use noindex meta tags or noindex headers for that. Robots.txt tells crawlers which pages they can't fetch, but they're still allowed to include those pages in search results if they want.
If you also know why we're making some basic mistakes here, especially apply here.
Much better to use x-robots noindex, nofollow for pages you don't want to be in the public domain.
https://developers.google.com/search/docs/advanced/crawling/...
If you are a neutral site (in this case, caters to a lot of people - especially in Asia, where Google as a search engine is a hit-or-miss proposition), then this might backfire sine other crawlers won't understand X-Robots. You can mitigate this by knowing which spiders knows X-Robots (Google, Bing and Yandex AFAIK) and whitelist them while still using a disallow directive for the rest of them.
Why is only poland blocked for all robots? Why is there no /deutchland?
Seems that Bloomberg hosts something for the Polish, but I'm not sure what.
I think it's play on the Nike's famous slogan "Just do it".
That's where the saying "it's not worth my time" comes from. If I asked you why you don't maintain and moderate reddit for free (I'm sure they would be happy to have you there and pay you nothing) I would assume your answer would be I have better things to do with my time or I don't have enough time. i.e. I need to be paid for that to be worth it.
I think moderators time is extremely expensive. They could be doing one million other things with that time and getting paid for it. They could actually be doing the EXACT SAME thing and getting paid for it if they worked for a Reddit/Facebook/StackOverflow or were a community manager for a brand.
You can't just pretend like the key fact isn't relevant.
Platforms don't grow on trees.
Somebody has bills to pay and its up to them how to decide how to pay them. If you don't like their decision, find another platform or moan incessantly about it.
I’m using Safari on iOS 15.2.1, running on an iPhone X.
I'd be interested in browsing these pages.
robots.txt does nothing for the bad actors, but for grey actors it's a bit more complicated as they can move from "grey" to "good". Adding something for that in robots.txt in addition to some other limits (if needed) can be useful.
I'm not even trying to read their files at all, just to check whether my link is 404. Good on W3C for respecting standards, but this means I just have to copy every one of those links from the report into my address bar, like a robot, to check them...