It's only after it happened to him that now he's suddenly against it. Until he removes the same type of blocks from his own website I have absolutely no sympathy for him.
It's only after it happened to him that now he's suddenly against it. Until he removes the same type of blocks from his own website I have absolutely no sympathy for him.
Lets read through that page for a second though:
Drop support for obsolete HTTP versions
Doesn't seem like that's going to cause much issue for any legitimate client from the past 10-20 years. He only recommends blocking HTTP 0.9/1.0, which fair enough Append a #hash to the form’s action URL
Hah. Clever man. I don't see how this is going to stop any legitimate user from loading your website or submitting the form, but I can see how it might frustrate bots. Include a hidden prefilled form field
This is just standard practice to mitigate CSRF. Verify the Host and Origin request headers
Yes. You should be doing that. Set a test cookie and verify it gets included in the submission
Another CSRF trick. Swap the name attributes in the name and email fields
This one's a little user hostile to folks who use assistive devices like screen readers. But still won't prevent you from accessing the site in the first place. Verify the POST/Redirect/GET (PRG) chain
As noted by the author, might cause some issues but again, won't stop anyone from loading your website. Block ancient versions of common browsers
Alright please just don't do this. UA blocking is gross and might prevent access through specialist software. But he also calls this out himself. I strongly discourage you from blocking or discriminating against unknown or uncommon browser User-Agent request headers
All in all, with the exception of UA blocking I don't see how any of these mitigations would result in users not being able to access said website, or having their loading times drastically increased.It's an arms race/defense-in-depth situation. If someone truly wants to automate your site in a targeted fashion, and it's profitable for them to do so, you'll have to invest a lot more in stopping it (and decide how much of it is worth stopping).
My college English professor was so obsessed with beating cliff notes that all of tests were hyper focused on the most obscure details he could dream up. It was just impossible to maintain a full course load and memorize what was the 1st, 3rd, 5th, 7th , last, etc word on every page and which character spoke it. Did the sentence contain any commas? How many times was the nurse mentioned in chapter X, etc.
Google can always follow his lead and make their data impossible to access, but impossible doesn’t increase ad views, so they will never do it. People using ytdlp is just the cost of doing business.
Using NGinx as an example:
if ($server_protocol != HTTP/2.0) { return 403 'Nope'; }
Another thing I have found useful to drops some bots is to become invisible to them. Many of the poorly written scanning tools do not properly set MSS for reasons I still don't understand. I use this to my advantage.Using IPTables as an example:
/sbin/iptables -t raw -I PREROUTING -i eth0 -p tcp -m tcp --tcp-flags FIN,SYN,RST,ACK SYN -m tcpmss ! --mss 420:16384 -j DROP
Any TCP packets setting a very low or high MSS or missing MSS will be silently dropped. I drop about 35K packets per host per day on average. This also drops hping3 floods.MSS issues attract me like a moth to flame [1], so let me ask some questions.
It looks like this is dropping syns with MSS over 16384??? That is indeed a pretty crazy high number. 9000ish seems reasonable for someone on a jumbo network without a mss clamping router, but above that is someone weird for sure.
Under 420 seems unlikely too, but technically acceptable, but sure, I'd drop it. In theory, a proper OS will send several SYNs with MSS, then assume your server doesn't support TCP options and send you a SYN with no options. Going to take a while, but if someone legitimately has a mss less than 536, their internet is probably pretty junky anyway, so ok, seems fine.
[1] I just built a browser based pmtud test site, http://pmtud.enslaves.us/
All of this said, I could set the range to 1:65536 and it would still drop most bots as they don't even bother to set MSS at all in their scans. I'm not sure which tool they are using.
Oddly enough most of the bots in Asia are exactly 1398 and most of the bots in Russia are 1424. I find that interesting.
As long as you're using a <label> or aria-label attribute, that shouldn't be an issue.
(Author here.) If I remember correctly, his browser of choice predates the Origin header.
In general though the whole tone of parent of “I am owed access to someone else’s computer system on my and my terms alone” just doesn’t jive with me. It’s also not remotely comparable to Cloudflare’s approach of sitting in the middle snd then appropriating end-user compute resources without their consent to fuel their business.
I just have no sympathy for Daniel since up until just now he was trying to get everyone to do this.
> Just about every website I visited from my home internet connection would result in a challenge page.
That capability is only available for paid CloudFlare plans.
He also specifically called out CAPTCHA as user-hostile.
False positives happen. They happen a lot more than you think. And they are a serious problem. Even more serious when it's cloudflare, but arguing for everyone to implement these algorithmic blocks "that won’t inconvenience users" individually, taken to it's logical end, does the same.
The blog post also calls out that you should not block based on user agent.
If a form post didn't respect the action property having a #, that name/email HTML names might be reversed (whole the type is correct, and the user displayed values are correct), or include hidden HTML form fields that have been standard since ~97? Back when I made my first few websites, I certainly would agree that they are likely bots.
Again, apparently this person has some hateful following, but I don't appreciate you limping me into this hatred for agreeing with his statements on this one particular issue.
You and others can keep quoting the legit and clever ways to mitigate bot spam but if you ignore the false positives the other checks create it kind of defeats the point.
> Bots often mimic the User-Agent of a common browser, but the version numbers used in the bots rarely change. Over time they drift farther and farther behind until a point (maybe two-year-old versions) where you can safely block them without inconveniencing legitimate users.
This supports the idea that browsers are subject to constant change and everyone should be forced to come along (rather than respecting and supporting standards). I have a Chromebook that stopped receiving updates some years ago (thank you for your very safe and sustainable product Google!), his heuristic would litteraly block me.