Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
knownagents.com
knownagents.com
I only mentioned it in the assumption they had a VM or shared hosting, in which case it's worth thinking about.
My "audience" is motivated by it being very outdated software. If non of the exploits work they will think I've patched everything.
There must be more interesting targets out there?
If you are capable enough to hack appache the choice of target makes no sense.
My best security layer is that one first has to even find my websites. If you do there are plenty of better targets among them. From the business end however the most secure look the least secure.
It's a shared hosting account, if I try to send bulk email they will immediately unplug it.
That leaves only the glory of pwning the last instance of some obscure cms?
It would probably make me laugh? Everything will be back up within a week. Nothing of value is lost.
Like you said, shared hosting and low value server so you don't mind either way.
My point though, in my prior comment, was attack surface.
With PHP you have outdated code in the ORM/MVC/whatever. Then you have your code. At least with only a web server and static HTML, and entire litany, and the most likely part to get compromised, is not there any more.
In my 30+ years of experience dealing with people getting compromised, it's always been some asshat not doing security updates (eg, using a distro and updating daily). Or worse, just compiling stuff then not updating builds on a daily or weekly basis.
Outside of that, it's been bad PHP code. Or perl. Or whatever.
I think once out of all the times I've been called to clean up a mess, has it been the web server itself. Bearing in mind "I didn't update my OS/web server for a year, and now I got hacked!!" isn't "it was the web server", it's "dumbass didn't do security updates".
Anyhow.
You're not wrong, yes everything is vulnerable. But PHP + framework + PHP code bugs == 9999, web server == 1 of the time.
If it's shared hosting, still a problem, just not theirs.
You're running the wrong stack - I, myself, find that simply having a static file website is enough to cut down on the traffic.
You need to run something other than static file serving to get bot attention.
You can watch a live stream of it here: https://bencevans.io/security/certificate-stream
They're either using Passive DNS logs or a historical dataset.
For my personal website it’s 10x more bots but I barely notice because it’s a few pages.
For the record I thought I had this site behind basic auth.
The scam of "everyone should have SSL" right here, ladies and gentlemen.
New hotness: DNS-less
See? It's quite short.
I have basically 180d entirely on view metrics, they are more or less noise to a small business owner. Did someone buy or not, that's all you actually need to care about.
Even big retail stores are pushing back on crap like KEPLAR/foot traffic tracking, since it doesn't actually change what you do, or impact sales.
Measure sales, measure customer delight, make those the targets.
/s
Monitoring WAN traffic really gets the paranoia juices flowing.
Just build your website yourself as deep in the stack as you can instead of piling up 50 abstractions on top of each other. Some decisions like having your page be accessible by IP can only happen if you use technology like generic http servers (like apache or nginx) from the 2000s instead of implementing the lower stacks and actually thinking about whether that makes sense for a second.
If when you build a website or a backend, your server responds to requests by IP address (for example), you are building a bottom 90% product, and considering most software markets are super top-heavy, (say 1% win), that's ngmi land.
So glad wireguard exists. It just drops all packets unless I authenticate with my cryptographic keys. It's like the computer is not even there.
> IMO that's the same as going on the street door by door and checking if one is left open to steal everything inside the house...
From experience: this does happen regularly in some neighborhoods of some cities in the US, and even that isn't always an enforcement priority. So lack of enforcement on the internet, where most the perpetrators probably aren't even in a jurisdiction with an extradition treaty, isn't exactly surprising.
I'm just collecting the data now to be used to secure some of my upcoming projects, but I would absolutely also like to take it in a direction where it sends the bots into an infinite slow loop, or preferably something that burns as many tokens as possible for them.
I don't really care about the morality of that. I'm a big fan of fighting fire with fire.
Not to throw water on your plan, but the bots I've written intentionally run very slow with respect to each target. When done in parallel, across a wide range of targets, it doesn't slow down the effort at all.
You'll see a lot of deepfield, censys-scanner, visionheight.com, shadowserver.io, etc., but also the usual suspects of Chinese or Russian IPs.
With OpenWRT I use something like this: `tcpdump -i pppoe-wan 'inbound and tcp[tcpflags] & (tcp-syn|tcp-ack) == tcp-syn'`, or alternatively `tcpdump -i pppoe-wan 'inbound and tcp[tcpflags] & (tcp-syn|tcp-ack) == tcp-syn and not port 44000'`, if we have some torrent client running (e.g. here at port 44000) which would mess up the result. I'm not sure it's the best way to handle this, but it's definitely enlightening what bounces off on the router.
Sure their packets will still hit your router, but if they are dropped immediately at least you're not wasting a syn-ack on them.
Keep in mind that this should be paired with an ASN blacklist - MaxMind also has an ASN mmdb for convenience - because IP address to country maps are almost entirely self-declared[0].
For example, Tencent (AS132203), which you almost certainly want to block, has ranges in 73 different countries per [1].
Thanks in advance.
Your firewall vendor should supply you with country lists, just select the known bad ones and drop their traffic. If you have a consumer grade router, you will probably have to configure the blocklists manually.
ISPs sometimes do trade IPv4 blocks and countries to which it belongs do change occasionally. That can become a problem if you were like literally Netflix and someone few nation states over started an ISP.
Eventually, my lets encrypt cert expired and it turns out certbot is run from USA, so the auto renewal failed me.
Here is a "simplified" version in various formats.
Why not just block all the inbound connections you don't need? Is there a particular reason your firewall policy needs to be xenophobic?
(You still could get poked from the other users' hosts behind the ISP's NAT, of course.)
We collect signals that help determine good vs bad networks. For example, large amounts of requests to .php endpoints, large amounts of empty accounts from the same /24 subnet, etc etc. All these signals let us automatically determine risk, and then put up a challenge. Authenticated users never see the challenge even if they are on a risky network (VPN 99.9% of the time), unless the network has been identified as 100% malicious, then it gets a full block.
Here's a small snapshot of the dashboard:
https://cos.ridewithgps.com/screenshots/6a7c54d0-12Aug26-358...
This was probably a total of 3-4 days of work, spread out over a couple months of iterative claude led hacking. I didn't know exactly what to build, but had some of the key architectural ideas in my head. Opus+Faable made easy work of it all, and ended up guiding some really slick improvements for performance.
I would say this has dropped about 20% of all traffic to our service, though it turns out turnstile is a massive target for bots, so replacing that with something custom is next on the list.
You know all these guys just switch to residential proxies if they detect a site is blocking data centers, right? Because that's a very common thing to do.
As for the latter comment....not sure what your implication is. Yes, bot/spam mitigation is whackamole, but there are consequences for not playing the game of whackamole. Luckily residential proxies are few and far between so far, but they will grow in popularity. When they do, and I can't get by with the occasional individual residential IP ban, we'll come up with other methods to handle.
Luckily the signal is strong with vulnerability scanning, which makes it pretty easy to automate. The only reason to put up whole ASN mitigation (captcha/turnstile, outright bans) is just efficiency. Nothing stopping individual IP banning. The scrapers are the tricky ones, since they more easily hide in legit traffic. However legit traffic has patterns that scrapers do not emulate (at least for a service like ours with millions of pieces of user generated content that's easily walkable), so you can still pull out the signal. It's just a little trickier.
Definitely a continual arms race though.
From a technical perspective, all this "china/russia" attribution is built on a quite shaky foundation. As a sysadmin you'd never know if it would be the British crown attacking your European company instead.
Not minimizing nation state cyber crime here, but the packet goes through many hands with different incentives.
My traffic from Germany passes through a British-owned hop on its way to Switzerland. My German ISP is British as well so either way it wouldn't make a difference, they basically have all traffic twice.
Blocklist download and configuration: https://knock-knock.net/blocklist
Honeypot dashboard, where you can see attempted attacks in realtime: http://knock-knock.net
It never ceases to amaze me that these ISPs don’t bother to shut down the botnets. They could do so very easily. For example, they could identify the IP address of every bot that hit this honeypot with their ASN with one API call: https://api.knock-knock.net/check-asn?asn=<asn number>. (See https://knock-knock.net/api). They just don’t care!
You take IP down, you kill the cancer but you also end up killing the patient.
However, I can see the argument for giving the customer 24-48 hours to resolve the problem.
I'd say its probably an acceptable casualty in the battleground that is the internet especially for little one-off sites hosting blogs, forums, chat servers, etc... For a bigger site I would expect that person may have to open a ticket with the platform such as Amazon accepting that some CDN's and firewalls may be harder to get the block removed. This is why we can't have nice things.
IP Blacklists, no matter how good can't stop this. You have to start using stats or deep-diving telemetry.
So appeal to emotion doesn't fly with me. If grandpa is 76 in the year of our lord 2026 that means he was 50 when the internet was getting popular and 59 when cell phones became very popular on the internet. He's not much older than I. He knows what's up.
God help the makers of that television if he finds out it has been spying on him and dorking around with his traffic. If they are lucky he will just take a baseball bat to it. If they are unlucky he will fly to their headquarters and end up on a viral bodycam video likely with a lot of supporters that will bail him out of jail.
IP Blacklists, no matter how good can't stop this. You have to start using stats or deep-diving telemetry.
I use a myriad of methods including IP blacklists. That's my choice and every site operators choice. I do not have to use deep-diving telemetry but you are free to do so.
[1] - https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Som...
Edit / Update: It was Apple's Private browsing mode that causes it not to respond. I can now see it when this is disabled.
> block http 1.1, real users only use 2.0
Chrome on android and Firefox on linux both appear to use 1.1 still...
There are some reader apps that act as a proxy that only support http/1.1. Be careful, some of those are not just readers and do not trust what they claim to be the source code. Some of them are created by cute and fuzzy bunnies.
There are a number of botters on HN, some that control residential and phone browser-hijacked systems. One was sending me playful messages the other day. I enjoyed the bot block-jousting with them.
No, because legitimate users do not just use residential and "commercial" IPs. Like me, right now
I am going to move full blocking to a test node that people can play with but I have to finish working with Claude to revise someones repo is is no longer maintained because one does not simply put an anonymous chan board on the great wide open internets without some critical thinking.
The complete hysteria people go to over the near-non-issue of bots is, well, completely hysterical. Don't be that guy.
If you have fail2ban or NGINX logs, you can use our CLI to summarize those IPs and identify the ASNs you want to block. But before you block entire ASNs, make sure they are not classified as "ISP" type. For that, visit our website's ASN page first.
I have quite a few community posts around this approach. https://community.ipinfo.io/
If you have raw logs, you can send them to me as well, and I can review them and provide some guidance.
https://en.wikipedia.org/wiki/Code_Red_(computer_worm)
I remember when 'code red' spread and it had the effect of crapping up the contents of my apache server logs. Fun times.
such as:
GET /default.ida?NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN%u9090%u6858%ucbd3%u7801%u9090%u6858%ucbd3%u7801%u9090%u6858%ucbd3%u7801%u9090%u9090%u8190%u00c3%u0003%u8b00%u531b%u53ff%u0078%u0000%u00=a HTTP/1.0
Best hypothesis I can come up with is to somehow make the AI companies look bad, but they seem to be doing an excellent job at that themselves already by scraping everyone hundreds of times per hour over and over.
bots hitting your site aren't problems per se
Oh, the link is entirely contained in the fragment. So it’s some SPA blog thing. I get it.
Links to posts no longer contain URL fragments, but I kept the fragment-style links working with JavaScript.
Most of the traffic is originating in GCP. We're seeing ~70k req/min sustained from Google Cloud IP space (AS396982). Reported to GCP Abuse, they've been non-responsive so far.
The main distinguishing factor is the reuse of a bunch of legit AI-training bot UserAgent strings. It's clear that the traffic is under the same centralized control because of how it changes volume across thousands of IP addresses simultaneously.
What if we're now moving into a world of strictly KYC. The same way "The Facebook" generated massive revenue by creating a KYC world.
Sometimes it's better to not fight with bots actively but harden environment and only react for the worst offenders.
Seems like the crawler companies would be incentivized to not want to take responsibility for people spoofing their user agents.
Any random bozo can trigger that.
The big problem with AIs is that they can try new paths/payloads very easily, adapt quickly, I'd be more worried about that.
Point a domain at your IP, use letsencrypt, post your domain on Reddit, github, x, etc. The bots will find you.
Thanks for the github suggestion; I'll add a repo for the bot blocker Apache config, and include the above server URL that I'm using to test the blocker.