Botspam apocalypse
memex.marginalia.nu
memex.marginalia.nu
I'd recommend:
- Rate-limit everything, absolutely everything. Set sane limits.
- Rate-limit POST requests harder. Preferably dynamically based on geoip.
- Rate-limit login and comment POST requests even harder. Ban IPs that exceed the amount.
- Require TLS. Drop TLSv1.0 and TLSv1.1. Bots certainly break.
- Require SNI. Do not reply without SNI (nginx has 444 return code for that). Ban IP's on first hit that connect without. There's no legitimate use and you'll also disappear from places like Shodan.
- If you can, require HTTP/2.0. Bots break.
- Ban IP's listed on StopForumSpam, ban destination e-mail addresses listed there. If possible also contribute back to SFS and AbuseIPDB.
- Collect JA3 hashes, figure out malicious ones, ban IPs that use those hashes. This blocks a lot of shit trivially because targeting tools instead of behaviour is accurate.
This breaks when multiple users are behind the same IP. I've seen services fail even in classroom, because the prof did something and a few tens of students followed (captchas everywhere).
I do in fact rate-limit everything, it is good advice, but the way you implement rate-limiting allows for traffic bursts. It's basically a reverse leaky bucket, where you start out with N allowed requests, which gets depleted for each request, and refilled slowly over time.
Search traffic is fairly bursty, people do a few requests where they tweak the query and then they go go away.
The only real way to properly police Internet communities is to keep them smaller so that botting is more obvious, and to involve carefully managed moderation. Reddit tried this, but also lost track of the human factors involved and now moderators collect side money and promote their own posts artificially.
The main problems facilitating the surge in bots are shammy creator funds and all the other measures sites take to boost their profit and market dominance. They have grown far too big and can no longer effectively manage their user bases effectively. Things weren't meant to be this way at all, the excessive quest for market dominance and profit has thoroughly corrupted freedom of info online in business, now many users are also following the same road map.
The answer needs a whole lot more finesse than this.
If such a limit would hinder classroom usage, but that's your target audience then other solutions should be found, fairly logical.
Don't ban IPs. Or if you do, let the ban expire relatively quickly (days/weeks, not months/years).
Or instead of strictly rate limiting, ask them a question that can't be "looked up" in a table and that requires human thought, philosophy, emotion, ethics. Maybe GPT could eventually adapt to this and in that case fall back to IP rate limiting and grow the set of questions.
You may be fine with banning those VPN users, or even want that - lots of bots will try to hide behind "legitimate" VPNs - but one has to be aware of this consequence at least, especially considering that more and more people seem to use them - probably also thanks to the aggressive "sponsoring" certain providers such as ExpressVPN do on e.g. a wide variety youtube videos.
It would also stop a not-insignificant number of my customers.
That said, I sympathize somewhat with the parent. "Done right, negative side effects are minimal" is an uncomforting statement. First, because things are often not implemented correctly. There are a lot of details and tuning that will often fail to materialize in practice. Second, because long tail usability issues can often go overlooked. The abuse->anti-abuse feedback loop is pretty tight. Abuse gets identified and counteracted. The anti-abuse-> UX problems loop tends to be noticeably looser. Often, it's just aggregates (revenue/AUD/etc).
Sometimes I just turn it on if I have to fire off a few searches because it'll make me complete a long, tedious captcha for EVERY search
There's a lot of outdated garbage bots out there. Not using HTTP/2.0 is also often the default with various HTTP libraries.
So will people who run older computers with older software. But I guess people who don't have money don't matter for commercial websites so screw 'em.
So, yeah. You're absolutely right. In a lot of cases the loss of revenue from users with severely outdated software will be less than the cost decease of cutting spam and abuse.
This gets back to an old question - to what degree should legacy systems be supported and at what level of expense? There's no one easy answer that works for everyone.
But yeah, corporate employees aren't going to care about those people. Governments have to. And human persons building personal websites should too.
Wonder if you could respond in a way to get them to crash, or even better, to hang indefinitely.
[1] https://blog.haschek.at/2017/how-to-defend-your-website-with...
Non-bots break as well. I have Firefox configured to use HTTP/1.1 only.
No reason to chase Google's standard-of-the-day, HTTP/1.1 has worked for ages and it will continue to do so for the foreseeable future.
Instead, you are often dumping 4.8+4.8+4.8 as you add block methods, with some overlap.
I'm trying to figure out what I would gain by configuring my browser to use HTTP/1.1 only.
Because it adds nothing to improve my browsing experience, and reducing the number of protocols supported by my browser from 3 to 1 also reduces the attack surface.
> Your web browsing must be awfully slow sans multiplexing.
And yet it's not slowed down at all. How many different resources must a web page use before it feels slow on a connection pool of keep-alive TCP sockets? Maybe people visit some wild experimental web pages with hundreds of blocking <src> tags that are not bundled/minified?
Either way, my experience is it doesn't slow anything down when I use both websites (forums, resources, youtube, social media) and web apps (banking, maps, food delivery etc).
HTTP2 is barely any better than http 1, if you want it to make a 1/10 of a second difference you have to be making 100s of requests.
> The only ones that can survive the robot apocalypse is large web services. Your reddits, and facebooks, and twitters, and SaaS-comment fields, and discords. They have the economies of scale to develop viable countermeasures, to hire teams of people to work on the problem full time and maybe at least keep up with the ever evolving bots.
This is not true at all. There are web forums that are not "web-scale" and don't spend all day fighting bot spam. The solution is real simple: it costs 10 bux to register an account, if you're a nuisance your account is banned and you pay 10bux to get back on.
Even the sites that don't require payment for explicit registration - often succeed by gating functionality or content behind paywalls. Requiring a "premium membership" to post in the classifieds forum is an extremely extremely common thing on small interest-based web-boards (photrio, pentaxforums, homebrewtalk, etc). That income supports the site and supports the anti-bot efforts as a whole. The customer isn't advertisers - it's the community itself, and you're providing the service of high-quality content and access to people with similar interests.
You need to bootstrap a community first, of course, but it doesn't need to be a large community, just a high-value one.
The twitters and facebooks of the world just don't like that solution because they value growth above all other considerations. They'd rather be kings of a billion user website with 200 million bots than a 1k-100k user forum with 100% organic membership and content. And they value engagement over content quality, which is the entire reason comment-tree/vote-based systems have been pushed heavily over web-1.0 threaded forum discussions as well.
This botpocalypse is the inevitable outcome of the systems that social-media giants have created, not inherent outcomes of the internet as a whole.
That doesn't work at all unless your service is already pretty popular. Who would pay $5 to access a new, empty forum?
You mention "you need to bootstrap a community first," but that's basically an admission that this solution doesn't solve the problem at all, because you have to solve the problem in some other way to use this solution. 10bux was a solution limited to a very specific time.
I feel like the article here talks more about the first kind. I do agree with your solution for the second kind of spam though.
I've never seen this referred as "spam". Denial of service, botting, scraping, sure, but does anyone call that spam?
Uhhmmm, I beg to differ and so do a lot of very smart people with many more servers and users than you or I are likely to see.
As with most 'Oh, its' Simple - Just Do XYZ' solutions there are often very good reasons for not doing the 'Easy/Simple/One-Liner' and here are a few with yours -
Firstly - The '10 bux' could exclude a vast swathe of the poorest. Skipping a couple of Starbuck coffees vs. the local currency equivalent of whatever you are charging equating to a month's worth of food or being able to send at least one of your children to the local village school. I mean - your forum / site so you can gate it anyway you wish, I'm just pointing out that it could and would be exclusionary (perhaps unintentionally so).
Next Problem: Accepting and Processing the 'Good Behavior' deposit. Congratulations, you now need to become a Payment Processor and as such have certain legal requirements regarding payment details and storage and also tax returns. 'Oh, just Off-Load it to Stripe' someone might suggest. Do-able I guess but anyone who has taken payments over the internet will tell you that its a Royal Pain in The Ass. Also, now all a 'Griefer' needs to do is run a few dodgy cards through your registration system and 'Poof' there goes your payment processor and/or the fees go sky high.
Most ‘oh its simple – why don’t they just…’ overlook (or are not aware) of the many, many good reasons why greater minds than yours or mine haven’t already implemented it.
Sure – sometimes people do come up with novel solutions to old problems so theres no harm in spit-balling and I’m certainly not directing any scorn or ill-intent in my reply.
It is intentionally exclusionary. Not necessarily of the poorest among us, but of those who expect free service. Botters and spammers are disproportionately likely to look for free service. Pretty much any level of required spending in any currency will have a similar effect. By cutting off the abuse-prone free tier that many bad actors depend on, you dramatically decrease your exposure to abuse.
The point is not to keep out the poor people. The point is to make it far more work to get over the hurdle than it's worth for abusers. If you have a way to do the latter without the former that doesn't hinge on pushing a bunch of extra work onto administrators, I suspect quite a lot of people would be very curious to hear about it.
The forums they are referring to have been operating with the "10 bux" model implemented on top of a highly customized version of vBulletin for over 20 years.
Many years ago there was a public server called SDF (Super Dimensional Fortress). It was a BSD system and anyone could get a user account for $1. The theory was even the least of us, a kid scrounging for money on the street, could come up with a dollar (and presumably the postage to mail it). To a certain person, access to this kind of server was invaluable - the only situation you could hope to get close to this kind of system. As time went on, the number of people interested in this was dwindling.
Jumping through hoops is a useful gateway, but if your hoops are too complex or arduous, you miss out on people who you genuinely want to include in your community.
SDF is still around though, and still operates on the same model - pay once to get in, and you can stay as long as you want unless you become a nuisance.
> Jumping through hoops is a useful gateway, but if your hoops are too complex or arduous, you miss out on people who you genuinely want to include in your community.
It is certainly not impossible for communities with this model to die - that's not what I'm saying at all. Small social media sites die all the time, including with Reddit-style gamification bullshit. Or they turn into cesspits like Digg or Voat.
But yes, increasing the friction of engagement is literally the point, you are losing some users but increasing the quality of the ones who remain. It's the old "fire your bad customers" routine, but for social media.
"Oh no, we are all losing out on your valuable shitposting, how will this community ever go on?"
I don't think so...
First and foremost, SDF is well alive and there is a constant stream of people registering on it...
That's mostly because access to computers has became easier - you can either get your own Linux box or get a proper VPS for extremely cheap (if not free - see cloud provider free tiers) nowadays so why bother with a non-root account on a BSD system?
IMO it doesn't have anything to do with the barrier to entry.
The highly-chaotic multithreaded model of Reddit/HN/etc is directly designed to be impenetrable and chaotic, where everyone is just responding to everyone rather than having a "flow of conversation" in which everyone is involved. The "everyone responding to everyone" is literally engagement, and that's what those sites/companies want to drive, not community-building. It's designed to suck as a medium for serious discourse, because making 27 slight variations on the same response to 27 different comments keeps you on the site.
As long as we persist in having engagement be the primary metric, that's what you will get, and as long as we persist in the idea that the objective of social media needs to be making a couple people into billionaires, engagement is going to be the focus.
And again, it's a fundamental shift in "who the customers are". Are the customers the people using the site, who want a great place to discuss the nuances of parrot taxonomy, or are your customers the advertisers? Those lead to different ways you build the community.
And you can still make a six-figure income being a webmaster of a smaller community too. You're just not going to make Reddit money off Pentaxforums.
I couldn’t disagree more with this characterisation. Part of the reason that sites like Reddit and HN are preferred to traditional fora (which have their own engagement mechanisms) is because it's possible to have a different set of discussions on a topic.
Single-threaded fora result in many conversations on various sub-topics and of varying quality being multiplexed in a single chaotic and incomprehensible comment chain. Comment trees allow sub-topics and sub-conversions to be grouped in a reasonable fashion, and voting allows junk contributions to be pushed to the bottom.
It's not just webscale entrepreneurs - users love comment trees. It's part of why sites like Reddit and HN succeeded in gaining traction, and why subreddits are now the de facto replacement for fora (unfortunately, as it centralises editorial power).
What are examples of websites that maintain a flow of conversation while involving everyone without letting everyone reply to each other? Being able to reply directly to others while discussing a topic has been the norm at forums, and before those, mailing lists and BBSs. The clients/interfaces just got better over time at sorting/collapsing replies.
Whose got the forum where "everyone responding to everyone" doesn't happen?
While they certainly do value engagement over quality, I suspect that the systems are put in place because they don't scale in terms of manpower, and they don't trust their users to do formalize the structure of the site.
There is a big problem with all sorts of activists though. Their modus operandi is finding forums they don't like and the go on post illegal material and reporting it to hosting provider. Many forums stopped accepting new users because of that and for instance only way to sign up is to find the owner and speak directly to them.
You've highlighted its biggest tradeoff which is that it creates an economic incentive to ban people. The only way to make more money, is to have more rules and culture for ostracizing people. It would have been smarter of Something Awful (since that's the site we're talking about) to charge $4/month or something.
The innovative thing off that model is driving the revenue off the misbehavers instead of good citizens. You don't want to have the shitters around even if they are paying $4/month, and you don't want to drive off good-faith users even if they're mediocre/hapless/etc. So run the site off the backs of the people you don't want to have around.
People don't like paying monthly (this is even true of, say, app store revenue today) and if you apply recurring charges then when people don't think they're getting enough value they'll leave. You have a hard enough time on the user-acquisition side, why make it worse on the retention side by driving away the users who you're actually trying to keep?
Billing good citizens works in some situations where you have some specific value that you provide to them - providing sales listings on classifieds boards inside interest-specific forums is a good example, since you are providing access to interested buyers, which is a value-add, same as ebay taking their fee - but just in terms of operating a forum, you aren't a big enough value-add that people are going to pay Netflix-level subscriptions to the Parrot-Ass Discussion Club. You need the users more than they need you at that point. But a one-time fee is viewed much differently by people. People will pay $5 for an app, they aren't going to pay you $5 a month for it though, or at least far fewer.
Does anyone know what's going on with that "Duke de Montosier" spam botnet? It accounts for more than half of the botspam attacks on my sites, and I can't find anyone talking about it online anywhere, except one tweet dating back to mid-2021. It's identifiable by several short phrases that it posts:
Duke de Montosier
for Countess Louise of Savoy
Testaru. Best known
And cryptic short posts that can assemble into creepy sequences:
Europe, and in Ancient Russia
Century to a kind of destruction:
Western Europe also formed
and was erased, and on cleaned
only a few survived
number of surviving European
55 thousand Greek, 30 thousand Armenian
Many of the IPs involved seemed to be in Russia, China and Hong Kong, though they're coming from all over (eg European & US VPNs, Tor exit nodes). From tracking the IPs on AbuseIPDB, the weird spam posts seem to be just one layer, while behind the scenes it also attempts SMTP Auth and IMAP attacks on the server.
I'm eager to know more if anyone knows, and especially if anyone is trying to shut this thing down. But I can't find anyone even talking about it. (Maybe there's a reason for that?)
I've found spam email that certainly is the above, because it has things like PUT_LINK_TO_STORE_HERE and other variables that obviously weren't updated in the config file.
I've seen it suggested that botnets use comment fields for command and control, maybe something like that?
Here's the one tweet I found in Swedish about the comment spam botnet, and it dates back to February 2021. She's the only person I could find who has mentioned it in public. Or maybe my search skills are failing me.
First, of course, you have cloudflare and recaptcha, which are free and very efficient, as the author say.
But even if you don't want to use them (some of my services don't), most bots are very dumb:
- require JS, and you lose half of the web ones
- silly tricks like hidden input fields in forms that worked in 2000 still work in 2022. Use a bunch of them, and you can yet again halve the bot traffic.
- many URL should have impossible to guess paths. E.G: just changing the /admin/ url to a uuid in django or the /wp-admin/ in wordpress, you save so many requests.
- bots are usually not tailored to your site, meaning if you require JS, you can actually embed anti-bot measure in the client code and they will work. E.G: exponential backoff + some heavy calculations if too many fast consecutive ajax requests.
- fail2ban + a few iptables rules (mitigate syn flood, etc) will help
- varnish + redis gets you very far to shave excess dummy traffic
It's not great, but it's not an apocalypse.
Unless you are under targeted attack.
Then it sucks and you die.
That way well-behaved search engines won’t be affected, but naive scrapers get auto-banned.
Recaptcha has been almost useless, in my experience. If you read the spam logs, you'll quickly learn about the spam software they (claim to) use to bypass Recaptcha, because that's what they end up promoting. I started tagging in logs if Recaptcha had validated on messges, and sure enough these spam posts had all successfully passed it. Great opportunity to rip out more Google dependencies from my website.
I've found my own custom written filters to be vastly more effective than Recaptcha.
Lots of the bots are running full Chrome with JS, lots of HeadlessChrome being used lately. The fact that they're using HeadlessChrome is something that makes them easy to detect, ahem.
Recaptcha was also filtering out some legit humans (I logged all posts regardless of captcha status to be reviewed later), so it just wasn't worth reducing the user experience when the captcha bot detection rate was so low.
I would add to that:
- block signups/comments from known throwaway email domains
- block known datacenter IP ranges, at least for POST requests. Honestly on our sites 50% of spam was coming from AWS EC2 IPs
- use a proxy/vpn/bot detection service like https://focsec.com
The 0.001% was that person using EC2 as a proxy or VPN server.
Guess the fraudsters were vendor locked lol.
Throwaway email services have caught up to the point where even bigger players don't manage to block them, i.e. they rotate the domains used for the emails.
because then the attack becomes DOS as they cycle through dictionaries of words?
TL;DR: I won't expose any of my projects or API directly these days due to spam, ddos and other abuse.
This is unfortunate, because they're amazing feedback if you write about bureaucracy. People won't take the time to write to you about their experience, but they'll leave a comment.
Have you considered an info box where the comments would be to push anyone interested to mail? Or allow comments but make them work more like mail with comments not being publicly visible until you approve them?
If you just want to cut down on non-targeted spam then it might be enough to have your comments work even slightly differently from the standard word-press solution, e.g. by adding a simple text field with a question that real human visitors of your site should be able to answer. This alone has worked for all of my sites that allow user contributions so far - won't stop anyone trying to spam your site specifically ofc but you can always take further measures WHEN that becomes a problem.
PS: Is your site supposed to have a default WordPress favicon?
Not a bad idea, thanks. Maybe I'll try that. I suspect the "problem" is I just have approx. zero readers :)
> If you just want to cut down on non-targeted spam then it might be enough to have your comments work even slightly differently from the standard word-press solution, e.g. by adding a simple text field with a question that real human visitors of your site should be able to answer.
Yeah, I'd like to do that, but I don't have the time/interest in figuring out how to make Wordpress do that. I tried googling for plugins or whatever a while back, but never found anything pre-made.
> Is your site supposed to have a default WordPress favicon?
Never thought of changing it!
Im thinking of using GitHub issues/discussions as a comment system. The website and everything will function normally without CloudFlare, but the comments are based on GitHub, which will deal with spam and hosting for me.
And i personally think using it is better from centralizing the internet view, as you don't increase the absolutely crazy 20% that CF controls of the web.
But that obviously doesn't solve the DDOS problem, which should be solved with big cloud providers which have included DDOS protection.
Another workaround is using IPFS, but for a normal user, he/she will need a gateway, and guess who operates on of the biggest IPFS gateways ?, Yes CF. That without considering the tradeoffs is using IPFS.
I think a static site + external comment provider like Github, might help with the attacks and spam without using Cloudflare, but I don't have any website close to your blogs size, so its all just a predication.
I do generally think federated social networking (ActivityPub) is a good way to handle commenting. Comments are associated with identities separate from your site and there's plenty of opportunity to layer a reputation system on top of that, though I don't know if it's been done yet.
If someone was unable to get past that captcha (it still happens I have logs!) I figured they were probably not that valuable a contributor anyway.
If someone wanted to target my site directly they could but hasn't happened so far.
In the error message I have a friendly texts for human to turn on javascript if it is off and a <a href="javascript: history.go(-1);">Go back and try again</a> making them not loose the text that they have typed.
So, some creative people set up bots which check periodically for them. They are paid services which will do that for you. Now we have bots hammering gatekeeper's website. Perhaps hundreds of bots.
Which results in the website is being unavailable, serving a serious qps to bots. I think it is only a matter of time before someone will write bots that will apply to application bots hoping that more entries with their information will provide them with better probability of success.
This is so dystopian and cruel to the average person, and I don't think there is a good solution besides a deep anti-bot expertise whithin the primary website development team
But there is a solution: the website team should get their act together and remove the "first come first served" aspect altogether.
Do you, citizen, want to register? Cool - leave your e-mail and we'll call you. Is the service optional? Then we'll pick at random from the pool of applicants and e-mail them. Is the service mandatory? Then sign up and we'll call you once you reach the top of the queue. Add a quick ID/credit card/whatever check on top (like good concert venues do), regular e-mail updates to let people know they haven't been forgotten, and you're done.
Any second year CS student could write such a system. The difficult part is accepting that the current approach doesn't work and looking for alternatives.
Half the point of these services tends to be giving users some choice in when they have to show up somewhere. Because not everyone can make time in the middle of business hours of an arbitrary day. Not to mention that you might simply be out of town.
This is where having your own email domain with unlimited accounts is useful.
If you don't do that, I 100% agree with you - scalpers could then register with hundreds of accounts for reselling, and we would be back where we started.
Wait a minute, wouldn't that be more work for "us"?
Let me think ....
Luckily, if you're technically inclined, editing the value of the input element via dev tools is accepted by the form.
The tax agency had grayed out the "calc tax" button until the declaration period started, but you could just enable it with the enable flag.
Did nothing but read querry their servers though.
At least you didn't need to worry about it all day: there wasn't a chance in hell to get something 3 minutes after the booking system opened each day.
Friends described how stressful it was for them, their parents being completely helpless and essentially fearing that their health depended on getting one of those elusive appointments, and being devastated each day they didn't succeed.
This would be entirely practical for some small-time operator trying to run a forum off residential broadband, while impractical for the reddits, facebooks and twitters.
However for a small scale thing I'd gladly go visit at a face to face meetup to fulfill this type of validation.
Even if it were 3 flights totalling 18 hours away? :)
Or even just from one coast of the US to another...
However there is a medium / large organization case, where each area has local 'chapters' or some other term for a small fragment of the larger group. In that case the local leaders each operate as a small group for their areas.
Does this extend to societies as well? One can think of a membrane that has selective permeability to ideas but resists antisocial actors and concepts. Alexander Bard has talked a lot about social membranics (it's a bit hard to search for).
As odious as the web3 charlatanry is, I'm starting to yearn anything that raises the transaction costs for the dumbest bots. I remember reading something about new ideas with distributed moderation at some point--maybe someone can refresh my memory.
Although, surprisingly, I host such a service, and while it gets a constant stream of random hits, they're a minor nuisance. Probably because it's just the back end for a web page, and nobody bothers to target it specifically. Random web browsing won't find it, and the API will just return an error if called incorrectly. Even if it is called correctly, it has fair queuing on the service, so hammering on it from a small number of IP addresses won't do much.
That did happen once. Someone from a university was making requests at a high rate and not even reading the results. I noticed after a month, and wrote to their department chair, which stopped the problem.
Which other broadly interesting services do exist? The owners of those services could come together and offer a VPN that gets preferred treatment for these services. This could be more precise than https://www.abuseipdb.com/.
Rate limit everything you can and use a captcha where acceptable, there are also a load of public IP and email blacklists that you can use to run a quick check. Working in a field where there is a large amount of bots and incentive to abuse we invest quite a bit of time and money in fraudulent traffic detection using a cornocopia of different services in tangent and at the end of the day we still see a small percentage of traffic getting through that is fantastically human like.
With that out of the way, I've been engulfed in AI and GPT3 functionality lately and I thought this post was going to be doomsaying the coming apocalypse of bot spam, because the level of human like quality coming from the AI is going (already has) to make deciphering human vs bot traffic/posts/emails/comments nearly impossible. It will be fun here soon when we see forums entirely dedicated to bots conversing and arguing with each other outside of reddit.
Interactivity is not a must-have. The world's first general purpose computer, ENIAC, was not built for "interactivity". It was built to calculate ballistic trajectories, which were otherwise calculated manually. Computers exist to allow automation, to reduce manual labour.^1 "Tech" companies need interactivity to support collection of data about www users and paid services related to programmatic online advertising. Generally, users do not need interactivity. Generally, users do not need to spend excessive quantities of time "interacting" with networked computers.
As a user, I want _non-interactive_ www services, whether it is data/information retrieval or e-commerce. I want to use more automation, not less. Automation is not reserved for those providing "services". It also should be available to those using them.
Provide bulk data access. Let others mirror it. Take advantage of "open datasets" hosting if necessary. For example, Common Crawl is hosted for free with Amazon. Upload the data to internet Archive.
"The API gateway is another stab at this, you get to choose from either a public API with a common rate limit, or revealing your identity with an API key (and sacrificing anonymity)."
Publish the rate limit for the public API. Do not make users guess. Do not require "sign-in" to use an API to retrieve public data.
1. Some folks consider having to "interact" with a computer as labour, not fun.
Yes !
I call this software literacy. And yes - no matter how cool the JS on a major site, the fact that the sites goals are to keep me there and clicking and my goals are to get what I want with minimal action are in conflict.
I would suggest that bots are actually not a problem. For most things I would like a bot acting for me. Telling me as and when that I need to visit the dentist, who has slots free next weds and friday. Friday is best because I am also WFH that day.
The bot apocalypse is only one because we are trying to make a "web for humans" when actually a "web for bots, and a bot for a human" is a much better idea :/)
Isn't that simply your calendar? Sure, you want it automated; but it doesn't need internet access, it doesn't need to crawl or search, I don't know why you refer to it as a 'bot'.
To my mind, the idea of personal 'bots' was that you could give it some general instructions such as "Let me know when the content at any of these URLs changes", and then leave it running. Were they also called agents?
I run a website about immigration. I'd love to reinstate comments and get valuable feedback from people who just tried my advice. Bots just make it too time-consuming.
Of course slowing down DOS attacks is a great goal in itself, and it's very often what captchas have been (ab)used for, but it doesn't seem to me to replace all or most use cases for a captcha. In particular, since it can be completed by an automated system at least as easily as by a human, it doesn't seem like it would limit spambot signups or spambot comment or contact form submissions in any meaningful way.
Or am I misunderstanding, @realaravinth?
I used "captcha" to simplify mCaptcha's application, calling it a captcha is much simpler to say than calling it a PoW-powered rate limiter :D
That said, yes it doesn't do spambot form-abuse detection. Bypassing captchas like hCaptcha and reCAPTCHA with computer vision is difficult but its is stupid easy to do it with services offered by CAPTCHA farms(employ humans to solve captchas; available via API calls), which are sometimes cheaper than what reCAPTCHA charges.
So IMHO, reCAPTCHA and hCaptcha are only making it difficult for visitors to access web services without hurting bots/spammers in any reasonable way.
I'm the author of mCaptcha, I'd be happy to answer any questions that people might have :)
Also, here: https://mcaptcha.org/, under the "Defend like Castles" section, I think you meant "expensive", not "experience".
Keep up the good work!
I wonder why this approach hasn't been widely adopted?
Allowing abusers to freely abuse would cost even more power than just forcing them to do the work.
disclosure: I'm the author of mCaptcha
In addition, using sha256 for this is IMHO a mistake, calling for ASIC abuse.
Glad you asked! This is theoretically possible, but the adversary will have to be highly motivated with considerable resources to choke mCaptcha.
For instance, to generate Proof of Work(PoW), the client will have to generate 50k hashes(can be configured for higher difficulty) whereas the mCaptcha server will only have to generate 1 hash to validate the PoW. So a really powerful adversary can overwhelm mCaptcha, but at that point there's very little any service can do :D
> In addition, using sha256 for this is IMHO a mistake, calling for ASIC abuse.
Good point! Codeberg raised the same issue before they decided to try mCaptcha. There are protections against ASIC abuse: each captcha challenge has a lifetime and also, variable difficulty scaling implemented which increases difficulty when abuse is detected.
That said, the project is in alpha, I'm willing to wait and see if ASIC abuse is prevalent before moving to more resource-intensive hashing algorithms like Scrypt. Any algorithm that we choose will also impact legitimate visitors so it'll have to be done with care. :)
A big part of what makes our captcha successful in fighting abuse is that we scale the difficulty of the proof-of-work puzzle based on the user’s previous behavior and other signals (e.g. minus points if their IP address is a known datacenter IP).
The nice thing about a scaling PoW setup is that it’s not all-or-nothing unlike other captcha’s. Most captcha’s can be solved by “most” humans, but that means that there is still some subset of all humans that you are excluding. In our case if we do get it wrong and wrongly think the user is a bot, the user may have to solve a puzzle for a while, but after that they are accepted nonetheless.
There is a blessed source-available version of the server that you can self-host [0]. It is more limited in its protection, but it is probably good enough for hobby projects.
[0]: https://github.com/FriendlyCaptcha/friendly-lite-server
Newer variations (such as argon2) are tunable so you can include memory footprint and cpu-parallelism. There also are time-lock puzzles or verifiable delay functions that negate any parallelism because there's a single answer which can't be arrived at sooner by throwing more cores at the problem.
Beware when choosing a CAPTCHA that serving "most humans" might exclude those with accessibility issues, like the visually impaired.
We need to redesign a web based on APIs, certificates, rate limits etc. And stop having "engagement" as a goal, and have "getting things done" as a goal
Edit: mucked up formatting
Yet, I'm a very good web citizen.
Because of this, I often have to solve the same captcha many times before it thinks I'm human.
So the second browser is the solution. But then the site will do all the bad things that I wanted it not to do in the first place. Like serving terrible French results instead of good English ones, or assuming Firefox doesn't work based on UA while their site work fine with it. And of course track me to death, sell my data, and so on.
The only solution that works is to chose services you pay money for: they have your card, so they know you are not a bot. For years now, I have been suspicious of anything free. But it doesn't solve the tracking problem.
That's going to be way worse for humanity than spiders and automated scripts sending too much traffic. This article isn't imagining apocalypse creatively enough.
Though we can all work against that by securing our own systems and preventing them from being abused. Used or unused domains should have a strict SPF policy, website registration (or newsletter signup forms) should have captchas, comments should have captchas. Wordpress or other CMS's plugins should be up-to-date and so on and on. Work on requiring 3DS everywhere, everything in-depth.
That way malicious actors would be limited to the services they pay for and that makes their life significantly harder.
> The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse.
I understand the drawback here but I would like to see monetized transactions employed as a defense layer a little more before we make a final decision. It is undemocratic, to be sure, but maybe for those of us who can afford it it's still better than the cesspool we currently sift through on every major platform. Anyone aware of any platforms taking this approach?
Maybe the fediverse will help - by fragmenting networks attackers may have less incentive to attack a particular one.
The Postal Service?
Sure, there's junk mail, but imagine how much junk mail there would be if it were delivered for free. It wasn't until phone calls became so cheap as to be "unlimited" that we ended up flooded with billions of junk calls.
Microtransactions (non-crypto, thankyouverymuch) would solve a certain number of today's problems.
You get either that, or yucky defi web3 crap.
In the interest of practicality: There's a way to go the web3 route without being laden with transactions:
- Mint a fixed-cost non-transferrable NFT to an address, with ownership limit of 1 per address. - Use SIWE (sign-in with Ethereum) to verify ownership of address & therefore NFT. - If malicious behaviour is detected, mark the NFT as belonging to a malicious actor at the server's end & block the account. - Require non-malicious-marked NFTs in order to use the site/app.
At most, the user only had to perform 1 transaction (minting the non-transferrable NFT) on any blockchain network where the contract resides, & the costs to do so can be made cheaply with Layer 2 networks. (Polygon PoS, Arbtirum, Optimism, zkSync 2.0, etc)
Can this be done entirely without web3? Yes, but the added friction imposed onto malicious actors to generate new addresses & mint new non-transferrable NFTs increases the costs for them considerably.
> If anyone could go ahead and find a solution to this mess, that would be great, because it's absolutely suffocating the internet, and it's painful to think about all the wonderful little projects that get cancelled or abandoned when faced with the reality of having to deal with such an egregiously hostile digital ecosystem.
In all honesty, there's no perfect solution, just hard-to-make tradeoffs: The prevention of botspam inherently requires tracking in some form to resolve said issue, as there's no immediately-recognizable stateless solution for botspam tracking. Someone has to do the tracking to prevent botspam, which inherently involves in state being changed in order to mark an actor as malicious.
Not commenting on op's solution's validity or effectiveness, just replying to your comment.
Do you mint one NFT per address? If blocking only applies to the one site, a malicious operator can just spam the next site using that address - after all, they own tons of addresses and have many sites to spam, and they can surely spam for at least a bit before getting caught (per site and per address).
If blocking is a public operation that gets you banned everywhere, well now one callous server owner can now disable your address’s ability to access anything.
I fail to see how Web3 NFTs solve any of these problems…
The only way out is to increase the costs for bots to such an extent that it becomes costly for them to operate. The methodology for performing this is still up for debate, but tracking of some form will have to exist (be it publicly-collated or privately-monitored or some blend of both).
In the real world that's called a paywall. Just requires users have an email address (wallet address) and proof of payment (NFT). Web3 contributes nothing but a different lingo for the existing concepts, yet has all of the same problems: nobody likes paywalls.
We also use some simple heuristics to reject obvious bot traffic.
One of the simplest is to have a form field that is hidden via CSS. Humans don't see it and it stays blank. Bots fill it in.
Bots tend to fill in every form fields with random garbage, even checkboxes. Validating checkboxes rather than checking they have a value is another good way to detect bots.
Many bots have a hard time with CSRF tokens in hidden fields.
Many bots also don't handle session cookies properly. If someone submits a registration form without an existing session, we reject it. (So we don't get as far as checking the CSRF token.)
After a certain number of failed attempts to register or login, we block the IP for a period of time.
My best guess is they're assuming it's backed by google, and are attempting to poison its search term suggestions. The queries I've been getting are fairly long and highly specific, often within e-pharma or online casino or similarly sketchy areas.
Many of these bots are exceptionally simple - zero tailoring, just scriptkiddie copy/past/run stuff.
They could simply iterate through a list of websites inputted by the user. Said list could be some curated list you find on blackhat forums, which is just a collection of websites with traffic over a certain threshold. Doesn't say anything about the content or what type of website, just the URL.
Then it does a simple operation to map the website - as well as checking for input forms. If an input form is detected, it does a POST operation. If there's an error, timeout, or whatever, it simply moves to the next one.
This of course sucks if you try to run a search engine - because even the simplest of bots will succeed when your website is literally just one website, with the text/input field right there in the front, and with no captcha or similar to dissuade legit users.
If the queries are not a megabit each, you're doing way too much processing before applying rate limiting. Rejecting traffic ought not to take more than 1-2 milliseconds, even if you need to look up an api key or IP address in the database.
I, too, host services on a residential connection: 50 mbps shared with other users. My domains must host hundreds of separate scripts, a few of which have a database attached (I can think of six off the top of my head, but there's over a hundred databases in mariadb so I'm sure there's more that I've forgotten about). This is a ten-year-old laptop with a regular "apt install mariadb", no special configs.
Yes, most traffic is bots, and yes sometimes they submit more than 1 q/s. But it comes nowhere near to exhausting resources to a noticeable extent. Load average is about 0.15, main peaks come from my own cronjobs. If you're having this much trouble rejecting traffic, you might want to spend some time looking at bottlenecks. You'll also notice the bots knock it off if it's unsuccessful.
You know how blogs sometimes can't handle being on the hacker news front page from all the traffic? Well my search engine has survived that. That was 1-2 QPS.
Unmitigated bot-traffic is roughly 10x the traffic of a hacker news death hug, as a sustained load.
Most of those has been either “worse google” or “utter trash”… this one returns some interesting results for some queries I have tried.
15 RPS is very far from an apocalypse.
I traced it once, and I got to admit there was not an obvious bottleneck (this was 2015 or so). Just millions upon millions of calls into deeper and deeper layers for things like translations or themes. Wrapping mysql_query in a function that caches the result (to avoid doing identical queries) helped a few % I think, but aside from major changes like patching out the entire translation system for single-language sites, I didn't spot an obvious way to fix it. You'd need to spend a lot of time to optimize away the complexity that grew from suiting a million different needs, contributed by thousands of people across many years.
Netlify seem to not really care after reporting it on their support forum. The spammer disables JS so no client side protection works. I’ve recently decided to (unfortunately) break the ability for JS disabled browsers to be able to submit the contact form. The form elements attributes are all wrong meaning the form won’t submit correctly. Instead, some JS on page load sets the attributes to the correct values. I will wait a while and see if this solves it.
While netlify does correctly mark all this as spam the fact is that legitimate messages sometimes can slip past these with false positives. So I have to check the vast amount of spam often.
Cloudflare is not the only CDN/protection. It's the most popular and the most evil one. You have a choice.
It is actually nuts just how much bot spam there is.
Now the author above has stated they dislike the crypto route and I agree that the whole web3 idea is bs but what if in the case that spam of some form is detected by the server, it requires the visitor to show some proof-of-work and combine that with the "mining crypto in JS instead of ads" craze. That way the bot would need to put work in which would slow it down and at the same time it would pay for its own visit.
No ofc no spam detection system is perfect and it would also hit human users but in their case it would be just a wait a few more seconds longer for page to load kinda case.
Payment per request is the long term solution, but I completely disagree that it’s in any way dystopian. The trick is to set the fee so low that humans, who make few requests, aren’t really affected while bots, who make a large number of requests, become unprofitable.
It’s exactly the same solution as email spam: at a hundredth of a USD cent per email, spam emails would no longer be profitable while regular consumers would spend 10 cents per year (assuming they send 3 emails per day).
The number of languages supported would be small to match whoever helped moderate this, but it would at least require speaking to someone. A PM thread or live chat would be an instant way to find out if someone can string two sentences together and might be worth allowing into the site. You could even have them create an account and solve a single captcha prior to getting access to the chat.
It's not perfect by any stretch, but might be worth exploring having humans-verify-humans.
I only agree when it comes to the system resources that can keep up with bots. When it comes to fighting spam, these services often do a terrible job because 1) their business model benefits from higher user & engagement numbers and 2) their monopoly status affords them to retain users even if their experience is degraded by the spam, something a small site often won't be able to do.
If one request to the site generates more revenue than it costs in resources, the bot problem is solved.
The author says that he is getting 15 bot requests to his site per second. That is about 36 million requests per month. How much does it cost to serve those? $1000 would seem high.
$1000/36M = $0.00003 per request.
How long would a crypto currency, that is suitable for mining in the browser, need to be mined before $0.00003 is generated?
If it turnes out it is a few seconds or so, the solution would be nicely user friendly. A few seconds of CPU time for access to the site. No ads needed to finance the site.
It is kind of telling, that Bitcoin started as a spam blocker. The original "hashcash" use case was to use proof of work to prevent email spam.
It might be enough for the request to require more resources from the requestor than from the server, even if it doesn't actually give the server any money. I mean the requestor probably isn't going to be willing to dedicate more hardware "horsepower" to taking your search engine down than you are to keeping it up. That was the idea behind Hashcash.
As for coins, the current Bitcoin hashrate is about 200 exahashes per second, down from a high of over 250 a couple of months ago, and the block reward is 6.25 BTC until probably June 02024. At a price of US$24000/BTC that's US$150k per block (plus a much smaller amount in transaction fees) or about US$1.25e-18 per hash. So your suggestion of US$3e-5 would require about 2e13 hashes. https://en.bitcoin.it/wiki/Non-specialized_hardware_comparis... says an overclocked ATI Radeon HD 6990 can do about 800 megahashes per second (8e8) so you're looking at about 3e4 seconds of compute on that card, about 8 hours.
Maybe one of the altcoins that uses a hash function with a smaller ASIC speedup would be a better fit, although I don't know enough about mining to know if there are any where GPUs are still competitive. Still, it seems like it might be more than a few seconds?
I did say "crypto currency, that is suitable for mining in the browser" for exactly this reason: That Bitcoin is not well suited for it.
One would have to look at what typical consumer hardware is good at. Maybe an algorithm that saturates one CPU core with serial calculations that need fast access to exactly 1GB of RAM. I think consumer hardware is pretty good when it comes to single core performance and RAM access.
EDIT: bad joke but maybe someone will get a chuckle.
Secure Scuttlebutt[1] doesn't have a lack of moderation / spam issue and it is completely decentralized and without monetary fees nor proof-of work. Why can't centralized services do better?
[1]: https://ssbc.github.io/scuttlebutt-protocol-guide/#follow-gr...
I’m hopeful that “decentralized identity” is a solution here.
In theory this would let people cryptographically prove whatever small fact about themselves they would like to share with a service provider like marginalia, without paying you, or sharing their identity, or anything else.
Eg: “I am a human”, or “I have written less that 100k words on the Internet” or “my comments have an upvote average above zero”
Trying to login to a linkedin account using a google account from an automated browser(like puppeteer+puppeteer-stealth or fakebrowser), will open a white empty window instead of the normal google login window, I could be a limitation of those libraries, but I doubt it, smells like something they detect, maybe looking into might lead some interesting insights on how to limit modern bots.
What type of queries are they generating? For what purpose are querying Marginalia? Scraping and filling internal search engines?
> If anyone could go ahead and find a solution to this mess
I would maybe trying to investigate why are querying your search engine. Is for the search results? Maybe from there you can create and sell an API service. Is for the wiki? Is for research purpose?
I would love to see some data, raw or with some behavior derived from it.
As I've mentioned in another comment, my best guess is they're betting it's backed by google, and are attempting to poison their search term suggestions. The queries I've been getting are fairly long and highly specific, often within e-pharma or online casino or similarly sketchy areas.
Like
> cialis 50mg online pharmacy canada price
Either that, or nonsense like the below, where they appear to be looking for CMSes to exploit (although I don't understand the appendage at the end)
> "Please enter the email address associated with your User account. Your username will be emailed to the email address on file." Finestre Antirumore Torino
> affordable local seo services "Din epostadress delas eller publiceras aldrig Obligatoriska flt r markerade med"
> "You are not logged in. (Login)" Country "City/Town" "Web page" erst
Point is, none of these queries actually return anything at all. I don't offer real full text search, for one. And the queries are much too long.
1. whitelist the finite IP ranges for the regional ISPs/country where you do business
2. blacklist the proxy and tor exit nodes
3. blacklist the list of published compromised servers
4. add spamhaus blacklists
5. add fail2ban rules to trip on common server security scans, and unused common service ports
6. publicly reply to those having access issues, and imply they have bad neighbors.
This will often take care of 99% of the nuisance traffic, but I still recommend live monitoring traffic regularly. ;)
A lot of (lucky) us have the luxury to live in real democracies.
Some others live in countries that use every single aspect of their private lives (DPI, mass surveillance) to put pressure on them and bend them to the regime's will.
In my opinion, Tor and anonymity should not be killed as a result of silly bots.
We have had exactly zero traffic from it at any point which was legitimate. Any user who ever showed up with a exit IP ended up being banned eventually, so we just proactively fraud banned anybody who uses one, and anybody that was related to them. There is zero value in allowing anonymizer traffic on your service, and a whole lot to lose.
That being said, a commercial site owes nothing to financially irrelevant bandits, sociopaths, or shills.
Try it for a week, and then weigh the liability again. ;)
>large resources are the solution
To those who have been recently pondering the history of antivirus companies of the 90s and 00s, and suspiciously wondering how they were always able to so quickly come up with definitions for the newest infections, this all feels so familiar. What a sad world we live in, sometimes.
captchas were designed to solve this (and only this, as opposed to requiring them to merely view content like modern ignorant web devs like to do [yes i know some web devs now require it to be able to make sure the people they're datamining are real, but this is a new practice from this year basically])
public services should be implemented by decentralized p2p. static content is solved by ipfs, freenet, etc. dynamic content perhaps can only be solved with smart contracts, which would be less bad than cloudflare if they weren't expensive, as they still provide protocol conformance (unlike cloudflare that requires you to have your packets look like a big 4 browser), anonymity (yeah, pseudonyms, you can still make one per query), etc. without smart contracts many interactive applications are still possible
> The other alternatives all suck to the extent of my knowledge, they're either prohibitively convoluted, or web3 cryptocurrency micro-transaction nonsense that while sure it would work, also monetizes every single interaction in a way that is more dystopian than the actual skull-crushing robot apocalypse.
centralized web hosting is and always was unsustainable and this is the reason most web content is commercial garbage, and the problem will only get worse. my concern was always what kind of garbage boomer protocol will become the new standard. i sure as hell dont want something that looks like email, web, or UN*X.
Has NLP progressed rendering Paul's plan a failure?
Am I a bot? How about you? Does it matter if I make valuable contributions?
In our experience (we don't have a forum), almost all of our bot traffic has been SEO spiders (or claiming to be so).
I don't really understand, is that a lot? 15qps does not sound like a lot, especially for a blocking/rejection function.
I’ve noticed that a lot of old popular forums disappeared in recent years, but I didn’t realize it was possibly due to bots. Why is that? I assumed that the admins just got tired of running them and moderating them.
What??? My phone can serve that easily, any modern server can handle 50-250 rps
All while we blame Russia and China, and say that their spambots and evil actions forced us to do this.
Another end game: some sort of ID is required to use anything (that ID being a local phone number, which can be tracked down to you).
For our app, we don't deal with spam in any novel way. We use honey pot, SFS, and Akismet.
However, by far the easiest way to stop spammers is a post queue. Lots of spammers will just create a burner account, fire off their spam, and start over. Given no actual reputation, give them the trust they deserve — none.
The other factor is building out a fast backend. Besides benefiting your own users, it also means Googlebot or Ahrefsbot won't absolutely cripple your site when they come knocking. Sometimes that is doable, sometimes not.
I think my backend is plenty fast given it's hosted on a PC off domestic broadband. Most searches complete sub-100ms.
Specific scenarios require creative solutions. For a search engine, how do you differentiate between robots and legitimate users? It seems a rate limiting step is likely the best solution.
If query rate from the same IP exceeds a threshold, throttle them creatively. +100ms the next time, +250ms the next, etc.
The upside is these bots will adjust their strategies to hit your site slower, which is the whole point, isn't it.
If they spread requests across IPs, perhaps try fingerprinting. I'm not sure how effective that is on the backend though.
1) The problem is centralization. Yes DNS is federated but there is a central registry. This means anyone can spam you@yourdomain.com or visit your web server listening for HTTP connections at www.domain.com
2) DNS is a glorified search engine. Human readable domain names are only needed for dictating a domain name (and listeners often make mistakes anyway). They only map to a small fraction of URLs) namely the ones with “/“ path name. For most others, the human readability adds little benefit.
3) Start using URIs that are not human readable. The titles, favicons and other metadata of resources should simply be cached, and displayed to the user in their own bookmarks, search engines or whatever. For Javascript environments, variables can easily hold non human readable URIs. Also QR codes can resolve to non human readable URIs.
4) There may be some cookie policy for third party hostnames etc. but just make them non human readable also.
5) We should have DHT or other decentralized systems for routing, and here is the key… in this system, you need a capability issued by the website / mailbox owner in order for your message to be routed to them. If the capability is compromised and used to get a ton of SPAM, they simply revoke that specific capability (key).
For HTTP websites you can already implement it on your side by signing the keys / capabilities ie session cookie balues with an HMAC, and there is no need to even do network I/O to verify them, you can upload the whitelist to the edges and check them there easily.
But going further, for new routing protocols, IP addresses should be removed after the first hop in the DHT, because the global routing system will send traffic there otherwise. See how SAFE network does it.
6) I don’t need a “real names policy” or “blue checkmark”. I can know who “The Real Bill Gates (TM)” is through some verified claims by Twitter or someone else. Just because I have the email billgates@microsoft.com doesnt mean I should be able to email him. There can be many Bill Gates. The names are just verified claims by some third party. Here on HN we dont have names or photos, and it works just fine.
7) Most of the celebrity culture, papparazzi, Elon Musk and Donald Trump moving markets and tweeting at 5am to 5 million people at once, are problems of centralization. Both a 1 to many megaphone and a many to 1 inbox. Citizens United is just a symptom of the problem. I have spoken about this (privately owning access to an audience) with Noam Chomsky in an interview I did a year ago:
https://community.qbix.com/t/freedom-of-speech-and-capitalis...
Fox News (Rupert Murdoch), CNN (Ted Turner), Twitter (Elon or Jack), Facebook (Zuck) are controlled by only a few people. Channels on youtube, telegram, podcasts etc are controlled by a few people. This leads to divisions in society, as outrage clickbait rises to the top. Nonprofit models based on collaboration like Wikipedia, Wikinews, Open Source and Science produce far more balanced and benign information for the public.
In short we need alternatives to celebrity culture, DNS and other systems that centralize decision making in the hands of a few, or create firehoses and megaphones. Neither the celebrity nor the public actually enjoy the results.
Yet Cloudflare still gets two paragraphs of complaints in the face? Because the author wants to “own” something instead of “renting”?