You probably don’t need ReCAPTCHA
kevv.net
kevv.net
I run 100s of small random low traffic low priority sites. Without some form of form control, the ALL get hit with customized and random other crap spam. I don't have decent experience with many things in life, but I can say this is one topic I have YEARS of experience with. I've never over-estimated the amount of any type of spam any form can get after being on the web for just a few months. Doesn't matter how big they are or what they do.
I'm not saying ReCAPTCHA is the only thing out there or even the best, but having an open form is just asking for trouble.
The company I work for makes a SaSS forum product, and while we do have multiple spam prevention methods (akismet, stopforumspam, honeypot, a hidden input), there’s enough stuff out there that has targeted our platform that a Recaptcha on the registration form is needed.
We haven’t need it on any other forms yet though. After registration it’s all handled by the other methods and various moderation tools.
I can and have defeated forms that tried to do all of those things very easily in the past.
Keep in mind that if you randomize across a few variations (i.e. 4-5 page layouts), that's easily discerned if you pull the page source down 20-30 times, doa complex diff, scrub out obviously random strings, and check the total unique variations you're seeing.
That may seem like a lot of work, but consider that if you don't do it all at once, but instead roll out small change after small change, the person or people using it are not weighing to cost to do everything required to bypass it compared to finding another open mail form, but the cost to bypass just the new fix you put in place. Also, they might think it's fun doing so...
And on the site dev's side, they can just choose to outsource it to a CAPTCHA (not that there aren't services to easily bypass CAPTCHAs at scale at sub-cent per CAPTCHA rates, see https://anti-captcha.com/).
Note: To forestall any assumptions, I wasn't doing any spamming or helping spamming in any way.
It's a "don't have to outrun the bear" situation, make yourself just difficult enough that some easier target gets snagged instead.
If everyone else is incorporating recaptcha, they're all running faster than you. Even with bypass services, cheap is not the same as free, especially at the scale spam runs at. I imagine a mail form that obviously doesn't incorporate a CAPTCHA is going to garner some attention. It might work for weeks or months if it's not being paid attention to, so that's probably worth them spending a few minutes looking at.
Spam doesn't scale on a small site. Say you can absolutely fill a small site with spam comments to the point that 99% of comments are spam. Very few people visit the site (it's small after all). Fewer still read the comments. Virtually none of those will click on the (usually obvious) spam links. And still fewer will buy, making you money. If you spend 2 hours customizing your spam script to circumvent anti-spam measures on a small site, you might as well flip burgers at McDonald's, you'll make significantly more money.
Spam works at scale only when you're not customizing. I'm involved with quite a few small to medium and a few larger sites (the largest getting around 4m PI/month) and though we use WP we get virtually no spam because of trivial deviations. We get an immense amount of attempts though. The little we do get is obviously manual spam: in the correct language, with content targeted to the individual page/post content (beyond "very interesting article, I wrote about the same" one-size-fits-all).
This is as primitive as it gets. I didn’t get a single spam mail in all that time.
The idea is not to outrun your competition, it is to become a special target that would demand special work to successfully get into. Bots are dumb as long as the humans behind them don’t give them a hint how to deal with your site.
And if you’re really that valuable of a target, you can step it up a notch or even switch to google’s data collecting solution.
I am much more confident in ReCAPTCHA of stopping bots compared to any roll your own solution.
I dont want to hope that an alternative is good enough for my needs. I want the best when it comes to protecting my site.
Any alternative needs to have a proven track record and support to make consider replacing ReCAPTCHA.
Do you want your site "protected" from those users, too?
For my personal blog I managed to be spam free with a simple question/answer pair for 5 years. Took me a minute to implement and leaves my user data where it belongs.
You can use common knowledge or simple ambiguity of language. You can use simple math arithmetic, written in properly obfuscated html. and randomly generated on each page load. You can use custom question about the content of the article (helps with informed answers).
On a small blog of mine just one question with one answer on the contact form prevented all spam for over 5 years already although it would be trivial to exploit in a targeted attack.
Targeted attacks are rare unless your captcha protects a juicy target that is worth a targeted attack at some point.
Are there alternatives in situations like this?
And I would even hazard a guess that the TOS specify that Google will not retain/link that information, considering that's how Analytics is run.
[0]: https://addons.mozilla.org/en-US/firefox/addon/umatrix/
I am as well. We enabled Recaptcha on one site and had spam signups drop by 99%. Unfortunately, regular signups also dropped by 20% because people give up when they hit Recaptcha and don't absolutely, seriously need what it's protecting. To us, joining the arms race against the spammers (which, so far, we've easily won) was much more profitable than turning away legitimate customers.
>>> You probably don’t need ReCAPTCHA
Probably being the keyword, because you probably aren't a big enough site for a dedicated attacker. Or for a dedicated attacker to be an issue.
And really, let's s/attacker/bot/g. Not every bot is a problem. Not every bot is an attacker, i.e. someone doing something malicious.
Can you please provide a few ready-to-use links?
You can pay for Azure and other STT engines to solve it for you an dthe results are usually a bit better.
I'm much more afraid of ReCaptcha blocking bonafide users. It's a harmful obstacle that punishes legitimate users for not sharing as much data as possible with Google.
Even if you really need a captcha, there are better solutions out there.
I recently spent time ensuring our Auth pages’ HTML could be easily cached outside of our application servers. They were a common target of DDOS attacks because we were generating a unique nonce for CSRF protection.
Randomizing form field names does not defeat a targeted attacker (and we have definitely been a target), prevents HTML caching, and will prevent auto filling fields by browsers and password managers.
Additionally it will be terrible from a usability and accessibility standpoint.
It’s trivial to target a form field by the text/label around it so those would need to be randomized as well.
I would MUCH prefer the recaptcha over this!
I wholly agree that this would not help, but for the sake of completeness, I want to point out that <input autocomplete=""> [0] is designed to solve this, by decoupling input field names from their intent.
But Chrome is playing dumb about it [1]. And of course, the spambots will just adapt to parse the autocomplete info…
[0]: https://developer.mozilla.org/en-US/docs/Web/HTML/Attributes...
[1]: https://www.reddit.com/r/programming/comments/ar1qj1/chromiu...
ReCaptcha is by definition terrible from a usability and accessibility standpoint too, just has all the privacy problems too.
The ball had an animal in it and I was asked to bounce the ball, causing it to rotate. I had to bounce the ball with just enough force to get it to land so the animal was positioned upright. After several failed attempts, I gave up.
On my own site I see 1 or 2 spam posts a week although I get the feeling it’s real people doing the registration. They sign up, make 1 comment, get reported very quickly, then banned.
We haven’t had to make our signup/registration system that strong in of itself though, because most of our largest clients end up using some SSO method exclusively and will have their own prevention methods.
`autocomplete="off"`
Huge pain
https://codesandbox.io/s/static-jkvzs
The other trick is to add a random string/number in from of the name attribute e.g. name="348349_name". This prevents autofill. Interestingly 1Password and LastPass are smart enough to infer that it's a name or email field.
For the honeypot, random number + word makes it ignored by autofill/1password
looked into it again and it seems Chrome enabled it again in Chrome 68: https://stackoverflow.com/questions/25823448/ng-form-and-aut...
Firefox had it disabled too but enabled it back again: https://developer.mozilla.org/en-US/docs/Web/Security/Securi...
And IE is just a cluster f.
My point being, Autocomplete off is not a valid solution as it can break at an updates notice, and the code hacks, while may work, are a pain to deal with
I'm sure this won't work for everyone, but if your small, I highly recommend giving it a go.
If you want to filter comments you could even make the questions reflect the content of the article, filtering uninformed TL;DR type of comments and giving the users the feeling you value onformed opinions.
It prevented 100% of bot spam for years. Granted, I was never a big enough target to make anyone rewrite their bot, but that's the same for most of us. I'd never use a captcha so long as something trivial like that works 100%.
https://www.phpbb.com/support/docs/en/3.2/kb/article/how-to-...
It works very well.
It's a small site, footfall in the 10s, so maybe that's the reason.
However, if you're working on anything with non-insignificant amounts of traffic, you'll get hit with some customized spam.
I've been dealing with these spammers, and if you do nothing, your forum will be filled with korean ads. We implement Akismet, StopForumSpam, Project Honeypot, and ReCAPTCHA, with the latter being the most effective (sadly). I'm pretty sure some of these spam agencies have customized tooling to handle NodeBB (they're using websockets to submit the posts, instead of HTTP POST).
Outside of these strategies, the most effective by far is reputation restrictions. Post queues if you're new or don't have enough upvotes, etc. However it does require manual effort, of course.
Would definitely love an alternative.
Did you tried some techniques from the article? Like hidden form fields, simple javascript checks or simple captcha?
Spam attempts would grow exponentially. So every time we cut it down by 90% via one of these tricks, it only gave us a bit of time.
None of this stops the determined troll though.. they can have all day to manually add offensive content. Shadow banning was good for this (1999) and group shadow banning was the best (bifurcated forum posts so all the banned people saw each other, but no one else did). Ah, memories. So good.
What did happen with that strategy?
In cases like these, someone loads up a huge botnet, a downloaded list of hacked usernames and passwords, and tries every single combination hoping to find a reused username/password combination.
In these cases, it is almost always extremely targeted. Log correlation has helped quite a bit, but it is still very painful since they alternate IPs with every request.
Automatically adding a blanket ReCAPTCHA on all login pages during a distributed brute force attempt is one of the few things that actually stops an attacker like this with minimal negative consequences.
I'm sure it is frustrating to users, but I think service disruption from what is effectively a DDoS is a worse user experience.
1. Rotate through several thousand to several hundred thousand noncontiguous, geographically distributed, residential IP addresses,
2. Associate each IP address with a single user agent and suite of cookies,
3. Associate each IP address with a particular target username,
4. Only attempt a few incorrect logins at a time, and a somewhat random (albeit realistic) number at that, within a given time interval,
5. Use random, apparently human delays between successive requests,
6. Issue requests using extremely high fidelity simulacra of web browsers, customized to the sequence and structure of HTTP requests on the website.
When the stakes are high this is the kind of opposition you'll get. Bank account takeover, social media account takeover, ticket scalping, automated sneaker buying, financial research, market research, etc.
Recaptcha introduces unpleasant user friction, but it usually works well. To invert a popular turn of phrase, it makes stopping simple attackers easy and hard attackers possible. The most sophisticated attackers will still lease reputable Google accounts and mechanical turk time to bypass Recaptcha challenges, but it will be expensive for them.
Technical sophistication is only one dimension of this game. The other is making adversaries spend more money than they can gain from being successful.
???
Please ELI5. I mean, why are sales bad, even if automated? Are they using stolen cards?
Back to the point at hand, I don’t like recaptcha in principle. But given my view from both sides of the table, it’s one of very few things that consistently works for sophisticated adversaries. It’s about as close to a silver bullet as they come, with the additional upside that it’s the absolute easiest thing to implement - in both an absolute sense and relative to the return. And once you have, most of what you can implement beyond recaptcha has diminishing returns in comparison.
All of that being said, I would be inclined to agree that most websites and apps don’t need recaptcha, simply because most of them aren’t worthwhile targets for the types of attacks recaptcha is singularly effective against.
The more successful quant funds will often build out internal research teams to do this. For example, both Two Sigma and Millennium have (not so well advertised) research teams devoted to this kind of data collection internally.
Or as ALittleLight says, auction them?
Edit: OK, I know, limited editions. Like numbered and signed prints. But it's arguable that people who want them the most will get them. Even if it's just for resale. Doesn't seem like the seller's responsibility.
sounds like to me that the seller doesn't want the scalper to sell outside the official channels imho. It might dilute the brand as well.
And why brand dilution? Scalpers sell at a premium, not discounted.
The real problem is supply. Popular tickets are scalped because there's only so many tickets. Then unpopular tickets are scalped because it was so easy to scalp the popular ones.
There's only so many sneakers that can be made: making more chews up the supply chain for something which isn't _truly_ being consumed.
https://www.digitalmusicnews.com/2018/11/02/post-malone-croc...
How are they getting residential IP addresses, compromised PCs?
> Monetize your mobile app or game with our SDK, without showing intrusive ads or requiring annoying subscriptions and in app purchases.
They approached nmap of all people:
> Hi,
> My name is Lior and I'd like to offer you a new way to make money off your software. The Luminati SDK provides your users the option to use your software for free by contributing to the Luminati proxy network.
> We will pay you $3,000 USD a month for every 100K daily active users.
> No collection of users' data, no disruption of user experience.
> I'd like to schedule a 15 minute call to let you know how we can start. Are you available tomorrow at 12:30pm your local time?
> Best regards,
> Lior
But yes, the whole cottage industry is sketchy. Almost all providers are leasing users’ computer with outright malware or shady TOS. The savvy play is to release a free game, app or even SDK which will then opportunistically route requests from the control server through the user’s device.
Recaptcha solving APIs are frequently bundled with the more reliable and premium services of this kind. They introduce a lot of latency since there’s a real mechanical turk across the world solving it for you, but they basically work.
You can also just pay people to solve recaptchas all day.
Never had your /login forum attacked with {uname,pass} combos? This is exactly what the traffic looks like.
Why isn't there a solid alternative offering yet?
> To pretend this is about slowing down bots is disingenuous as best.
I'm not sure you have a good understanding of what happens to internet services when they don't throttle spam. They become completely unusable.
If you have access to the server's information, it gets even better. Origin makes it much easier to identify likely spam, previous interactions with the site and their speed ("hits the page and 1s later submits a comment") provide more info.
Sure, all of that can be worked around, but that makes it more complicated and increases the cost for the attacker. If they spend money on faking actual user interaction with your blog, routing all requests through a residential IP in your country etc pp, they are likely spending as much or more than they would on a recaptcha solving service.
If clearly spam, block it, if clearly okay, allow it. If unsure, leave it for a moderator. It might even train users to write better comments if badly written ones need to wait for moderation.
Also, anyone who mentions <something the service wants to censor> can be, of course, blackholed as spam. ;)
The submission itself should be done by the site, so the user remains anonymous.
Any spam detection engine that asks for that information will outperform ones that don't.
For example, one week of stop-sign recognition, another week of pedestrian labeling, etc.
The latest version of recaptcha doesn't even prompt users. It loads on the front-end and uses a scoring system. It's likely you've used it but didn't even know because it's invisible.
It's the older implementations that have the slow loading images.
Recatpcha v3 by itself NEVER shows anything to the user and its entirely up the application itself on how to deal with users who are likely to be bots (which also includes users with anti-fingerprinting measures)
As far as I can tell, it just checks to see if your browser is Google Chrome to give you your score.
This is just splitting recaptcha into two pieces and giving you the first half. Okay, fine, but it's the second half that was causing all the problems!
Humans can do a lot of bad things that computers can do. Think of armies of low-wage people in Asia, that are paid to click on ads, spread spam, or write reviews.
And also consider that computers can actually do good things, for example, allowing humans to automate their work on certain websites, or providing better accessibility for certain users.
Therefore, instead of introducing CAPTCHAs, why not focus on the actual threats. If you want to protect against spam, then build a spam filter. If you want to prevent bots from bulk-downloading your data, then build a rate-limiter, etc.
I solved that in past by actually charging for my service. I think the internet would benefit from having more paid content and less ads driven stuff.
One thing that captchas do protect from is brute force attacks on user passwords. Although there are other possibilities (like making the connection slow after a number of attempts).
It's easier to fight against a bigger opponent (botnets, etc) if you mitigate their superiority in number first.
Recaptcha has saved the internet as far as I'm concerned.
[1] https://esolangs.org/wiki/Special:CreateAccount "Which number does this Befunge code output: [...]"
We have a small set of anonymous edits every day, which go through recaptcha to be allowed.
Sometimes I wonder if there was a real person on the other end writing code to combat the code I was writing at the same time, and finally gave up on the 3rd iteration.
Classic example of collecting more information than what is needed.
Is it? What's stopping AI from developing an identity in the eyes of Google? An AI that might behave exactly like Google's ideal user. Searches random stuff on their search, looks at and clicks their ads, logs in to various Google services, and when it encounters a ReCaptcha, it clicks the "I'm not a robot" checkbox.
At some point, caring about privacy might turn out to be the distinguishing feature of humans.
EDIT: I now see the article does actually mention this, though I still do wonder how far the fingerprinting goes.
Let the user create account, let the user create first comment/post, save it somewhere but hide it, then require ReCapcha to make the post visible and activate the account.
The issue with requiring ReCapcha before is that the website owner will never know whether a legitimate user was turned away or whether it was spam.
If you turn both spammers and annoyed users with ReCapcha before any interaction,you will never know how many were legit and you won't be able to manually accept interesting legitimate content written by a user that is annoyed by Recapcha.
Literally thousands of spam comments per day would stop instantly on deployment. We put this on over 400 sites and never had an issue of customized spam.
Definitely agree with recaptcha not being necessary. However, it's probably needed for popular sites which you're more likely to use. Do we see more recaptcha because of that?
It will also annoy users of password managers with auto-filling capabilities. "password" is normally used for actual passwords.
Besides, nothing stops the attacker from replacing your code with a faster implementation.
There are many methods that are easily dismissed with "but not all spam is like that" and "someone could work around it easily":
• blocking of links (if you don't need them, or in fields that are not for them).
• blocking of obviously spammy keywords (or bayesian filter)
• invisible fields and syntax to trip up dumb implementations
• requiring JS, properly-functioning cookies
• blocking of IP ranges that belong to VPS providers and a couple of 3rd world telecoms that allow spam
Each of them is surprisingly effective and combined they block 99.8%.
You really shouldn't flatter yourself thinking that a spammer will even look at your page. They have literally millions of sites to spam, and they couldn't care less if they get yours or not. A bot will find a 1000 other sites to spam quicker than it takes a human to click "View Source".
Most of spam is done by amateurs who take shitty off-the-shelf spam software, seed it with a target list copied off some forum, and run it on a couple of spam-friendly or incompetent VPS hosts. By volume, this is the vast majority and it's very easy to block.
This is what we did on one of our sites. 5 minutes to implement it using a few lines of code. Same result, couldn't be happier.
If I know my email address and password I should be able to login without being logged in to Google as well ...
Hidden inputs, honeypots, etc. None of them worked long-term.
Our solution: focus on the content.
All content now runs through the following filters, and takes care of 99.9% of spam: 1) Auto-approve any and all content immediately, unless: a) The content contains any HTML (url included). If it does, send it to a moderation queue. a1) If the submitter had submitted content previously and gotten approved, auto-approve their content, even if it contains HTML.
Admittedly, so is ReCaptcha, so the trade-off may be necessary as much as it sucks. But, it's probably at least worth a mention?
https://gizmodo.com/thai-click-fraud-farm-busted-using-wall-...
I’m thinking about how Overcast uses a token linked to the user’s iCloud account and doesn’t require a username and password if you only use iOS devices. You can optionally add a username and password to access the web client.
Many of the more sophisticated ones prefer emulating mobile application requests to web requests, so yes.
I found that some uncustomized spambots are able to solve simple math challenges like this one.
And the color question is bad for accessibility.
Many of the suggested alternatives are also terrible for accessibility, because they require solving a visual puzzle, without an audio alternative.
That's a valid concern, but not a reason to use ReCaptcha, because ReCaptcha is worse for accessibility.
As long as a browser provides a cookie that is present on your server, there is no need for another ReCAPTCHA.
If the user misbehaves, e.g. too many wrong guesses of password(s), remove the cookie from your server.
You can even email the user the cookie as a url e.g. example.com/?cookie=12345678901234567
Are you sure it is HN that uses ReCAPTCHA?
If most users don't even know that ReCAPTCHA is used, that's a good sign that it is being used as little as possible, though.
Either HN has changed or they now conditionally load it. For example, I encountered it every time I used Tor on HN, though I haven't hit /login in months.
It just makes use of the non-transitiveness of friend.
there were a couple of feints as if he were about to get to the content, then ha! back to diatribe.