CAPTCHAs decrease conversion rates
90percentofeverything.com
90percentofeverything.com
We operate a forum with 250k members and ~800k posts per month, a new registration every minute and we get so many spam bots even with captcha (mechanical turk etc) and without captcha it's unworkable. Captcha is a necessary evil, but it does help.
This seems to be coming from someone dealing with a site where spam wouldn't be that much of a problem, who would sign up to animoto to spam? Very silly post.
What's so silly about that?
> For some reason this article has hit the front page of Hacker News and is getting quite a lot of traffic. I should mention that yes, I acknowledge CAPTCHAs are of course sometimes unavoidable. That doesn’t mean, however, that we should ever feel good about using them, nor should we fool ourselves that users don’t mind them.
Which was my point. When spam is a serious issue then captchas are unavoidable 99% of the time.
Spammers are probably not targeting your website in particular, rather the software your forum is run on. If you add atypical anti-spam measures you'll separate yourself from others using the same platform, defeating the typical phpbb or vbulletin bot which probably accounts for most of your spam.
For juicier targets, something more sophisticated is necessary. Captchas are one answer.
So yes, who would spam animoto? But you know what, now their spam filter is good enough, their user registration has been increased. Better more "foolproof" techniques will always be needed, and directed attacks are hard to prevent, but getting to a good enough point is great as well.
And the problem you face may be a smaller percentage of sites than animoto, who do need captcha especially if you are the target of a directed attack.
I honestly believe that CAPTCHA's are one of the most evil things on the internet and that there are many valid and better ways to avoid spam.
BTW im not just some random guy with an axe to grind over this, I wrote http://www.wausita.com/captcha/ as an example of how trivial 90% of the CAPTCHA's on the web are trivial to decode.
- Honeypots: Add a field to your form that is styled to be invisible to normal human users, such as being located off the screen, sized to 1 pixel, or placed behind/under images on the page. Bots examine a page through HTML rather than through eyesight and will not distinguish these fields. Reject submissions which have entered text in the honeypot fields. - Timestamps: Some spambots operate by 'playback' - a human fills the form out correctly once, then copy-and-pastes the form output into a script that replaces the comment text/etc. with desired spam links. Place a hidden field in your form that contains a timestamp (possibly hashed or combined with other form output). Reject submissions which contain a timestamp far in the past, indicating a bot which is 'playing back' an old submission.
The idea with defeating spam is not to be 100% accurate with unbeatable security, since no matter your system, a bot tailored to your site can defeat it. However, putting several simple techniques together can defeat general-purpose bots that shotgun spam across many sites. This reduces spam to levels that are manageable by hand.
I'm sure blind readers would prefer this approach to a CAPTCHA.
"Thank you for filling out our form. Press <hotkey> to submit the form or <hotkey> to review."
"If you are not a real person, input your state now. Otherwise press <hotkey> to skip State : <input field>
This is absolutely the wrong thing to do - not that I'm doubting the conclusion, but the data does not support that confidence level.
The results are good only if you do a set number of observations (say 500) instead of waiting for a significant result (say it happens at 623). But what if you had decided to run 623 tests at the beginning?
for i in range 623:
data.add_result()
s = calculate_significance(data)
if s > 0.95:
publish()
for i in range 623:
data.add_result()
s = calculate_significance(data)
if s > 0.95:
publish()
break
The second one gives you many more chances to succeed, which must result in your confidence in the answer going down.Instead, I use a combination of human detection scripts, bayesian filtering, and moderation. Combined, this keeps the site pretty much 100% spam free from the perspective of our end users, and more importantly, Googlebot.
More details here:
http://www.expatsoftware.com/articles/2010/03/care-and-feedi...
Here's my algorithm:
/form-show
fieldhash = hash(ymd(today))
valuehash = hash(remoteip + ymd(today))
<input type=hidden name=fieldhash value=valuehash>
<input type=text name=email value="" style=display:none>
/form-validate
field0 = hash(ymd(yesterday))
value0 = hash(remoteip + ymd(yesterday))
field1 = hash(ymd(today))
value1 = hash(reomteip + ymd(today))
if(post[email] != "")
// reject form
if(post[field0] == value0 || post[field1] == value1)
// accept formThe latest thing I'm seeing on my site is a robot that automates real web browsers, jumps between ip addresses, scrapes real user content off the site, then posts it back using some form of Markov generator to make the content look unique. It'll do that on new accounts for weeks before trying to insert any links.
It's amazing the lengths spammers will go to to get their content onto your site. In this case, the crawler is clearly written specifically for my site, even though it's only PR4 and nofollows all its links. It's no wonder 99% of the content on big sites like Blogger is spam.
You can "sign" a timestamp by appending a hash of that timestamp with some sercre value. This way, whenever the user submits your form, you can reliably determine when it was requested without storing anything on the server.
Even hiding the CSS externally doesn't help, many bots now use inline browsers to decode javascript and stylesheets.
They say they successfully used timestamp/honeypots to keep out spammers; if so, how many spammers did they keep out? If it was tons, then say so, that's useful information. If it wasn't very many, then they didn't need the CAPTCHA in the first place.
I think not being a lazy developer in order to allow your customers to not make as much effort is a good thing. Only at a point where other methods dont work should you then employ CAPTCHA.
There is very little distinction between writing a phantomjs unit test and writing a spambot.
With "invisible" JS form validation of any sort:
1. you run your existing spambot software through phantomjs.
2. your unmodified bot fills in all the visible fields without changing a single line of code, and the webkit backend transparently computes your hashes and other automated javascript "human" tests.
3. again, your existing "stupid" spambot code submits your form, and your site is now overrun by spam.
With Captcha, you get an image and a unique ID that is validated at the server. Sure, you could run it through mechanical turk, but I'm guessing that a few CPU cycles to load a webkit backend is still vastly cheaper than farming work out to MechTurk.
My point is that you wouldn't even have to change your spambot software to defeat these "new" validations, and they can be trivially overcome, as opposed to MechTurk+reCaptcha. Add to that the benefits of targeting sites that are relatively spam-free, and you have a real incentive for spammers to simply plug-in phantomjs instead of using WWW::Mechanize or what have you.
captcha = sign of clueless or lazy, or both, developer. I don't put up with it anymore - I have yet to meet a single registration that I actually need that uses a captcha. I'm not the only one, either.
I think the goal is to make spammers lose the arms race simply because the payoff has become too small. If we can do that without CAPTCHAs, so much the better.
So start with a pessistic view they they are, and that they need to be shown a CAPTCHA. Then do some analysis to try figure out if they're legit, e.g. time spent on page, mouse/keyboard interaction, geo-location, referrer etc.
If they're all good, don't show them the CAPTCHA (perhaps just rely on honeypot inputs), otherwise show them a CAPCTHA as a next step after posting content (and apologise in case it's a false positive).
Any thoughts?
Add 100ms-of-2011-avg-cpu computation and tie it to the submit button (avoiding any complications interleaving with user activity). So that deals with first-order dumbbots and makes life a little harder for Javscript-executing (but still volume-based) folks. Marry to a bayesian system to handle the third-order mechanical turk-style miscreants.
For curious people, Wikipedia also has related information: http://en.wikipedia.org/wiki/Proof-of-work_system
All it will do is increase the latency, the overall throughput will be pretty much unaffected.
Some of my forms also have a CAPTCHA. I think it's got to be case-by-case. Do you have something desirable to bad guys (like the signup for a new Yahoo account, or a high-ranking blog about pharmaceuticals)? Do you have tools in place to deal with spam submissions effectively when they do occur? Will a bunch of bots signing up for accounts degrade service for legitimate visitors?
For example, the Contact our Sales team form definitely does not have a CAPTCHA. The sales team will gladly sort though a pile of junk if it means one more inbound lead. But the Post a Comment form would be an absolute disaster without a strong CAPTCHA. A surprising amount of junk gets through anyway, in fact. (As far as I can tell, it's actual humans in developing countries copy/pasting into comments by hands. Blocking referrers from Google that have the phrase "post a comment below" made a dent)
Can you elaborate? I haven't heard this technique (I don't personally have a lot of need for spam fighting), and I'm very curious as to what you mean.
Edit: obviously you could just avoid using this phrase on your site instead.
Thanks
However flawed the experiment might've been, it's obvious that if you add barriers (e.g., CAPTCHAs) before some end goal and detract from user experience then you decrease your conversion rate.
P.S. I'm working on the project to make CAPTCHAs more usable. We will have some updates soon. :)
If he wants to increase conversion rates, he should get rid of the irrelevant fields such as date of birth, zip code, country, gender, and check-to-agree to legal contract.
Ha, checking the actual site, "sign up" leads to "pricing" and not a sign up page. So much for their grave concern about losing sign ups at each stage.
On the other hand, his link to an article about including Honeypot fields is good advice and valuable. Timestamp analysis is not so great since it requires javascript and cookies. The more stuff you require the more users drop off. The problem with captchas is bad captchas that are impossible for humans to decode. Sometimes the reason these are used is because simpler captchas are implemented in a faulty manner that allows spammers to decode them without even having to do OCR. So the site developer upgrades to more complex captchas rather than fix the underlying problem that is breaking the captcha security.
Couldn't this be used to increase the security of computer systems? What if one could extend this to be able to tell particular machines from humans, human/machine combos, and counterfeit machines. I suspect one can do this. I have been working on this problem for the past 3 months, and I'm about to implement it and publish it on the App Store.
If it's less than something reasonable for a person (say, 20 seconds or something), then it was clearly auto-filled.
With a little help from javascript, you could even expand this to the individual fields.
As someone mentioned in the comments above, it would be pretty trivial for spammers to adapt to this if they thought it was common, with a few random pauses. Perhaps they already have...
I believe the Honeypot concept that has been discussed on here is referring to creation of a honeypot field on a web form, tempting the bot to fill it in. Many bots will blindly try to submit something into each field, just to make sure that they get all the required fields on their form submission.
By adding a honeypot field, and adding text that instructs humans to leave it blank, a very high percentage of bot submissions will be detected, with few false positives.
Furthermore, you can hide the field from humans, with CSS tricks, as others mentioned. Make it 1 pixel. Make it hidden. etc.
"How does a honey pot catch comment spammers?
In addition to including specially tagged spam trap addresses, some honey pots also include special HTML forms. Comment spammers are identified by watching what information is posted to these forms."
Here's a list of comment spammers they've caught:
http://www.projecthoneypot.org/list_of_ips.php?t=p
You're absolutely right that fake fields like that are a good way to catch bots, though, and that making your site unique is a great way to avoid being targeted by mass attacks that go after, say, all MediaWiki sites. Of course that doesn't help when you're big enough to be worth attacking specifically, but it makes things a little harder for the spammers.
I assume the spam is there in the first place to increase search engine rankings; so why not update the Google ranking algorithms (for example) to identify this spam and immediately give the targeted site (but not the site with the spam on it!) a terribly low rating?
Then, hopefully, the incentive to spam in the first place is gone.
1. Whois IP address of spam accounts.
2. Identify bad blocks of IPs. If it's a datacenter, someone is probably running spamming software on a dedicated server or VPS. Maybe get your hands on some of those open proxy lists that are floating around.
3. Use your data to prune bad accounts, throttle or block creation of new ones, etc.
This exploits the fact that bots don't usually run javascript or load all resources on a page.
http://docs.jquery.com/Tutorials:Safer_Contact_Forms_Without...
Pretty sure that 33% was bots, lol.
And they do train the bots to avoid negative-fields and timestamp analysis - all they have to do is look for type=hidden or display:none/visibility:hidden on the CSS
I use simple math instead of word captchas, seems easier on people.
I like to do this as a game, to see what I can get away with, adds some fun to the drudgery of typing in a captcha.
Of course your 'game' is hurting reCaptcha's goal of digitizing books.
Such complaining doesn't accomplish a thing, unless you tell them about an effective alternative. If you don't change anything about the trade-off they have knowingly made, nothing will change. To have any chance of convincing anyone, you at least need to explain the alternatives. Everyone that reads this post just shrugs their shoulders and ignores you, because their captchas effectively solve a problem they and their clients would suffer from without those captchas.
In this case, if you open with
Using a CAPTCHA is a way of announcing to the world that
you’ve got a spam problem, that you don’t know how to deal
with it, and that you’ve decided to offload the
frustration of the problem onto your user-base.
then I think it is very dissatisfying[1] to follow up later with They replaced the CAPTCHA with honeypot fields and
timestamp analysis, which has apparently proven to be very
effective at preventing spam while being completely
invisible to the end user.
which indicates that you have no idea about alternatives for fighting spam, apart from some measures that have 'apparently' helped in one particular case. It's not better than someone in a bar complaining about stupid government rules, without any idea or suggestion for how to improve things.[1] it said 'hypocritical' here. That is not the correct word for it.
That he offered up the word "apparently", even with strong evidence of proof shows that he's being an objective reporter and a good scientist. I'm disheartened that this would earn somebody ridicule here.
Your analogy does not stand true: security labels in real-world stores don't cause a percentage of customers to give up their purchases in frustration.
Having said that, in a forum like HN, most readers would expect both a statement of the problem and some proposed solutions. Frankly, when I posted this article, I didn't expect it to get onto the No. 1 spot on the front page. It must be a slow news day.
As for the security labels: a few days ago I wanted to fit a belt with an awkward security label that prevented a proper fit. It was an additional bump that wasn't overcome and may have been the only thing preventing me from buying it.