Hacking and defeating Google's reCAPTCHA with a 99% accuracy
dc949.org
dc949.org
If not, then it isn't useful work, really.
Useless for transcription, because it 'validates' phonetically. This is part of the attack : Whether audio is "wagon" or "van", the 'word' "wagn" validates. Same for "Spoon" and "Teaspoon", which both can be validate by entering "Tspoon". (Since Ts can be said "Ss" as in "Tsunami" (Sunami) and "T S" as in T-Rex.)
1/ Homophony. Knowing how to write a word you hear often needs contextualisation. Giving two whole sentence for the user to hear is too long and takes too much time for the user. 'Right, but no need for a whole sentence, a few words suffice' Right, but if the computer knows where to cut these sentences, it can also transcribe it itself. 2/ You have to assume good spelling from user. 3/ Is it useful ? I mean, 100M recaptcha are done everyday, what part of these 100M are audio recaptcha ? 0.0005% ? Less ? Transcribing 2 sentences every week makes no sense. Keep in mind that homophony+bad spelling are two factors increasing hugely the number of times the same 'unknown clip' will have to go through 'human validation' until we can assume with a certain level of confidence that it has been transcribed.
Funny 4/, take a look at the [cc] button on some youtube videos : on the fly transcription. Thanks Google. :) Oh, and also on the fly translation of on the fly transcription, btw. Google even said they were working on Voice to voice translation for Google Voice : English speaker calls Chinese speaker, english voice transcribed, then translated, then synthesized, same the other way. :)
It's very difficult to generate some sort of noise via algorithm that a) humans can filter out and b) can't be removed by some algorithm. As a result, audio captchas are a huge vulnerability and the weakest link in almost any captcha system, although you can't get rid of them by law.
Hypotheticals aside, the code was easy to patch - note the footnote: > In the hours before our presentation/release, Google pushed a new version of reCAPTCHA which fully nerfs our attack.
All of this aside, removing background noise is not a huge issue anymore. We have pretty decent noise-cancellation technology. Speech recognition - the other big component - has advanced a lot in recent times and is actually pretty good, although not for every company/product.
Even if it would be helpful, you'd have to record an incredible amount of noise in the first place, seeing as you're getting millions of hits a day and if you have a small sample set, the attackers will just figure out the solutions to that sample set and be done.
I'm not saying it's impossible, but I am saying it's probably not worth it at this point. Captchas (in their traditional forms) don't make sense as a long-term strategy anyways.
No it doesn't. reCaptcha only checks one of two words it displays (the other one being what OCRs can't handle themselves), so naturally you only need to crack one and input garbage as the other, thus actually making the world a worse place.
To goal of good research (and one that differentiates researchers from criminals) is to present a proof of concept and to advance the state of security. The fact that security is a perpetual arms race is incidental.
Well, at least that's what security researchers tell themselves anyway to avoid going mad. :)
I spent the whole article wondering why they were so interested not only in presenting a proof of concept, but in getting as many people as possible actively breaking captchas. Then I got to the end and switched to wondering whether the whole thing is an elaborate prank.
On the other hand, it looks like they provide a corpus (http://www.dc949.org/projects/stiltwalker/stiltwalker-corpus...) [1.5 GB!] that you can still use to run the program.
However, every day five or six make it through. Observing their patterns and timing, and the fact that they make zero mistakes while interacting very slowly with the site, I can only presume that humans are directly involved as well, either as mechanical turks or simply manually posting the spam.
notably, they follow normal, guided flows through the site - not like robots at all, but hit pages that would obviously be interesting to a human (say, have a certain type of image content linking them versus another), while avoiding the least popular pages - never hit urls that are hidden via CSS, etc.
If you'd like to see it you can usually get it by putting in a bad password to a gmail account too many times (though I don't know if that has other consequences).
Edit: Here is an image example http://www.techian.com/wp-content/uploads/captcha.png
Most software developers are under 40 and therefore have little to no appreciation for what happens to people's eyes after they are past 40.
Anything that I can think of like adding to the CAPTCHA such as a half-line of letter above and below or adding noise would probably make cracking them programmatically even easier.
I find that pretty worrying for the future of the internet. An internet without working captchhas will probably be full of bots and spam.
Like many others, I can barely get through their captcha service. I'm actually happy people circumvented it. Maybe someone will think it through this time around.
There are symbols for visual impairment, but they're not international.
(http://commons.wikimedia.org/wiki/File:Pictograms-nps-access...)
(http://3.bp.blogspot.com/-HfEx4Y_O_Gs/Tf0huVPZBXI/AAAAAAAAAC...)
Anytime you're ready, we're listening.
http://arstechnica.com/security/2012/05/google-recaptcha-bro...
the youtube video in the OP does a great job of explaining and the thinking behind the attack, and even though it's an hour long, it's worth the watch.
Plain text wins again, at least for one-way technical communication.
I will never understand why people think that the crappy audio channel of a youtube video of some guy "umming" and gulping through a poorly prepared speech, all recorded on a $5 microphone, is a suitable way of passing information.
Audacity can do this, and there's almost certainly automatible tools out there. Worth a shot, anyway.
But perhaps if one could learn two sparse representations that incorporated phase information for backwards and forwards speech using matrix factorization - it might work as a method for removing recaptcha noise. This idea is so simple though, I assume someone tried it and it doesn't work well.
Seriously, who adds reCaptcha to a login form?!
It's a good lesson in a form of social engineering. Sites have to provide this alternative access for the visually impaired...yet I bet the resources/creativity put behind it is not at the same level as the kind put into the catchpa used by 99% of the userbase. Furthermore, the most important client -- your boss -- is likely to not be blind him/herself, which eliminates that extra critical layer of oversight.
I take it that "fully nerfs" means this defeat of recaptcha is no longer useful?
Deciphering obfuscated text is one of very few verifiable tasks that humans can do which computers cannot.
Alternatively, you can look at other methods of identity verification. If you ask for $5 to create an account, you'll have few spam problems.
It may be easy to do today, but, going forward, how do we determine which bots are "good" and which ones are "bad"?
Clearly, simply being a "bot" does not imply "bad" intent. If it did then we should all be blocking search engine bots. Yet this is what reCAPTCHA does: it blocks not based on intent, but based on the characteristic of being a "bot".
There will be lots more bots and lots more usages. Things might not be so simple. For example, some sites might exclude all search engine bots except a chosen few despite the fact they all honor robots.txt and behave essentially the same.
(BTW, if by "text file" you mean robots.txt, isn't that an exclusion list? You seem to be saying it's an inclusion list.)