Uncaptcha2: Defeat ReCaptcha with Google Speech2Text
github.com
github.com
[1] https://cloud.google.com/speech-to-text/pricing
[2] https://2captcha.com/, the first hit I found with the search [captcha solving serving price]
Disclosure: I work for Google on security and cloud, but not on anything related to captchas or speech to text.
The unCaptcha paper and the team's research is very much relevant, because it informs the public about the effectiveness of these security systems, and it helps website admins consider these threats and possibly adapt to them.
It's a bit difficult to miss two words displayed on an image.
It's incredibly easy to miss a traffic light 30 meters away from the camera on such a tiny photo.
Also, the automated method is probably more reliable than humans and much faster. And the cost of the speech-to-text API could be lowered by using cheaper services or an in-house model.
(Looking at other services, they all seem to agree on the $0.20-$0.40 range, mostly dictated by the hourly wage of their workers)
One of these days I'll gird my loins and go into battle to convince a bot that I'm not a bot. One last time.
However websites must specifically allow the noscript version to be used; by default it's disabled for all websites.
4channel.org is one site that allows noscript captcha, you can try it out there. But I've rarely seen it possible on other sites.
It seems so. That's not the only factor but it seems to be strong one.
(in most cases, i just close the tab if i get a captcha challenge-i’m sick and tired of the tracking)
> Google doesnt care about it, they are totally OK with it
Google hasn't said they don't care about it, where did you see any of that ?
They merely allowed the code to be released despite it still working against the current. Previous experience (namely the original uncaptcha) prove that they intend to find a way to fix it.
> You can see here that captcha doesn't block robots, but blocks people and makes browsing inconvenient.
Total BS, remember it's not Google that uses it, it's website owners (us), if what you claim was indeed true we wouldn't be using it, we would use something else that did what we wanted.
> reCaptcha is a way google mines data from us for free.
Of course, through the visual selection it displays when "unsure", although I do not know the detail it seems pretty obvious that once it's sure you're human it sometimes ask you to detect things in picture anyway so as to provide training data (for maps, waymo, image search, whatever ...)
> Google hasn't said they don't care about it, where did you see any of that ?
Check this out:
> The Recaptcha team is aware of this attack vector, and have confirmed they are okay with us releasing this code, despite its current success rate.
Source: https://github.com/ecthros/uncaptcha2/blob/master/README.md
> They merely allowed the code to be released despite it still working against the current. Previous experience (namely the original uncaptcha) prove that they intend to find a way to fix it.
Not saying this is the case here, but cargo culting does exist in tech, so this is quite a weak argument in my opinion.
I can say for my current needs right now that if created a non-subscriber posting content page tomorrow I would use a captcha because it removes enough bot to be worth it, and I would go with recaptcha because I find it the better one for end users (I as a user prefer to see it on websites compared to other solution).
So it's not at all obvious that known humans are being asked to solve captchas just for the purposes of training.
I have it semi-regularly (like once or twice a week); but I also have some automated tooling using my account AND I travel quite often so location testing probably flag we as weird.
Remember that the original recaptcha also did that with text to help train OCR (it would send a known word and a unknown word, if you succeeded at the known word it would record the answer for the unknown one, and after enough people gave the same one train it as the proper OCR'ed text).
A few months ago, my bank added the image-clicking one to its login screen. I've always gotten past that one on the first attempt. But with the same profile on the same Firefox, all other sites always take multiple tries.
Because believe me, if I get asked to click another 50 cars without good reason, (3 failed logins would be a good reason) I'll blame your site for being dumb and not google.
I have only been served image captchas since forever. I literally thought the warped text captchas had been phased out. I literally never see anything but image captchas.
(I work for Google, on nothing related to browsers or recaptcha, this is purely my impression from encountering it logged in and out.)
It has been since March.
Allow 3rd party cookie, log in to google and I only have to check a box.
Same when I use VPN, it does not accept any of the correct answers.
So when Google sees that I am trying to protect my privacy, it punish me by having to work for them.
One more thing. If try to use audio challenge in the first case, it directly told me that I am using some method to solve captcha and they won't allow it. So much fun.
From my experience, it works perfectly in a default session and not at all in the private browsing mode. I've never bothered to figure out why is that (possibly some other add-on interfering).
That's just code for "it rejects correct answers to frustrate you." If you manage to get the noscript version of the captcha with otherwise the same browser state it will accept a correct answer the first time nearly every time. Presumably this is because they didn't bother to implement their "hassle the user" code in the noscript version; it's probably neglected by google since it's disabled by default.
For instance, the sloooow fade in of challenge tiles... what legitimate purpose does that serve? That's not there to make it harder for bots. That's there just to hassle and punish real humans that google dislikes because they don't buy into the google 'ecosystem'. The more they dislike you, the slower the fade in gets. The fade-in can be several seconds long in severe cases.
Why on earth did they publish this?
I've kept it secret because Google will close this loophole and probably make it more difficult for disabled people to verify that they're humans. And Google is not dumb: They already know that speech recognition "breaks" their bot detection, just like screen readers - this is about accessibility. Publishing stuff like this will increase the pressure so they will be forced to "improve" their bot detection system - which simply means that even more people won't be able to solve those captchas.
Heck, some weeks ago I've tried to solve a ReCaptcha for literally 10 minutes! My answers were right, it was a matter of discrimination. My point is: My bot automation is able to solve a Captcha faster than a human being. This is silly and ineffective.
And about the people who've published this: they think they do someone a favor with this. But I can't see how it's in anybody's interest to release this into the public (especially on a site like HN where Googlers are reading). If they would propose a better solution for website owners to secure their sites, fine.
But everyone who's talking about "vulnerabilities" like this makes it more difficult for real people to access the websites that they want to use. I know disabled people who can't solve those captchas - it's just too much of a hassle while it's easy for my bot automation to do it.
We should really ask ourselves what we're really trying to improve here.
Bonus: the talk is HIGHLY entertaining. Their approach gets counter measured by Google an hour before the talk and so they can’t demo it anymore.
This was entertaining. Thanks for posting it.
The point of recaptcha is blocking "captcha farms" or automated bots from abusively creating accounts, buying tickets, etc.
The author hasn't demonstrated that this attack is effective in those scenarios. The only thing he has shown is a very convoluted way for a human to solve a recaptcha (harder for 99.9% of humans than the standard recaptcha experience)
That would explain why Google didn't care about them publishing this.
2. Navigate to a RECAPTCHA.
3. Click the "I Am not a bot" checkbox.
How is this not an exploit? Is Google doing something extra when it detects a screen reader?
Steganography prevents for people to use their exact service while making it more difficult in general prevents people from using any existing service.
Of course the next step will be to resample the sound to remove the steganography...and the arms race continues.
Verification seems really hard.
How do you go about verifying very other not the user is an actual person?
their own sites are protected by additional measures to detect bots like monitoring mouse movement
That was v1, they shut it down in March 2018. You won't see it anymore anywhere.
> "Google's current version is either the simple checkbox (which I assume is checking various things) or the image based version where you have to click on traffic signs."
That's v2. v2 will present you with a simple check box if you're very compliant with the google surveillance system, or will present you will image challenges if you're not (or if it's just in the mood.) v2 is very capricious and will reject correct answers from users google wishes to punish for, e.g. using firefox, using adblockers, using resistfingerprinting, blocking google's cookies, etc.
The recently released v3 is the worst of them all; it does away with the image challenges of v2 completely. The user never interacts with it directly, never has an opportunity to persuade v3 that they're a real human by answering any sort of questions. It's nothing more than a measure of how compliant you are with google's surveillance.
I did a very quick experiment with reCAPTCHA v3:
* using Firefox in private browsing mode (not logged-in to anything)
* with a VPN
* using uBlock Origin
* Do-Not-Track on, disabling 3rd party trackers
* Only went to one page and filled one form with garbage data
My score was 0.7, which is pretty decent I would say.I did a similar experiment using Ghost Inspector (a platform for automating browser testing, something similar to Selenium, but not sure what they use exactly), and my scores were consistently 0.1.
I'm also a bit suspicious of Google, and have trouble with the fact that this is the only solution on the market, and it's free for websites to use. But I'm not sure your statement is entirely accurate judging from my very limited experience.
Edit: If try to use audio challenge in the first case, it directly tells me that I am using some method to solve captcha and they won't allow it. So much fun.
Google is evil.
I don't know why twitter blocks my Android Firefox, but I feel as if we both benefit.
For some reason, it is exclusive to the links opened from another apps and doesn't appear when accessing a link directly (via a refresh). That might narrow down your search a bit.
I think they cap guest views/minute.
Also t.co links from within the Twitter app don't load at all from me until I hit "open in Firefox" on the drop-down.
First off permission was never needed to release the code. Second, Google's interest in captcha is not to protect websites but to further their machine learning algorithms. Unless the captcha mechanism was destroyed to such an extent nobody used it they will be happy to accept captcha requests from automated systems. Google may seem like it sometimes but it is not your friend.
Captcha -> request Audio verification-> download mp3 and send to google API for recognition -> 91% accuracy