I’m not a human: Breaking the Google reCAPTCHA [pdf]
blackhat.com
blackhat.com
I don't work on reCAPTCHA but I imagine they could easily + significantly beef up image captcha by adding adversarial image examples that trip up would-be automators.
Note that adversarial examples generalize across models. Meaning a good adversarial example will work against SVMs, convnets, etc just the same.
If you're interested in the basics of how it works, Julia Evans has a great post: jvns.ca/blog/2015/12/24/how-to-trick-a-neural-network-into-thinking-a-panda-is-a-vulture
The key is that deep neural nets are actually very (piecewise) linear and that makes them susceptible to adversarial training.
That's more impressive than it sounds. I'm pretty sure 70.78% is more accurate than I am with reCAPTCHA manually. A lot of the captcha's presented are very fuzzy, or have ambiguous questions, etc.
I think that means you might be able to get source code via FOIA or similar :)
The best I've seen are the "which of these photos show mountains"-type challenges. I'd imagine that solving 5 rounds of those would take too long to make it worthwhile for spammers, but I'd also imagine lots of legitimate users getting irked at going through that to fill out a form.
exactly. Many reCAPTCHA are beyond simple recognition and make me start guessing. I expect we'll see new type of reCAPTCHA - you're a human if you make mistake and robot if correct answer is typed in :) Similar to those 1x1 images not visible to humans, yet visible to the robots.
If they were open, it would be a lot better (considering all the training is done by volunteers, too)
If they want to make a profit from me, they should pay me, or they can license correct CAPTCHA results from me under GPL.
A healthy system needs some kind of motivation. In economies, that is profits/money. What's wrong with that? (I know I am being simplistic here but...)
When doing a captcha, I don't really get anything. It's something I have to do because the website I'm using has the problem that they can't find a better way to identify bots. So I do a captcha, fine. But, if there's benefits for whatever company offers them, i.e. I'm doing work for them without getting anything in return, or without everybody getting anything in return (the GPL option), then I'm again not so happy.
Simple.
I buy a thing, I get something – in that moment the contract is over.
VW doesn’t come to me every 3 days with "You bought a car, to continue using it, take this and drive to Hanover and deliver it there".
When I bought my computer, or its parts, I bought them, I put them together, and that’s it. The manufacturer has never asked me to do work for them, or pay for them again.
I use ARCH Linux and Firefox – projects done by volunteers, and they profit by having a better product for themselves and others.
Google profits from ads, and from selling data. That’s a tradeoff, and a reason for me to try to use as few Google products as possible.
But when government organizations use ReCaptcha, and I have to work without pay for Google, then I have no choice, and it is not something I agree with.
This ended, and all trace of public good was erased (Check archive.org if you dont believe me) when Google execs mandated that only things that made money can be supported.
I just had a chat with our book guy today about OCR corrections using captchas, one of our volunteers suggested it. We expect to scan 500,000-700,000 books this year.
This is so true. When I first go an image captcha, it took me 3 minutes to learn how to solve it. I had JavaScript disabled so I did not realize I had to do 3 pages to complete the challenge (after two, I reloaded and got a new route thanks to tor the page since I though i message up). Also I'm not sure what to click when there is a partial match, is that still a match? And don't even get me started with solving signs (in my first challenge I had to tell apart different kind of signs that were in the image). Is a no trespassing sign still a traffic sign? What about if it is a sign but I don't know which kind of sign (due to domain knowledge – especially when they're foreign signs). Another problem is that they require you to understand the language the challenge is in. You can still solve a text captcha if you don't speak any English.
I wish there was a way to force text captchas
From the paper: 'Assuming a selling price of $2 per 1,000 solved captchas, our token harvesting attack could accrue $104 - $110 daily, per host (i.e., IP address). By leveraging proxy services and running multiple attacks in parallel, this amount could be significantly higher for a single machine.'
The last such analysis I paid attention to was some years ago so the situation may have changed, but I suspect the hassles of running a porn site and CAPCHA proxy still aren't worth it:
* obtaining content sufficient to attract interest
* paying for bandwidth & other resources)
* writing the authentication system
* then maintaining it (every time the CAPCHA service(s) change their process you potentially need to make and test changes to your code)
* and you need to work around rate limits (depending on the CAPCH design it may not be possible to make the relevant requests client-side so if the services has rate limits you'll have to route through something that sufficiently randomises your source address).
* providing support
* dealing with bad press
By understanding how reCAPTCHA worked – the team was able to double their productivity (since they usually only had to enter one word instead of two). To further optimize their voting they created a poll front-end that allowed you to enter votes quickly while giving you an update of the poll status (and since it is a 4chan kind of crowd, they also provided the option to stream some porn just to keep you company while you are subverting one of the largest media companies in the world.
https://musicmachinery.com/2009/04/27/moot-wins-time-inc-los...
However, this is slightly different as people were deliberately solving CAPTCHAs (and watching porn) rather than wanting to watch porn and also incidentally solving CAPTCHAs that they had no direct interest in.
From the paper: "We compare our performance to that of Decaptcher, the (self-reported) oldest captcha-solving service. We selected Decaptcher for two reasons. First, it supports the image reCaptcha, charging $2 per 1000 solved captchas. [...] Interestingly, some of our summitted challenges rejected due to the service being overloaded, and had to be resubmitted at a later time, and received a time-out error as the solvers did not provide an answer in the time window allocated by the service. 258 challenges (36.85%) were an exact match. When taking into account the flexibility, 321 (44.3%) of the captchas were solved. The average solving time for the challenges that received a solution was 22.5 seconds. While the accuracy may increase over time as the human solvers become more accustomed to the image reCaptcha, it is evident that our system is a cost-effective alternative. Nonetheless, our completely offline captcha-breaking system is comparable to a professional solving service in both accuracy and attack duration, with the added benefit of not incurring any cost on the attacker."
edit: Thanks for both responses!
Since their inception, captchas have been widely used for preventing fraudsters from performing illicit actions. Nevertheless, economic incentives have resulted in an arms race, where fraudsters develop automated solvers and, in turn, captcha services tweak their design to break the solvers. Recent work, however, presented a generic attack that can be applied to any text-based captcha scheme. Fittingly, Google recently unveiled the latest version of reCaptcha. The goal of their new system is twofold; to minimize the effort for legitimate users, while requiring tasks that are more challenging to computers than text recognition. ReCaptcha is driven by an “advanced risk analysis system” that evaluates requests and selects the difficulty of the captcha that will be returned. Users may be required to click in a checkbox, or solve a challenge by identifying images with similar content. In this paper, we conduct a comprehensive study of reCaptcha, and explore how the risk analysis process is influenced by each aspect of the request. Through extensive experimentation, we identify flaws that allow adversaries to effortlessly influence the risk analysis, bypass restrictions, and deploy large-scale attacks. Subsequently, we design a novel low-cost attack that leverages deep learning technologies for the semantic annotation of images. Our system is extremely effective, automatically solving 70.78% of the image reCaptcha challenges, while requiring only 19 seconds per challenge. We also apply our attack to the Facebook image captcha and achieve an accuracy of 83.5%. Based on our experimental findings, we propose a series of safeguards and modifications for impacting the scalability and accuracy of our attacks. Overall, while our study focuses on reCaptcha, our findings have wide implications; as the semantic information conveyed via images is increasingly within the realm of automated reasoning, the future of captchas relies on the exploration of novel directions.
"""
Live attack To obtain an exact measurement of our attack’s accuracy, we run our automated captcha-breaker against reCaptcha. We employ the Clarifai service as it shows the best result amount other services.
Labelled dataset. We created a labelled dataset to exploit the image repetition. We manually labelled 3,000 images collected from challenges, and assigned each image a tag describing the content. We selected the appropriate tags from our hint list. We used pHash for the comparison, as it is very efficient, and allows our system to compare all the images from a challenge to our dataset in 3.3 seconds. We ran our captcha-breaking system against 2,235 captchas, and obtained a 70.78% accuracy. The higher accuracy compared to the simulated experiments is, at least partially, attributed to the image repetition; the history module located 1,515 sample images and 385 candidate images in our labelled dataset.
Average run time. Our attack is very efficient, with an average duration of 19.2 seconds per challenge. The most time consuming phase is running GRIS, consuming phase, as it searches for all the images in Google and processes the results, including the extraction of links that point to higher resolution versions of the images.
"""
It's like reading a security paper authored by Gollum.
I'm glad to see research around how this captcha method is actually as breakable as any, but the real problem continues to be that there is no separation between the data Google collects for advertising, and the data Google collects from the captcha service.
Their privacy policy says they can link the captcha to your other online identity and use it to target you. It's especially cringeworthy to see lots of sites of dubious legality implement reCAPTCHA (like thepiratebay).
[0]: http://www.businessinsider.com/google-no-captcha-adtruth-pri...?
However, after 4 submitted URLs, I get a reCAPTCHA. From then on for every URL, I have to complete it with additionaly visual quizzes.
This may or may not help with your quest.
Google triggers recaptcha when i use their search using one of my digitalocean servers as vpn and incognito mode, thus my IP address belongs to a datacentre, there isn't a cookie header and the user-agent is linux.
(I used to use my OVH VPS for some time as VPN, so it has gained quite a reputation for a "normal" browsing history)
There are also plugins some SEOs use to pull lots of requests. These tools are quite old now and not very useful, but people still use them.
Google has confirmed it is removing Toolbar PageRank: http://searchengineland.com/google-has-confirmed-they-are-re...
Googles reCAPTCHA is hardly an effective solution... Also if google wanted to they could just automatically verify people without you clicking that checkbox. Because at the end of the day they already know if they are going to auto-verify you, or make you pass a test.
Just like with DRM, captchas only seem to punish legitimate users and perhaps very small-time bad actors.