NopeCHA: Captcha Solver
chrome.google.com
chrome.google.com
Today, CAPTCHAs serve a similar purpose, except they’re used to train self-driving cars’ image recognition AIs. I always try to be a little subversive and correctly identify the images that are clearly unambiguously classified by the AI, and then purposefully screw up identifying the image that the AI struggles with. It lets me through the majority of the time, which indicates that my bad input made it into their training data.
Unlike the CAPTCHAs of yore, when machine vision simply was not advanced enough to solve them, anyone has access to pre-trained vision models easily capable of identifying the unambiguously resolved buses or crosswalks in the CAPTCHA image. The deterrent to spammers is no longer that actual humans need to solve the CAPTCHA, but rather that it’s too computationally expensive to solve them at scale. Today’s CAPTCHAs are basically Hashcash proof-of-work [0], but with the added benefit to Google et al. (and annoyance to users) that they help train computer vision models.
The only reason we still solve those stupid image recognition puzzles is because Google/Waymo and other self-driving car companies have managed to trick us into helping them do their training work for them.
The techno-optimist in me says because I want to force them to improve their underlying models. When their engineers notice that their model struggles with weird edge cases that I purposefully mislabel (e.g. when prompted to select images containing motorcycles, I also pick a mountain bike with fat, motorcycle-sized tires), perhaps they will contemplate how to rigorously encode the concepts of “motorcycle” and “mountain bike” into their model, rather than simply pushing an abundance of training data through a black box classifier and hoping that by adding more crowdsourced data, it will eventually arrive at the right answer.
I think it's reasonable to believe that real self-driving cars are not inevitable, or even if they are, deliberate disruption of this process is healthy; e.g. it shouldn't rely on something this dumb.
If I ever learn that they release that dataset to the public, my position on this may change.
edit: I get twice as angry when I have to fill in captcha in a service that I pay for.
I’ve always assumed each new input is tested multiple times on different humans for validation, but this might be incorrect.
i.e. even if the saboteur rate was as high as 10%, and I only showed images three times, only 10% * 10% * 10% = 0.1% of data would have three people intentionally picking the wrong answer. I suspect the rate is much lower, and 99%+ people just want to pick the right answer to get the captcha program to go away as quickly as possible.
Images with less than 3/3 matching results in this example would presumably be retested until you got the desired confidence level.
Then assuming your ML model isn't overfit/overtrained, you could even then assess your original input data to detect/flag anomalies for manual review.
4chan had a lot of fun with this when it was first implemented. Perhaps unsurprisingly, there very quickly developed a campaign to have everyone insert "n**r" in place of the unknown word. Many threads were dedicated to education, onboarding, and, of course, sharing 'trophies' when such a replacement was found to have taken effect in one of Google's products (Books, iirc?).
Maybe the system has determined that you are human and you are intentionally attempting to mess with them. Since the primary goal of CAPTCHA (confirm that you are human) has been fulfilled and you appear to not be a good source for the secondary goal (crowdsource training data), the system decided to not waste any more time with you.
Was I the only person who has always inserted some nonsense word as my answer for the clearly scanned word? It was very obvious which word was generated and would be checked, and which one was scanned and which wouldn't be checked just accepted as-is. I always just typed in something else other than the word - I think it was just being contrarian against being used by a corporation to do their word recognition for them for free.
But yeah, I also cross the street on the red light.
Want reliable input? Pay me.
Typically, the API is a screen recorder and the CAPTCHA is sent to thousands of workers who essentially mini-remote-desktop in and solve them for about 80 cents/1k CAPTCHAs. Here are some other, similar services: https://0captcha.com/, http://bypasscaptcha.com/, https://deathbycaptcha.com/
I'm surprised these players are still around. They've been operating for nearly 20 years back when I had discovered them.
The entire industry is actually not completely as black hat as you might think. Yes, it's used for spam and botting, but at least at the time a lot of people used it for bulk downloading, which is how I discovered it. Additionally, it does provide work for the poorer parts of the world.
Going to the Google SSO page for their signin flow and clicking on the blue domain name for their app, the Google auth page shows the email of the GCP account that started the auth project, which in this case is jaewany@gmail.com
Looking that up on Google shows that it corresponds to Jaewan Yun.
Looking him up on GitHub gives you his profile which contains some captcha solver extension code for this very website and also many TensorFlow-related things.
His personal website[1] also lists the solver under "My Products"
Once the AI is good enough, I can buy a bunch of used GPUs from former ethereum miners, throw them in a cheap DC somewhere, and undercut everyone else! Sounds like a decent side project that could yield a bit of passive income. Somebody else has probably done it already. Maybe OP is that somebody.
You don't need few million samples, with 500-700 images per category you are more than ready to solve current captchas.
here is the link https://dashboard.hcaptcha.com/signup?type=accessibility
*edited typo
Please elaborate.
“workers (ie humans)”
Do you know what we call a process that takes tasks from a queue?
> We may share Your information with Our business partners to offer You certain products, services or promotions.
Incidentally, in almost all cases, if I'm faced with a recaptcha, I just don't do the thing. I have foregone purchases and charity donations, and not used products, because organizations care so little about their customers that they think making us solve a puzzle before we give them money is acceptable.
The main purpose of recaptcha is to prevent bots from abusing services, and has much less to do with exercising “monopoly power to … track you”.
It's funny when people think they are adding to the conversation by contradicting a thoughtful and interesting comment (even if it may be a bit conspiracy-theory-ish), by simply re-reciting the corporate line.
It's not free. It increases friction, and at least in my case, results in abandoned transactions. I'm not well versed in the different options for spam protection (or the attacks) but I do know that most merchants don't make their users solve a puzzle, especially at a critical point along the purchase workflow where is it most likely to get derailed.
The fact that google is (probably unintentionally) particularly appealing to small providers or charities, pretending they offer a "free" product, makes it even worse.
Edit: not an endorsement, but elsewhere in the discussion someone posted a link to cloudflare's captcha solution, which they say specifically addresses the privacy and annoyingness concerns of Google's captcha. So there are options: https://www.cloudflare.com/en-ca/products/turnstile/ (I'm not actually familiar with this, it may have a downside I don't know about)
(Also, disagreeing with something is generally a poor reason to downvote. It's much better to have a discussion, and I appreciate your comment)
True enough. That’s why I don’t use it. Google’s solution is absolutely awful, and I’m positive you aren’t the only one abandoning important flows on non-profit websites because of it.
A simple case of "subject does X and gains Y benefit" can, at scale become something like "subject is tasked with X, some fraction cooperate, some fraction defect, plus there are other induced effects such as cost of provisioning / supporting service S under various attack modes".
So:
- Without CAPTCHA, the service might be entirely nonviable.
- CAPTCHA tends to come with a large set of additional data-tracking elements and aspects. (E.g., I've got to enable multiple Google-domain JS in order to log in to several non-Google websites.)
- CAPTCHA itself directly consumes people's time, and thwarts legitimate use of numerous sites by many people.
- CAPTCHA and other countermeasures often mean that basic HTTP-based Web access is no longer viable. E.g., Internet Archive and Worldcat (two domains I make heavy use of) are no longer accessible via a terminal-mode browser. As I'd had (and still have) numerous terminal-mode query quick-lookup tools, this means I've now got to 1) break my terminal workflow and 2) invoke the full resources of a GUI browser (and usually a very limited set of very-heavy-weight such browsers) rather than run a quick one-liner on the terminal / command line.
(I'm not going to remotely pretend that this is a frequently encountered use-case from providers' perspectives. It's a frequently-encountered use-case from my perspective, however, and impacts strongly on various command-line, terminal, batch, script, automated tools, etc., and the value that these provided for Web interactions. Yes, in many cases, because of bad-faith / bad-actor abuse of those capabilities.)
- Measuring the net beneficial value of interactions is ... hard. A doctor looking up information probably has greater societal value than a bored pensioner or a pub's quiz-night team looking up answers to a game question. Discerning those at the Webserver level is ... difficult. Total requests is easy to measure, if not necessarily informative. W. Edwards Deming rolls in his grave....
Sure, you have to emulate or simulate the client JS challenges but when bots are running browsers in the background you can only do so much.
I wonder what the future of captchas, if any, will look like.
0: https://blog.cloudflare.com/eliminating-captchas-on-iphones-...
1: https://cloudflarechallenge.com/
2: https://blog.cloudflare.com/introducing-cryptographic-attest...
Open google.. captcha... every page has a 5 second cloudflare page before opening the page itself.
Bots have the time, they can wait and do other stuff in the meantime, but we, humans get bothered by that.
Google accounts give you a good score and tend to deliver easy captchas while dealing with Recaptcha; however, for this reason, google accounts are being sold and bought constantly.
People have tried similar fight tactics in the past. SMS and phone verification have failed because the return on investment is far greater than the price barrier it adds to get any of those "virtual identities".
iPhones might work but then, for how long? If you guarantee that an IPhone won't get captchas, it's a good investment to buy many old(or new) ones and sell token access to skip any captcha.
Many farms already have thousands of phones scrolling through youtube videos to get views, likes, and other stats for videos/channels.
The same "logic" applies to yubikeys and similar auth hardware; attackers can exploit it similarly.
Companies will tell you that they have abuse policies and actively fight abuse/bot farms, but again, they are not solving a problem but solving the problem with tape.
ReCAPTCHA was very useful for a while, it did genuinely stop bots reasonably well, but none of the "newer" versions seem as efficient as the older versions used to be. Progress stopped after V2.
[1]: https://addons.mozilla.org/en-US/firefox/addon/buster-captch...
https://www.cloudflare.com/products/turnstile/
Don't get any goofy puzzles which is nice.
Plus, ReCAPTCHAv3 makes this entire attack irrelevant by making image classification not a part of the CAPTCHA.
There is no way this doesn’t get abused, including, probably by the Company making it.
So, I’m dreading recaptcha v4
We may already be passing the captch event horizon where machines actually outperform humans on the damn things.
Maybe given a large enough input, but do you want to spend 10 minutes solving a captcha?
Simultaneously humans can be less likely to want to pay than bots which can skew the bot to human ratio.
How would the company making it abuse this? I feel like maybe I'm missing something obvious.
The extension probably has a hidden limit of 50 solves a day or something
For example ask him to upload his video with his ID. This video will be verified by another human operator.
In the end, user will be given some kind of identifier. He should present that identifier to anyone asking if he's a robot.
Of course that kind of verification will be paid. So you're paying $100 to get a verified identifier and then you keep that identifier (probably in the form of private key with signed public key).
There will be multiple certificate authorities who will issue those certificates to people. Rest of software companies will trust those authorities.
You need to renew that certificate every year.
If someone spotted your certificate being used in a nefarious schemes, your certificate will be revoked and you'll need to pay $5000 fine next time you'll ask for new certificate.
If you don't possess certificate, you're not qualified to be a human.
<https://toot.cat/@dredmorbius/104371588129861216>
It's faster at granting site access than cooperating with the intended challenge more often than not.
I would say this is useful for spammers and snipper bots.
However some shitheads like Discord also wont tell you how many, and will also outright fail you if you click too few, forcing you to restart the whole multi-test process. So fuck all of it. I fully support this extension, they deserve what they get. They need to figure out how to make it hard to fake, without making it a nightmare for legitimate users.
If there is a lion in the box.... Click it.
twitter/fb/google with their vast ml knowhow still cant figure out how to weed out the bots, even the ones that are doing obviously bot actions, at this point we are just increasing the tracking and are in diminishing returns, so maybe its time for a paradigm shift of some sort.
show your id to read a blog because the blog has ads and we want to know you are not a bot, is this where we are heading?
maybe even worse send your faceid and touchid hashes or crypto signed heart rate?
> Follows recommended practices for Chrome extensions. Learn more
Is this actually how it's done? If it is, how can an AI beat that? Thanks!
Might anyone have any links or resources regarding how these work?