An Empirical Study and Evaluation of Modern CAPTCHAs
arxiv.org
arxiv.org
In fairness, the company best positioned to harness user input to an AI that avoids crashes would probably be Rockstar. OTOH, that AI would definitely not obey stop signs or pedestrians.
I don't know if it is or isn't, I never drove one, but those are two completely different standards
But I guess people these days just love to jump on the opportunity to hate whatever is trendy to hate at the moment.
0 - https://www.progressive.com/lifelanes/on-the-road/understand...
But they have? For years Google Street view has read signs, house numbers, phone numbers of businesses, etc. from the environment. It is safe to assume they have this built into Waymo as well.
I assume you might be trying to reference "vision only" self-driving, which is a fantasy made up by Elon Musk because nobody would sell him LiDAR sensors cheaply.
https://www.thedrive.com/tech/43779/this-tesla-model-y-dummy...
“Sour grape Elon, touting vision because no one will sell him LiDAR sensors. Which are the gold standard sensors that solve self driving.”
How exactly does LiDAR tell you whether the thing in question can move (dog) or not (trash can)? How does it allow a neural net to infer intent?
You’ll actually have to solve vision. Even if you had LiDAR. There’s no way around it. And once you’ve solved it, LiDAR becomes superfluous.
Chesterton’s fence.
LIDAR is continuously scanning, usually multiple times a second. It is irrelevant if the object is a trash can or a dog if it has a trajectory into the street.
> And once you’ve solved it, LiDAR becomes superfluous.
Only if you have the low Musk level standards of simply being equal to a human. There are plenty of jobs robots can do better than humans and driving is one of them. But it does require LIDAR and/or radar.
https://abc7news.com/tesla-s-autopilot-self-driving-car-offi...
I have an older tesla S with the pre-ai so called autopilot. It has one camera in the front and a radar and the system detects a few things like speed limit signs. The main extent of what it can do is follow the current lane pretty wall, even when it curves, slows down if it comes up to a car going slower than its preset speed. The good thing is it works on any road. It does a shockingly good job.
The later systems with onboard special processors are like a crazy beginning driver to has way too much confidence and drives in dangerous situations willy nilly. There are many other people who have explored it and written long posts. It's not safe. You can try to use it be you have to be constantly paying extreme attention. It's like watching your kid drive the first time. I know you should be watching the stupid ai all the time, but it's far from being safe.
"Good" drivers see questionable situations and slow down or position themselves farther from potential issues before they get to the issue so they don't have to react at the last minute.
Hardly the AI uprising, though definitely a good tool for anyone, good or evil.
What’s you point, here? That you should be lied to when you ask, or that it should refuse to tell you any kind of fiction?
I agree with your larger point that there will be ways to circumvent these systems, my only argument is that the lie/fictional story divide is a bad example because the line between them can be made clear with a single statement.
If I tell you that I watched C-beams glitter in the dark near the Tannhäuser Gate that is a lie. If I write the same in fiction I receive accolades.
If I tell you on the street “watch out there is a T-rex about to eat you!” That is a lie. If i say the same thing sitting at a table with too many dice that is just acceptable DMing and everyone rolls initiative.
Humans are weird this way.
So it's not an example of "going rogue", but it's not like a researcher told GPT-4 "oh, and make sure to lie to an online gig worker to get him to solve catchas for you". GPT-4 generated the "hire a gig worker" and "claim to be a human with impaired vision" strategies from the basic instructions above.
I don’t think it should be OpenAI deciding what is allowed or not though.
Avoiding lawsuits is what they are trying to do. They don't actually care about what you use their products for.
You're kinda saying if you hire Bob's Handyman Service you should be able to tell him to break down the neighbors door and cart out the contents of their house.
You're in a desert, walking along when you look down and see a tortoise. It's crawling toward you. You reach down and flip it over on its back, its belly baking in the hot sun, beating its legs trying to turn itself over. But it can't. Not with out your help. But you're not helping. Why is that?
Might do that unobtrusively for the average person, by using projects like mCaptcha [0] for instance.
Two versions that I experimented with. One is where the incoming POW hashes contribute to hashing power for some blockchain mining. An alternative "pay as you use the API" system.
The other using hashcash. Just a way to slow down abuse.
Both, however, suffer from the downside that many/all "ASIC resisting crypto mining" suffer from as well: the cheapest CPU power is CPU power from machines/power you don't own. Botnets, viruses, trojans etc.
So such a mechanism to throtthe or protect APIs won't hold back spammers and abusers for long.
You might correctly claim clean energy is often cheaper, but you must also consider the regions in which they'll get away with nefarious activity, and whether those areas have made the investments into making clean energy cheap.
Hmm, I don't get this, surely all actors will want the cheapest energy, no? The problem being the underlying one, that the dirty energy doesn't pay its externalities and is thus cheaper than renewables.
Yes, the only differences are that mCaptcha is 100% FOSS and uses variable difficulty factor, which makes it easy to solve Proof-of-Work under normal traffic level but becomes harder as an attack is detected.
mCaptcha uses PoW and that is energy inefficient, but it not as bad as the PoWs used in blockchains. The PoW difficulty factor in mCaptcha is significantly lower than blockchains, where several miners will have to pool their resources to solve a single challenge. In mCaptcha, it takes anywhere between 200ms to 5s to solve a challenge. Which is probably comparable to the energy used to train AI models used in reCAPTCHA.
The protection mechanisms used to guard access to the internet must be privacy-respecting and idempotent. mCaptcha isn't perfect, and I'm constantly on the lookout for finding better and cleaner ways to solve this problem.
Are you comparing the energy it takes to train a model which is bounded and defined with unbounded inference which can (in principle) go multiple order of magnitude depending on the usage? Or maybe I misunderstood what you are trying to say? then I apologize in advance.
From what I understand of reCAPTCHA, the model isn't static and is continuously learning from every interaction[0]:
> reCAPTCHA’s risk-based bot algorithms apply continuous machine learning that factors in every customer and bot interaction to overcome the binary heuristic logic of traditional challenge-based bot detection technologies.
I don't know the energy demands of such a system.
mCaptcha, under attack situations, will at most take 5s of CPU time on a busy (regular multitasking with multiple background process) smartphone.
I had not considered that. Naturally, we're just speculating here, but yeah that does sound plausible.
I was also no aware of the "hard" 5s bound (which you seem to have tested on a normal smartphone setup); sounds neat.
You could already see the writing on the wall with image identification years ago, when the obscuration techniques became more elaborate. It was an arms race. I was having trouble with them. I can see less technically inclined being able to use them. I imagined how much worse it was for people with color blindness, disabilities, or people forced to use them at public library computers because that is all they have.
Open source capcha projects have either not been clued in, or don’t have the resources to pull this off. Google didn’t just switch out which signals they tested, they also wrote an obfuscating virtual machine executing within the browser environment (if I remember that article taking about this correctly). That was years ago and who knows what they do now — for all we know, the “byte code” running the test is now a neural net of some kind.
I mean, I’ve sometimes had to try three or four times with certain captures and I have perfect eyesight (with my glasses). I feel so badly for those with vision or hearing issues with an empathy I never had when I was younger. They are so often simply forgotten.
I'm kinda surprised that ADA doesn't allow them to sue site owners about this.
I contacted several attorneys, none of whom would consider taking the case, or even bother to discuss the details with me. One of them told me that, at least in North Carolina, an employer would effectively have to get on the stand and explicitly confess taking adverse actions against me specifically because I had been diagnosed with MS. Any other remotely plausible excuse would provide them with all the cover necessary.
It was only much later that I learned that I would have had to have filed a complaint with the EEOC and NLRB within 180-days, and allow them to investigate my claims fully before authorizing such a lawsuit to begin with, as without such a determination I could not file the suit anyway. None of the attorneys I consulted even mentioned this absolutely critical first step, which suggests that they had even less faith in a successful outcome.
Maybe it’s different for facilities and regulatory enforcement, but in my experience, at least for labor, the protections are incredibly weak.
Funnily enough, AI may be better at solving them than people. I've encountered many Google captchas which reject the correct answers, because you know... bots trained it to accept incorrect ones. Anyway, at least it's not stop signs anymore. It must have been truly embarrassing that Google was simultaneously selling "self driving" cars but at the same time demonstrating that stop sign recognition couldn't be done by robots.
With coffee N(1,5 minutes, 20 seconds), toilet N(4 minutes, 30 seconds), ...
I wonder when we'll get to the point that employers can't tell the difference between transformers and real humans anymore ...
Things like Private Access Tokens: https://blog.cloudflare.com/eliminating-captchas-on-iphones-...
I have no problem with Joan over the road curtain twitching. It doesn’t scale. I have a massive problem with the 24/7 surveillance from ring though.
I have wondered if they keep the scan or does the state? I asked and the random hourly worker there said they don't.
Personal Data has to be treated as a liability, but too much of the economy treats it as an asset.
But yea, so many people are nieve of what the authoritarian types would do with data like that (looking at you Texas with your civil laws on abortion now).
The entire difference is that from my mobile phone I can send more traffic in an hour than most services will ever see legitimate traffic in their entire lifetime, and the cost to me is minimal.
The comparison is as invalid as comparing piracy to theft - piracy isn't theft, it's piracy, and understanding the difference between them is the key to dealing with the problem.
There are very few places in the real world which can handl 1,000 people per second.
In the real world I rarely need to identify myself. I can see a movie, visit the library, buy groceries, go to a restaurant, and more.
Hobest question, are you being serious here? The sxale of fraud and automated traffic is disproportionately large, and has a significantly lower barrier to entry than other forms of abuse. That's the entire reason.
> There are very few places in the real world which can handl 1,000 people per second.
Exactly, and if someone started sending thousands of people per second there, they would make it significantly more difficult to do so.
Most of the real world does not require identity, so how does "The real world doesn’t allow that" make any sense?
Yes, some parts of the real world require you to identify yourself, and the same for some places on the internet.
Is that really the point? That if you have to use your real identify to log into your bank's web site that you don't have "unconstrained anonymity"?
Because I don't think even the cryptopunks of the 1990s required that sort of anonymity.
> and if someone started sending thousands of people per second
So, 100/second is okay but 1,000/second not okay?
I ask because it looks like 100 people per second enter Manhattan during the peak morning commute time, and I don't see massive calls to make it harder for commuters to enter the borough. (Go to http://manpopex.us/ , go to statistics, "Estimated Pop. for Wednesday, 9 AM: 2,888,116", for "10 AM: 3,284,591" gives 110 people per second.)
And these people aren't all required to identify themselves.
Question for you: does the internet currently have more anonymity than the real world?
Question #2: how much fraud is done on the internet vs. fraud in the real world, measured by dollars?
And what if you develop this very sophisticated system of reputation score, what if bad actors find a way to still perfectly abuse it, e.g. they pay for desperate people for the IDs and then stay just within the limits ever so slightly.
Would you be able to easily iterate on the system when that happens to make it more secure?
But if you also track IP addresses then doesn't that already mean loss of anonymity?
And ultimately with something like IP address, a bad actor could offer you to download an app where they could simply use your IP address to post content/propaganda from under your ID and IP.
It would be more expensive for bad actors, but also I think there was period when Facebook accounts were bought and sold, and there was very active market for that. I imagine teenagers for example are really easily tricked into selling their creds etc.
Also Reddit and other social media accounts are being sold a lot, so definitely there would be market for that.
Regarding the IP address question, I’d assume you could decouple the IP address verification portions from the “know who the person is” portions with some clever multi-party computation. Someone always has to know your IP address, but it doesn’t have to be the same person you’re talking to. (Think of Tor as an inspiration here.)
People have been able to write anonymous letters and send them through the mail for a long time. Still can.
No one checks my id before I stick an envelope in the mail box.
I would not be surprised if there is some country that has a facial recognition camera network faced at mailboxes these days.
Even then, here is literally the first post box I found looking in the UK, in a small town: https://www.google.com/maps/@52.0936599,0.0761217,3a,75y,165... . No CCT in sight, no power, good solid iron.
Plus, think of how difficult it is to match a person to the physical envelope.
At best there could be a distinctive envelope.
Otherwise, yes, you can get a list of people who use the box. But for that to be useful, the mail from different boxes can't simply be jumbled together into the same pickup bag as that would broaden the number of suspects.
Why is it that you have to lose your anonimity when you are on the internet? The real world always allowed that until it became dependent on surveillance capitalism. Of course you need to prove you're yourself for some things, but that should be the exception. You could always look things up at your local library while being anonymous (for checking out you'd need a card), you could call from a payphone while being anonymous, you could use coins (cash in general) while being anonymous.
Anonimity was the rule and should still be the rule
on large cities everybody is anonymous to some degree
I'm not denying reCAPTCHA is a source of training data for Google -- surely there's no particular reason that every single reCAPTCHA V2 challenge is about identifying traffic objects, and it's not like Google is building a self-driving AI or anything.
But that's the business model, not the core feature.
And, that training data isn't just given to the developers of captcha solving bots.
And also completely incidentally making the web browsing experience a wee bit less pleasant for people who refuse to have google track their every click.
Like users of non-chrome browsers, adblockers etc.
Totally incidental I'm sure.
checkbox = getPos(checkbox='notRobot')
button = getPos(button='submit')
cursor()
.transition(pos=checkbox)
.click()
.transition(pos=button)
.click()
They now
checkbox = getPos(checkbox='notRobot')
button = getPos(button='submit')
cursor()
.sleep(time=random(distribution='human_captcha'))
.transition(pos=checkbox , method='human_captcha')
.sleep(time=random(distribution='human_captcha'))
.click()
.sleep(time=random(distribution='human_captcha'))
.transition(pos=button, method='human_captcha')
.sleep(time=random(distribution='human_captcha'))
.click()
Where sleep and transitioning are sampled from some random distribution that is close to actual human behavior, which should be pretty trivial to model.
Solving reCAPTCHA v2/v3 requires more than just clicking the box and an image puzzle. If that was all it was we would be overrun by now.
Lots of folks commenting that the title's statement makes sense because CAPTCHAs are meant to train AIs. While this is broadly true, that's a nice side effect. The way modern CAPTCHAs like reCaptcha V2+ work, is they monitor behavioral analytics-- from things like your browsing history to how your mouse moves on the page. This is why most of the time, most people only need to click a box. I'm not sure there's a LMM out there that includes mouse movement as a modality.
The kinds of AIs that are designed to beat CAPTCHAs also don't have the data from Google et al to use to train, unless we're concerned Google is training it's own bots to bypass CAPTCHAs, I suppose it's not inconceivable?
https://ieeexplore.ieee.org/document/7467367
From the abstract:
> Through extensive experimentation, we identify flaws that allow adversaries to effortlessly influence the risk analysis, bypass restrictions, and deploy large-scale attacks. Subsequently, we design a novel low-cost attack that leverages deep learning technologies for the semantic annotation of images.
I'd suspect reCaptcha has been updated in the 7 years since to address shortcomings.
Another entry in the table (citation 45) is from 2020 and talks about using an object detection AI to solve the image tests. This again looks like it's focused on the task, not the primary mechanism (behavioral analytics).
...Yet, I suppose.
in upi system, you are presented with a QR code or you input your UPI ID, you click pay and it gets through.
if you are worried about "fraud protection", why rely on an intermediary like ebay or credit card company and instead should take up with your bank or the seller or courts.
The Internet will become a dark forest, and since that is where all of our communication and transactions happen of any significance, that’s pretty much game over for the significance of human activity.
Think I am overstating the fact? It already happened with wall street trading. First, institutions prefer bots to human. Then, you will come to prefer bots to humans. Then every human will be surrounded with 999 bots and unable to change anything or appeal to any significant number of humans to change anything.
It also does not help that the shown busses, water hydrants, pavements look totally unfamiliar to me. (Why aren't captures taken from all over the world Indian busses would be fun - London ones would be too boring)
(Amusingly, pain was proven to be preferable to boredom... and CAPTCHAS are boring as hell.)
Something that can work on any browser can be like this: Scan the QR code in your iPhone or Android device that supports attestation. Will ask you if you approve login, then will attest for you. If you turn out to be a bad actor, the website can ban this device - so no flooding with a single device.
These things are always cat and mouse games.
Which is not a bad thing
Pretty sure it's an AST interpreter too (metacircular eval - apply, as in SICP)
The “outages” that are common are slowdowns for logged in users.
It really gets me when I have a 8 year old account that has made purchases and I still see them across the app.
The annoyingly common one is on login pages. If I am giving you correct credentials you don't need a captcha. If bots are an issue you should be doing per-account strong rate limiting, not a captcha.
This is very different from many other sites where the potential to make a buck is much more pronounced and direct.
Like other people reported, if you ever use tor, it's very common for the captchas to just not load. They just kind of hang without showing the pictures. Regular websites generally just work fine on tor, it seems to be a captcha problem.
I ask myself this every time.
Submitters: If you want to say what you think is important about an article, that's fine, but do it by adding a comment to the thread. Then your view will be on a level playing field with everyone else's: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
An Empirical Study & Evaluation of Modern CAPTCHAs
* For nearly two decades, CAPTCHAs have been widely used as a means of protection against bots. Throughout the years, as their use grew, techniques to defeat or bypass CAPTCHAs have continued to improve. Meanwhile, CAPTCHAs have also evolved in terms of sophistication and diversity, becoming increasingly difficult to solve for both bots (machines) and humans. Given this long-standing and still-ongoing arms race, it is critical to investigate how long it takes legitimate users to solve modern CAPTCHAs, and how they are perceived by those users.* * In this work, we explore CAPTCHAs in the wild by evaluating users' solving performance and perceptions of unmodified currently-deployed CAPTCHAs. We obtain this data through manual inspection of popular websites and user studies in which 1,400 participants collectively solved 14,000 CAPTCHAs. Results show significant differences between the most popular types of CAPTCHAs: surprisingly, solving time and user perception are not always correlated. We performed a comparative study to investigate the effect of experimental context -- specifically the difference between solving CAPTCHAs directly versus solving them as part of a more natural task, such as account creation. Whilst there were several potential confounding factors, our results show that experimental context could have an impact on this task, and must be taken into account in future CAPTCHA studies. Finally, we investigate CAPTCHA-induced user task abandonment by analyzing participants who start and do not complete the task.*
@dang, could you please correct the title? Thanks.
Captchas are purposely not made too hard as people like pex.com need to be able to bypass them for copyright enforcement. Note I’m biased as I was a founder of hcaptcha
As prices for bot operators decrease, website operators will increase the challenge and drive up effort for the intended website audience (humans) who are solving captchas instead of paying bots.
In the end, the website operators will have to stop using captchas as the intended website audience will no longer be willing to solve harder captchas.
Website operators can use alternatives, like asking for micro-payments, high enough to discourage most bot operators.
Similarly to how dApps work in ethereum-like blockchains?
https://uk.pcmag.com/macos/138058/not-just-iphone-how-to-use...
I've had a few instances on Windows 11 and surrounding software where ctrl-C as well as the context menu entry for 'Copy' were greyed out for this reason when skimming through logfiles, presumably because there was something about the line that triggered the MS "that's a password!" regex; stuuuupid stuff.
It's honestly annoying as it frequently interferes with dragging images out of safari, except on those occasions when I do want the text when it's super useful. I think the iOS interface just tells you there's text in an image or photo and gives you the option to copy it rather than cursor based selection you get on Mac.
[edit: from other comments it sounds like windows can do this but it's not always present, and not present in all circumstances, which makes me wonder how many cases in cocoa/uikit/swiftui it does not work]
If your concern is "apple is harvesting my data" then no. All of apple's various analysis systems ("AI") are entirely local. This does mean you get a bunch of duplicated work as every device redoes the same analysis but on the other hand it saves you from "how do we defend against a compromised network".
Even if that were true (I couldn't say, and I don't think anyone who doesn't have access to Apple source code and production systems could either) that wouldn't preclude Apple harvesting the results of said AI analysis.
In fact, doing the analysis on users' devices would represent a shift of that processing from cloud to edge, representing a significant savings for Apple or anyone else in a similar position.
The use cases we're talking about also don't work on a cloud based analysis, as you can't have text selection block on network uploads (generally slower than downloads), and it would require uploading every image you open to apple which would presumably be a lot of traffic, and an obvious privacy nightmare. It would also break for users who turn on the e2ee everything mode for iCloud.
I guess there is a silver-lining in the premise of AI generated one-time-use games for that sake, but then there is a significant "can a human even do this?" problem to conquer at that point.. and worse the same AI tech is going to be established on the opposite side of the wall trying to defeat the thing.
I think it'll all boil down to some sort of state-license fallback method like "please enter a CC or ID number to continue" -- which is ultimately a defeat of the user, unfortunately.
Or are bots somehow able to do those too?
Just because I emit to other clients does not obligate me to emit to yours, any more than my emission of ads obligates you to accept and render them (but if you don't, or if you choose to ignore my CAPTCHAs, I may choose not to emit to you).
It's also not really that much of a losing battle. Cloudflare will fight the battle for me quite well for free, and even better for a pittance.
Businesses that can't afford the expense will close or adapt, depending.
Maybe fewer hobby projects will be launched.
Which is why my hobby projects will continue to use bot detection and CAPTCHA recognition. Especially since I'm routing through Cloudflare, so that's invisible for 99% of my users and the remaining 1% can just get off Tor if they're tired of solving the captions.
may be the premise is wrong.
Why prevent non-humans from registering/using/viewing?
If website operators don't explicitly introduce micropayments as a captcha alternative, there will be browser plugins that outsource captcha solving to AI for a micropayment, which has the same effect.
Option 2: Using a means of authentication that can't be obtained cheaply at scale by bots, e.g. Twitter accounts, Gmail accounts, government ID, ...
deathbycaptcha does not let anyone simply sign up to work.
That's rough, but you could scale it up I guess. Didn't know that about dbc, thank you.
Websites will very quickly pivot to alternative solutions like payment card verfification, etc.
A lot of bots also use VPNs and Tor so captchas being a pain in the ass is probably working as intended, that way most people won't bother using services like that? This is different from regular internet users, there is no reason to make their life more difficult than necessary.
IPv6, 3G/4G/5G or public Wifi can increase that to about every 10 queries on Google for a CAPTCHA. I guess VPN too increase the probability to get a CAPTCHA.
I think in the same way AI can beat us recognizing unfocused photos.
Although it's not clear to me that the humans all really were humans.
However, I'm not entirely sure what kind of system a superintelligent AI would need to access which would be protected by a captcha.
Consumer devices have a lot of spare CPU and RAM. So a proof-of-work algorithm which consumes those resources for a minute might work?
If it generates $0.01 for the website owner in that minute, maybe that would work?
There are exotic solutions like the 'Idena Network'. But sadly I have to admit the best solution I've seen so far is Sam Altman's Worldcoin. Not that I'm a fan, I still hope we can find something better than scanning everyone's eyeball.
Tor has such a feature for denial of service protection.
https://blog.torproject.org/introducing-proof-of-work-defens...
A benefit of a token is you can recycle previous proof of work by using a small amount of Bitcoin, which could be transferred using Lightning. The value could also be transferred back some amount of time after registration given no bad behavior, allowing for larger sums than a cent, which could provide better protection.
If you only consume resources on the client side, then you hope that an attacker thinks "I won't invest $0.01 of resources just to log in here".
If you also transfer the consumed resources to the server, you get an additional benefit: The server thinks "$0.01 is enough to cover the costs of a fake signup".
And the second benefit is probably even better than the first. The server will never really know how cheaply attackers can access resources. But they probably know how much a fake signup costs them.
The drawback is it gets a lot more complex when using a token, because of the additional state, communication, costs and security.
A one shot proof of work can be very simple, but probably not effective enough, given that mobile users likely do not want to wait what may have to be many minutes and drain their battery.
Freezing a cent or a dollar for days seems like a better option. Might very well be that VISA/MasterCard figures this out before the crypto bros build anything usable. It will be far easier to do without decentralization and would also be great to spy on and control people.
Fucking A HN.
For any Juniors using this site, this is exactly what you don't post. Especially if it's just to cathart cynicism. I assure you, Poe's law guarantees this will find it's way into some PM's or exec's mind somewhere.
If the captcha is to prevent overuse of a free trial, then nobody will operate a lot of devices just to get more free trials if the paid version is cheaper than those devices.
If the use case is to improve democracy, then it gets more complicated.
For more complex cases, not AI but consider the attack in https://www.usenix.org/system/files/conference/woot14/woot14...
As an American, I have a similar experience when I travel across the Atlantic. It's always funny to me when I land in the UK, start using websites I use normally at home, and get cookie verification modals from hell to breakfast.