Bots can complete CAPTCHAs quicker than humans
theregister.com
theregister.com
Does this set off alarm bells for anyone else? Of course the best way to know if a visitor is a human or a bot is to deeply analyze their behavior. But that's at odds with the right of us humans not to be analyzed by every website we visit. What happens if we do reach a standoff where the bots become good enough at mimicking human behavior that the only way to tell us apart is unacceptable and illegal behavioral analysis?
But they just trade one privacy issue for potentially another, depending on your view
It does seem sadly unavoidable. Perhaps the internet has to go full circle and we need real identities if we want to ensure we’re not talking to machines?
[1] https://blog.cloudflare.com/eliminating-captchas-on-iphones-...
I’m can’t begin to theorize how the future will play out if you need a PAT to access most web destinations. Do cloudflare or Apple engineers use Linux machines ever? Surely they do, and either know this is bad or have some plan to make it work?
If that's the case, then I'll be done using the web entirely.
Although arguably most major web players already 100% know who you are - just maybe not your name.
As am I, which is why it's not that big of a deal to take this stance.
> most major web players already 100% know who you are
I have no doubt about this, but there's also a whole internet full of others who I want to remain pseudonymous with. I've been using a handful of online identities for over 30 years now, and have never tied them to my real world identity.
The reasons for avoiding that hold more true now than ever before. Having to tie my online identities to my actual identity is unthinkable.
If you think about it this is no different than showing your ID to get into a bar.
Also once you are sitting on a few dozens phone numbers, you can use them again and again to spam or abuse different services (possibly for sale). It's not like CAPTCHA solutions that you have to do every time.
You'd think so, but no.
I signed up for T-Mobile service early this year with no ID, and paid cash.
The store is so eager to complete the transaction that it keeps a government ID document in a drawer and the sales people whip it out whenever anyone looks queasy about providing information.
I didn't resist giving my information. All I did was pause because I wasn't sure if I brought my ID with me. Even that little hesitation was enough for the clerk to say, "Don't worry about it. I got you covered" and he pulled out the ID.
So I have a T-Mobile account that I can pay for with cash and no ID on file, and someone else's address.
Now, if a government was really interested in me, it could probably pull the security camera video or follow the signal around or whatever. But it turns out that KYC is easily bypassed when the incentives are right.
There are services that allow you to verify for as low as $0.03/activation and their stock is massive and diverse so that's not a solution
Human or bot is not really the problem, spam is the problem, and bots makes spam so cheap that admins can't deal with it. So, bots are banned. Human spammers can still get in, and you can pay people to solve captcha, but humans are more expensive, so there are less of them and moderators can deal with them.
If we had people (or bots) pay a few cents to access a service, it could be enough to keep spam to a manageable level.
The problem is, people don't like to pay, and unlike with phone numbers, the web doesn't have a good microtransaction architecture so behavioral analysis it is.
This is alosing idea out of the gate.
This will only work with mechanisms that let you pay like 0.05 cents. Should be enough to deter bots that practically run for free these days.
Too bad any intermediary will want 0.30 dollars per transaction for the 0.05 cents :)
* note that the word just does an obscene amount of lifting in that sentence.
You can slowdown bots with proof of work. A crypto-miner seems like the only possible payment method that would resist tracking, I think Brave tried something similar. Not sure I like that idea!
But now, when I see a captcha, I hit back, press unsubscribe, and find a new vendor. The work is harder for me than it is for a computer and so I won’t do it.
Accordingly, we can see that proof of work is the opposite of a solution.
The question there is if an acceptable amount of work (e.g. cpu/gpu work) on a mobile phone is a large enough deterrent for a bot.
I let my my cryptoskepticism run a little too freely here. Thanks for making me think harder.
Microtransactions solve the issue of bad bots, and possibly websites monetization. But then do you want to give free pass to search engine crawlers? The big ones will be strong enough to refuse to crawl your site if you don't. The small ones will be financially unable to crawl if you don't. If you allow them all, you're back to step 1. If you allow only one or a few, you basically freeze search engine innovation.
Not to mention credit card fees making sub $1 payments a no-go and crypto being it's own barrel of nightmares.
You don't need one, but if you don't have one, you're at a disadvantage against those who do.
Sophisticated bots are already good enough at this that a variety of behavioral-based bot analysis tools exist and are in semi-widespread use. They're not illegal.
Apple devices already support something like this when connecting to websites behind cloudflare and fastly, and as cloudflare explains this "vastly improves privacy by validating without fingerprinting"[1].
https://blog.cloudflare.com/eliminating-captchas-on-iphones-...
They also 'care' more than actual customers in many cases. Real customer -> "Stop sending me review reminders. It was a comb. Block." Bot -> "Dutifully review all kinds of products. 500 words on the life changing experience of hair brushing with this comb. A+ reviewer."
I find it difficult to believe that the bot networks would not have just immediately rolled every single generative AI advance into their networks (write convincing reviews, generate convincing product examples without buying, beat captchas more reliably, automated screen clicking, human eye scan impersonation). Need to be better than every other group doing paid reviews. Need to be better than actual humans. They might write critical reviews.
Also, lot of sites already doing some behavioral analysis. Popup every time you consider clicking 'leave' on websites lately? "Before you go..."
Probably, but if these are just like, Amazon reviews/etc, they likely violate FTC regulations. Enforcement is lacking, but I'd still be very hesitant to break the law.
What?! I've never seen this, is this proprietary to The Site Formerly Known As Twitter?
Users with cognitive impairments would struggle with this I speculate. (All humans, really)
genuinely, what did you mean by this?
The joke is intended to be funny because "twat" is a vulgar and generally derogatory term, and the author almost but not quite applied it to either a large company or (transitively) to its users.
I hope explaining the joke made it even funnier!
“Twat” is a mild to medium insult (idiot, asshole etc depending on tone)
The fact that bots can solve them -- and solve them fast -- is apparently a well-established fact in the literature. There is a table in the article comparing its (human) participants' solve times to a number of previous studies which examined how fast/accurate bots can be.
The Register (and the New Scientist, which most of this is cribbed from) is looking for a headline, so whatever. But the study's authors say that the "surprising" part is that "solving time and user perception are not always correlated" for human users. Game-based CAPTCHAs with sliders may take longer, but the users in the study still enjoyed them more than image-selection-based ones.
Better to deploy some light measures (tarpitting, RBLs etc.) on entry, then weed out the bad actors once they start acting bad inside the system, no? I mean CAPTCHA for everyone? Come on.
You may not have been around for it, but it's not like everyone was super duper excited to put these things on their web sites. It was something people were dragged into kicking and screaming, and even today there's a lot of those technologies deployed even so.
You are probably underestimating the willingness of bad actors to make efforts to avoid these things. Is your model of a "bad actor" on the web some malicious guy writing a program and running it on his personal laptop from his home connection? Because in 2023, your threat model should be something more like a guy who rents a botnet out with millions of computers of all sorts on it (the difficulty of this rental being somewhat higher than AWS, but only somewhat so, it's not that hard at all really), collaborates with other bad actors to work out how to best bypass filtering, creates websites to do things like CAPTCHA proxying so that humans fill out the CAPTCHAs in return for free porn or something, trades rootkits and other exploits around both for home computers and for compromising web servers for their campaigns (for the URL cred), and so on. You're not up against some guy, you're up against a honed and tuned machine with years of experience, internal division of labor and skillsets, basically an entire parallel predator economy.
Tarpitting and RBLs are not dead, but they became just one layer a long time ago.
So developers just install a Captcha and outsource the problem to Google.
I think the primary way to deal with the problem should be to design services in a way to make them unsuitable for spammers.
I'm guessing part of the answer is (and most likely already implemented in things like reCAPTCHA) is rate limiting and detecting bots when they solve these too quickly.
reCAPTCHA is one of the better captcha because they do a decent amount of browser fingerprinting and their captcha are interactive.
Still there are services for solving them. Fun thing is you only need to pay those services for first ~50K captcha and then you can train your own solver using the data you collected.
Ultimately, captchas only serve to increase the cost of running bots. If what ever your trying to protect is worth more you will fail.
;)
You can tell it because it's not actually Google or Cloudflare installing captchas on websites of third parties. They in fact cannot do it. It's done by the people operating the websites and who desperately need to protect it against abuse, and for whom letting a company track you is not even a hypothetical motive.
Even in the case where you're getting some kind of behavior or reputation verdict from past behavior (and possibly across multiple surfaces), you probably want a progressive set of outcomes rather than just a binary allow or deny. Even if some requests are clearly best just blocked and others should obviously be allowed, there's always going to be a grey area where you're not sure. You need something to do with those requests. Making an arbitrary choice is one option, but pretty harsh on the legit users. A captcha in another.
Sometimes you have options for that gray area that are much better than captchas that you can do, e.g. request a phone number and do an SMS challenge. But that's both expensive and will lead to a massive dropoff for most sites as people won't be willing to give out their phone number to every site.
(Also, the act of solving the puzzle can give you additional signals of whether the request is from a bot or not. Signal collection is kind of the entire point of the slider captchas in the first place.)
You can significantly increase your success rate by adding some "random human" actions:
- wiggle the mouse a bit, don't go in a straight line
- click multiple times in quick succession on a image
- wait a bit before clicking the Submit button after selecting matching images
- match a wrong image and then immediately unmatch it (as if a mistake)
- click randomly on the page around the CAPTCHA
- resize the window a bit, rotate the scroll mouse a bit
And don't do the exact same thing every time, improvise, pick randomly 2-3 actions from the list.
I was surprised how rarely that I have to make more than one submission in spite of intentionally making incorrect selections.
I'm looking forward to Google getting sued after a Waymo tries to make a right on red at a pontoon boat.
That google does this is not really a surprise, since they earn money by letting bots through (bots are then counted as humans and google can bill for ads shown to the bot), but hCaptcha at least advertises the fact that their interest is actually detecting bots.
surely the best option nowadays is to make your own or find an obscure one, and hope it's unusual enough that ready-made software doesn't exist that can easily solve it. then if and when it gets cracked to the degree it's impacting your content, move onto another one
putting in a credit card?
I'm wondering if paid services have problems with bots, or if they mind, since they are being paid.
Can also just use old empty VISA gift cards, ethically sourced from relatives and friends of course.
Can’t remember the last time I clicked reCAPTCHA and didn’t have to do the challenge, come to think of it - I can’t remember a single time, so if it has happened it is very rare, whereas the Cloudflare one always lets me through.
>getting users to pay for gold
Gold... a 4chan pass?
When you get used to it it doesn't seem that hard... Twitter on the other hand is incredibly confusing with the visual one but the audio is easy
they increase the processing/energy cost/set-up time/difficulty to the point where it may no longer be profitable to access that content, but no one really thought a powerful computer with the right software couldn't actually solve them at pace, right?
Profitability. That is what it comes down to. The price of solving these have plummeted in last couple of years.