ReCAPTCHA and the Anonymous Experience
blog.repl.it
blog.repl.it
> anonymous
Pick one. reCAPTCHA is based on tracking users in countless unknown ways, which is why it's basically unusable with any kind of known proxy (i.e. Google usually classifies Tor users or users from 3rd world countries as bots by default). It may be stopping "dark web hackers", exactly because it's unsolvable by anyone Google decides to be a potential threat, and that also includes bots.
Honestly, I feel like everyone would be better off if instead of reCAPTCHA, websites just used some service that automatically blocks proxies. That way you aren't getting Google to track legit users, and still forbidding potential users from joining who would otherwise be blocked by reCAPTCHA. And maybe if Google isn't watching me on your website, perhaps I'll be more inclined to visit you from my actual IP.
Or just use something that is actually solvable by humans, like hCaptcha. Cloudflare did that and it's wonderful (well, compared to reCAPTCHA), I can solve it in a few seconds, unlike reCAPTCHA which I can keep clicking on for countless minutes only to get told that "my network is sending too many requests" or something.
Captchas are in their infancy. We need to move beyond something that involves simply clicking to identify images - we're providing free training for AI, not to mention the fact that it's incredibly inconvenient for blind people.
I think a measure of uniqueness is required based answers to questions that don't necessarily have correct answers. The distance between the question and answer would be the measure. Identifying questions that provide consistent distances for individuals would be the hard thing.
Though if you aren't considered suspicious, image captchas are way less ridiculous, too. The scoring system goes all the way from .1 to .9, but usually you're either in one or the other category based on who knows what (probably mostly ip addresses, but Google doesn't tell us so who knows?).
So in the end the service is mostly just something that arbitrarily classifies users as one or the other, bots can absolutely solve them when classified as "probably human" and they're impossible for humans when classified as "probably bot".
If you want to actually present users a challenge to verify they're human, you're even better off giving them something like captchouli[2], or just a good old scrambled letters on a picture thing (which is what Google uses for their login as well, BTW). hCaptcha is also a great alternative (anything is, really) since it's still mostly about classifying users by making suspicious ones solve challenges. reCAPTCHA is absolutely not, it judges you before the challenge.
And that's fine, some websites want that. It's their choice, I'm not mad if they ban me for not wanting to be tracked, it's probably for the better. What does make me angry is that reCAPTCHA is advertised falsely, that the challenge is a joke and a crime against humanity at once, and that despite that, people seem to believe that it's the challenge that is magically blocking the bots without realizing they're also blocking legitimate users.
My guess about the future path is as follows: your government ID allows you to authenticate to generate a keypair you can use online for services such as Google and Facebook. It would include things such as a nickname, and an e-mail address as a derived public/private keypair you are able to use, including a scope (domain). Then, only such authenticated logins get full access (they are not going to phase this in like a wrecking ball, it would be slowly). Its perhaps a tad dystopian, but it has its pros and cons, and websites like GAB will continue to exist on darknet.
A centralised ID system is one solution for authentication online but it's far from ideal.
A much better system is a decentralized pseudonymous ID system where getting an ID is in some way expensive but still anonymous, e.g. proof of work equivalent to $5 in cloud CPU time. Then legit users can create an ID once and use it indefinitely, but also create as many as they like and replace them as often as they like for a relatively low price.
Meanwhile spammers pay $5 for an ID that only lasts ten seconds before it gets banned and has all its messages retroactively deleted, so they go out of business.
It's a tough nut to crack. It goes under the term proof of individuality in terms of Blockchain.
This is always the claim but I don't really see it. Obviously if you put something on the internet then you disenfranchise anybody without a computer and an internet connection. But who is it that can afford a $50 device and a $10/month internet connection but not a one-time cost of $5 worth of CPU time?
There's always a way around incentive based methods.
If they had that amount of CPU time they would have the opportunity cost of using it to mine cryptocurrency instead, so it's still costing them $5 per ID.
Maybe they are studying security in a country where that is considered illegal hacking.
Maybe not everyone wants their data sold to parties that sell the ability to manipulate them to the highest bidder.
Honestly people that even think about blocking Tor reek of privilege and an assumption that they are above being manipulated by well trained AI fed a lot of their data.
Personally I use Tor on most devices and refuse to use sites that block it.
My only browsing activity that exits from my direct ISP's IP is my online banking VM, and that's only because one site got insistent about hard blocking everything else. These companies don't understand that they're screwing over real customers with this crap.
It's especially galling when my long lived static IP VPS exit [0] gets hassled. It's clear there isn't abusive usage from this IP, since I control it. I've had it for years, so it doesn't have a bad reputation. Yet these sites still want to fuck with me for their snake oil.
[0] Which doesn't win any awards for nym rotation, but at least hides my location and disassociates activity from my direct ISP IP.
> no sites for tech with less than (arbitrary) 10,000 users
> no nudity, exploitation, drugs, copyright infringement or sketchy-content sites
Keep trying however you think best to convince people, but I don't see why this list would be very persuasive. With these rules, it's impossible to find anything that would make someone go "oh, that's something I want to use TOR for" if you're only considering sites that can be easily reached without TOR. In fact, I'd argue that sites on such a list are not really part of the "dark web" at all.
edit: Downvote and be mad if you want, but payments made through known tor/vpn services are not getting through fraud review.
That's even worse than reCAPTCHA. At least as a human using Tor you have some shot at solving a captcha, even though you usually have to try 5 times or move to a different end node.
It would have saved me quite a lot of time if those sites had just flat out said "403 Forbidden".
I'm starting to think that Cloudflare is even worse than the GAFAMs, due to its impact on so many of the other websites :
https://www.gigablast.com/blog.html
(If it can even be separated from them, considering that Cloudflare has received up to $110M from Microsoft, Google and Baidu !!
https://techcrunch.com/2015/09/22/cloudflare-locks-down-110m...
Any disincentive for sites to force users to train AI for Google is a good thing.
Every time I run into it, it makes me hate the site I'm on. Google punishes my crimes of Firefox use and not being logged into Google by giving challenge after challenge - along with those infuriating "slow fade in" images to take bigger bites out of my day.
The world would be better off without ReCAPTCHA.
Guess what, folks, every visit to your website involves some mix of humans and software. Nobody uses the web without a browser, and every browser was written by a human. You aren't entitled to make hair-splitting distinctions and dump the enforcement burden on the public.
Also it forces users to perform unpaid labor for Google.
You don't want them scraping 1000 articles an hour? Just put a limit on viewing 20 articles an hour, no need to differentiate between human and robot.
The boundary between humans and robots will blur over the next few centuries. At some point it will just be a spectrum from all-inorganic to mixed to all-organic beings that roam the planet. We might as well prepare for that future by abolishing chemistryism (discrimination against organic vs. inorganic chemistry of a being) today.
(Yes this sounds stupid, but 200 years ago, abolishing racism sounded stupid too.)
I'm mostly human and I don't store cookies.
0: https://www.google.com/recaptcha/about/#combined-table__tabl...
I might try navigating the web while blocking reCAPTCHA. How limited would my reach be? And what about blocking CloudFlare's solution?
While it's not the perfect captcha either (which I think is impossible), it makes a better tradeoff in terms of UX, price and privacy.
You can look around for "useful" work, similar to how recaptcha was originally about transcription. If you can find some problems of the right difficulty that people want solved (e.g. I dunno, protein folding or something), then the electricity isn't wasted and you might even be able to sell the solutions.
> It's broken
>Tasks that are easy for all humans but difficult for computers may no longer exist.
>Using machine learning or even browser plugins one can solve ReCAPTCHA in under a second. There are even CAPTCHA solving companies that offer thousands of solves for $1.
This is probably a bad argument when your proof of work captcha can be solved for much cheaper. Your site says "Solving it will take a few seconds on a desktop computer", which I'll interpret as 5 seconds. The spot price for a c5a.2xlarge instance (8 thread zen2 CPU) is 21.6 cents/hr. That works out to 0.03 cents per solve, an order of magnitude less than the 0.1 cents per solve for commercial recaptcha solving services. It probably gets even cheaper if you get your compute through non-cloud providers, or through GPUs.
The difficulty can be scaled in a predictable way - it's similar to rate limiting but less all or nothing. We're about to release automatic difficulty scaling per IP, so if many CAPTCHAs are requested/submitted from a single IP the difficulty increases exponentially. Also being able to set the initial difficulty for your usecase and audience is something that should help.
Aside from that there's some more measures on the roadmap: using lists of known-to-be-datacenter IPs, and reputation lists such as [2], as hints to increase the difficulty.
But you're right - it will still be affordable to attack any CAPTCHA, FriendlyCaptcha is no exception. Proof of work approaches have downsides too.
The main ideas behind FriendlyCaptcha vs ReCAPTCHA:
* The user experience is superior. It can happen in the background while the user is doing something else. There is no labeling task.
* We don't have any incentive to collect user data or track users (GDPR compliant, no tracking cookies etc)
* It's as easy to add as ReCAPTCHA to your website. The API is a near copy of ReCAPTCHA's API. You can host the JS code yourself, or even bundle it. With recaptcha it must be third party.
* It works in any browser less than 8 years old (IE>=11), although of course it's much slower in old browsers that don't support WebAssembly.
* It doesn't have inherent accessibility problems (poor eyesight/hearing doesn't matter).
* Open source at its core [3], the SaaS wrapper is not open source.
[1]: https://news.ycombinator.com/item?id=24921288 [2]: https://www.stopforumspam.com/ [3]: https://github.com/friendlycaptcha/
All those upsides are not compelling if it doesn't effectively stop abuse.
There are ASICs for crunching blake2b designed for mining siacoin. One ~$2000 card [2] can do ~4 trillion hashes, or 30 million captcha solves, per second
Right now we use standard blake2b as nobody has repurposed a miner to solve hashes for spamming yet.
The thing is, a determined spammer will be able to attack any CAPTCHA - even in labeling tasks there is always the fallback to human-in-the-loop which is cheap at scale (or even free if these are MITM'd users..).
Any (new) CAPTCHA system will have flaws and break in some way at scale, we're open to ideas and of course will try to address any (future) concerns. We are trying to provide a viable alternative to ReCAPTCHA that respects the user - and we will iterate on these problems as we go. Without some new thinking and openness to new approaches we'll be stuck with ReCAPTCHA.
Small nit: the difficulty is set to require around 2.5 million hashes, not 115 thousand. Your point still stands though.
A trifecta of doing evil from Google (well at least two out of three, 1 should cover for 2). And I say that as a relatively pro-capitalist with no problem charging money for services but I'm also pro-privacy and don't like training their AI models for free with them charging the hosts on top of it.
I had tried it some time before but IIRC it was either invite-only or enterprise-only or had some "size" requirements.
Just saw that it's available for all. Thank you!
Privacypass itself is a privacy violation and not available on most browsers and I don't want to spend an half hour a day doing free labor training someone else's AI to just use the internet anonymously.
what? https://support.cloudflare.com/hc/en-us/articles/11500199265...
It has addons for firefox and chrome, which makes up 90+% (by market share) of the browsers out there.
This add-on needs to:
Access browser tabs
Access browser activity during navigation
Access your data for all websitesBusiness tier (we needed custom certs) + advanced ssl is like $220/month and you get a WAF, a world class cdn, and DDOS protection.
I was thinking about looking for recaptcha alternatives since October and until now I wasn't aware of hcaptcha.
If by superset you mean it's as anyone as ReCaptcha plus 10 times more, I concur.
Recaptcha is annoying but at least it gives me a relief short a while after solving it. hCaptcha is dumb as rock and challenges me every 10 minutes even when browsing the same fucking website.
[1] https://support.google.com/recaptcha/answer/6223828?hl=en
[2] https://user-images.githubusercontent.com/20207154/29577170-...
I actually can't think of another user interaction I've had with a computer that had less respect for me as a person.
So... a web/company has to pay Google for the service... while the company's users/visitors will still do free AI training for Google? If that is how it works, that's some bold business plan (money on top of free labor).
The same is true for recaptcha, I usually just don't use a site if they make me use it. I have stopped donating to charities because they use recaptcha and would never buy from an online store that uses it. The same as I would never go to a bar where you are searched on the way in, or buy a pizza from a place with a central call center that makes me wait to make an order. If a business wants to treat me like shit, I just wont use it.
If you don't want to deal with cryptocurrency then this could be simply a provable burn - purely an IT thing with no accounting team involved.
Monero and ZCash supposedly offer real anonymity, but IIRC all implementations so far have been found vulnerable to (at least partial) deanonymization.
Crypto currencies may have good uses, but so far, anonymity is not one of them - most definitely not for the masses (and hard-to-impossible for the very disciplined and knowledgeable pros)
VPN or proxy users are not really anonymous anyway and would certainly welcome an option to pay $0.1 in cryptocurrency instead of fighting infuriating CAPTCHA-s all the time.
If I would have to pay 10c to visit a website I would like it to "burn up". Most websites I visit I don't support in any sense.
Unless you actively try to hide your tracks, every single transaction you make is related. You may pay your friend at work $5 back for coffee using Bitcoin -- and at that second -- since he knows your wallet id -- he can check the blockchain for every transaction your wallet has ever participated in - every website you paid for (as a captcha, as a registration fee for that totally-legal-but-morally-questionable site, the money you contributed to support/oppose a political cause, etc.)
This is NOT paranoia. People who aren't the NSA are constantly analyzing and making public identities related to wallets and transactions they made. Psudonimity is not anonymity - it's one step away from being an identity; and the fact that the blockchain is public makes all history public for an identity once that step was taken.
Take cash from ATM. Put in wallet. Use cash in wallet to pay for drugs. Use cash from same ATM from same wallet to donate to church. Only NSA/FBI has the means to track this, and they have to work for it.
Put money in bitcoin wallet. Use wallet to pay for drugs. Use same wallet to donate to church. Now church knows you paid for drugs, and your dealer knows which church you contribute to.
Brandan Eich was forced to resign because of his political beliefs, as evidenced by his monetary contribution to some political cause[0]. If pseudonymous blockchain payments become mainstream, such events are going to become an everyday occurrence. For some things, that's a net positive for society (you want to know who the hypocrite politicians are). For some things, it's a net negative (losing privacy for individua).
[0] I'm tryting to word this as neutrally as possible.
Law enforcement is increasingly targeting these mixers, there was a recent case were one operator was criminally charged[0] for not having a "money service license". I am willing to bet that, as part of a plea deal, all past logs they keep (which is likely all of them to day one) will be provided to FinCEN, which may or may not make it public at some point.
FinCEN is a few notches below the NSA. It's not your cousin Jerry, true, but it's a lot less than "only the NSA". And the history is immutable - it's possible, even likely, that 20 years from now, many mixer logs will have become public enough. Anything on the blockchain stays there for posterity.
[0] https://www.fincen.gov/news/news-releases/first-bitcoin-mixe...
Although I like services like hCaptcha a lot more myself, it may be possible that this notably bothers users and decreases conversion rates.
So either you're in a country that Google has decided is "good" (i'm in Germany), or you're not blocking Google's tracking as well as you think.
Without disclosing my location, let's just say I'm in a much less powerful (and hence, I would imagine, less reputable) country than Germany. I doubt it is in the "good" set, but who knows.
I also doubt I'm not blocking it well enough. I'm running Firefox with CanvasBlocker, uBlock Origin (with strict rulesets) and uMatrix. I also turn on restrictFingerprinting.
(I'm using uMatrix too, but this doesn't solve the issue that if I don't allow reCAPTCHA, I'm stuck on the first step of account creation…)
https://dashboard.hcaptcha.com/signup?type=accessibility
As a user that can't answer the challenges, you have to register with hCaptcha in advance of using the website you want to.
This lets you bypass the verification checks for a while.
The discoverability of this is poor
They've since seemed to put a lot of work into it. To the point where I don't notice it anymore.
They're iterating, and that's what I want to see.
I rarely get contacted so I wouldn't have to pay any time soon, but I want to remove all third party services from my site anyway and I have to find a way to prevent spam that would nevertheless be less annoying for someone who would actually want to contact me.
Can one still use v3 and not pay for Enterprise?
Or does Google only count successful human detections towards the 1000000-request limit?