hCaptcha now runs on fifteen percent of the internet
hcaptcha.com
hcaptcha.com
I realize anything connected to the internet will be subject to automated abuse, and it's impossible to run some types of services without taking some steps to defend against it, but it seems to me there's usually a way to handle that without invading the user's privacy or wasting their time. The exact details will vary based on the type of service, of course.
One particularly egregious misuse of captcha in a service I use presents one after I enter a correct username and password. An incorrect login says so without presenting a captcha. The potential reward for an attacker who successfully gains access to an account is high, so it seems almost certain anyone running a targeted attack would defeat this by handing it off to a human upon detecting that they had a good account.
As much as I agree with your dislike of captchas, I don't think this is true at scale (unless universal online identities existed, which could and should include anonymous identifiers by design). When you need to accept information from anonymous users (comments, votes, forms, registrations), there's no way to not invade users privacy and not waste their time, unless you are manually filtering / moderating all the input data, in which case you can't really say it scales. You might say emails can solve the problem. Well, they don't really solve the problem against dedicated attackers / spammers, and they do invade privacy for the average user. You can use statistical approaches to try to reduce privacy invasion or others, but I don't know of anything that really solves the problem without manual identity verification at some point.
Also with multiple requests from the same IP in a short timespan, the difficulty increases.
There are downsides to to any captcha, but in my opinion make a much better tradeoff. Accessibility and privacy are respected, and there are no annoying tasks.
For someone who has little expertise in this specific field, how are you calculating this?
But seriously I like the idea, although it seems trivial for someone to attack a protected site by exhausting its subscription level? Are there any protections against that?
As someone working in the field, I also doubt your claim "will prevent 99.9% of spam" is based on real data. Modern headless browser spambots are not deterred by this kind of approach.
(Edit: looks like the poster admitted this number was entirely made up later in the thread.)
Maybe not everyone but a lot of people use captcha services to prevent automation from being used to extract/insert data. I know as a developer that there is always a chance of bypassing this, even with Google's reCaptcha, but your service seems to make this trivial, so many won't even go beyond your demo.
I'm using latest Firefox on GNU/Linux. Admittedly I've got a lot stuff blocking all sorts of things, and I'm not really sure what's kicking to block background workers, but I'm glad it's blocked. Anyway, after disabled literally all blocking tools that I have, it still refuses to load.
It's reaching a point where encapsulating a VPN with anti-captcha is something I'd pay for.
Yes but no. Anonymized identifiers can be deanonymized. They should utilize zero-knowledge proofs in such a way that they can prove "yes, I have an identity verified by entity X (and Y and Z) (based on passport/phone number/...)", without disclosing any of those details.
It could, optionally, yield an identifier unique to each requester and unlinkable to others unless an explicit proof of the link is provided. Though if this is included, there has to be some mechanism to avoid huge ad networks sharing the same "requester entity".
This is a solved problem. All that's left is politics, implementation and alignment.
I have actually discussed the concept in the past [0], and I exchanged some emails with the guy in that thread to talk more about technical details. We all seem to agree that design and political will are the problems, not technology.
What I was basically saying in the comments, in general terms, is that you might have one primary identifier, and then somehow you can get more identifiers that are tied to your main one, but that might have different expiration periods, might grant access to different levels of information about you, and might be limited to a certain number for each service you use. Of course, there are quite a few ways to implement such a system. And that's precisely why I'm more focused on the design, usability and characteristics than the underlying technical implementation; I think the best we can do if we ever want to see this happen is to spread the idea in terms that anyone can understand [1]. I mean, I'm interested in the technical details too, so I'm just complementing and contextualizing a bit here.
[0] https://news.ycombinator.com/item?id=22180120
[1] ...or discuss more the idea among those that are interested and setup a demo website to make it easier to spread the word, even if there's no actual implementation behind it and it's just a mock-up. I'm quite busy at the moment, but I'll definitely do something along those lines when I have some time.
Actually, I would prefer: "Yey, this one-use temporary ID is tied to an identity which is known to behave on public websites." Or "... is tied to an identity which is known to be knowledgeable on topic X". Or whatever information is needed at the time.
Next time a new ID will be generated and the identity provider will vouch for it. No passport or phone number should be required.
Edit: fixed spelling.
I agree, there should be better ways to do anti-abuse. Yet I find myself coming up empty when I try to find better options for the common scenario where people would really rather invest deeply in their service than in anti-abuse.
I would love to hear some ideas about how to solve this nasty general problem while also respecting user time and privacy. Unfortunately, I've found that entirely too often the vague sense that there must be a better way fails to translate into substantive better way.
The number of things that are "wrong" with reCatcha etc, have been mentioned on here ad nauseam. In fact, I'll quote myself from another debate on the subject, a while back:
>1: It's never made clear exactly what you're supposed to click on. For example. If I'm told to click on "traffic lights" does that mean just the lights?... or the poles as well?... and what about a square that only has a tiny bit in it? Does that count too, or is it only squares which are mostly filled by the object in question?
>2: They make no concession to non-US English speakers. I've been asked to identify things before, where I had to guess what the word means because the same thing is called something completely different in UK English.
>The only thing that approaches the level of rage that reCaptchas instil in me are those captchas where you've got to transcribe what's in a photo of some letters & numbers and where they NEVER fecking tell you whether it's case sensitive or not, or where they use identical characters for zero and letter O, one and letter I, etc. >One particularly egregious misuse of captcha in a service I use presents one after I enter a correct username and password
Is it eBay by any chance?That recently started randomly showing reCaptchas to me when I'm already logged in and have been using the site for some time. When this happens, it descends into a never-ending cycle of more login screens and then more reCaptchas.
But thankfully eBay have taken note of the dozens of complaints about this on their user forums, dating back to 2018 and rushed their best people in to fix it.
[That last sentence was dripping with sarcasm, in case anyone unfamiliar with the company thought eBay ever took any notice whatsoever of their users' concerns]
I'm not a violent person at all. But if I ever meet the person who spawned reCaptcha and all its equally annoying clones, which are a pox on the internet, I won't be responsible for my own actions.
Contrast that with today's form of reCaptcha where you identify stop signs/crosswalks/et c. for Google's benefit, but at the same time you're also improving...oh, wait, Google again. It almost seems like forced labor, in a sense.
It is additionally resource-theft, when recaptcha-protected sites are used for business purposes. You are stealing valuable business time (possibly very valuable business time, if the person in question is a high-paid role like a CEO or surgeon) to power your pet "spot the crosswalk" project.
>I'm not a violent person at all. But if I ever meet the person who spawned reCaptcha and all its equally annoying clones, which are a pox on the internet, I won't be responsible for my own actions.
be careful, else they start sending you dead horses head and planting gpses in your cars https://www.justice.gov/usao-ma/pr/two-former-ebay-executive...
What a bizarre world we've made for ourselves.
I simply refuse to waste my time and drive up my blood pressure by doing unpaid training work for Google's AI, in order to visit some crappy website. I really wish more people would start boycotting any site which uses reCaptcha [or its derivatives], so we could get rid of this blight on the internet.
I've spotted this new hCaptcha junk show up recently on a couple of sites I used to frequent. I don't visit those sites any more. So well done webmasters. Apparently annoying the shit out of visitors to your site tends to drive them away. Who'da thunk it?!
One particularly egregious misuse of captcha in a
service I use presents one after I enter a correct
username and password.
That's nothing.eBay will CAPTCHA me after I enter my e-mail address, and then again after I enter my password too. Every time. And I'll be damned if I don't "fail" this CAPTCHA at least once a week, with it telling me to try again.
Come on, there are only so many mountains/hills, taxis, traffic lights, bicycles, and cross-walks I can look at before I go cross-eyed.
They even have the nerve to suggest that I can avoid this by using the latest version of my browser (Firefox), which I already am and always do.
Then it may surprise you to know that simply preventing automation makes many types of account takeover attacks infeasible in practice. It won't mitigate the attack if you are personally a high value, named target. But most account takeover attacks operate en masse and are coordinated after large security breaches, so having to hand over accounts to a human operator as part of the auth loop would make the campaign uneconomical. It also introduces another step at which an attack can be logged, recognized, fingerprinted and stopped by an incident response team.
This is something your security team would probably gladly tell you about if you asked them. There's also a bunch of talks about this presented at conferences like Blackhat, DEFCON, USENIX, etc.
Stated in another way: not all potential rewards for successful account takeover are high. The modal account in the modal campaign is low value, which is made up for by volume and particular purpose of accessing accounts. If you model these campaigns economically, you can eliminate entire classes of "low margin, high volume" attacks simply by introducing friction that mitigates automation.
Then there is a natural cost-benefit tradeoff as to how much friction is allowable on a per-user basis to prevent the most common types of account takeover attacks.
I run a problem validation community platform. Couple of days back an individual launched automated spam/DDOS attack by commenting an abusive, demoralising text on every single thread by creating different users.
Fortunately, I had systems in place to identify and mitigate it with Cloudflare. So, in this case even genuine users would have received captcha. I found out soon enough who the attacker was from the firewall, he had earlier created an account with his own name and was using the same IP to attack, after I blocked his IP he tried with couple of other IP addresses incl. Tor; but stopped with his activity after couple of hours.
I generally don't like re-captcha because it takes cultural background for granted(e.g. 'Pie' is not a common food worldwide), Accessibility as a disabled person myself and has no mitigation for captcha-solving farms.
But in nuisance cases like the one I detailed above, captcha is the easiest method available en masse.
Or wget to save a set of pages for later.
I understand protecting commenting with captcha, or contact forms. But captcha on regular read-only access to public web pages in the style of Cloudflare is a bit ridiculous.
One thing contact forms should have is a static indication there's a captcha in use. I've filled all too many forms that just sent my written text to void, because I block some domains.
Sadly, there still doesn't seem to be much in the way of micropayment infrastructure.
If they'd like me to pay for access, they should return HTTP 402 Payment Required instead of letting me download the page for free. Perhaps they could also rate limit the network connection to prevent denial of service. Why straight up block automated user agents though? That sucks.
Anti Money Laundering regulation killed it: KYC doesn't scale down to micropayment levels.
If you want to fix the web, you have to roll back the AML/KYC insanity. Until that happens, the web will stay broken, because paying with attention (ads) is magically exempt from the AML/KYC insanity, whereas paying with money or anything money-equivalent (fungible and transferable) is not.
Plus, it's not just benign read-only scrapers. Have you looked at the spam folder of your email recently? That's what every comment section and user bio and god knows what else would look like if you just blindly allow all automated traffic.
They used to be completely local and even some DIY solutions, evolved to signature updates, but eventually the attacks grew so advanced that only online services could be updated and aggressive enough, which is of course how gmail took over the internet with near perfect spam filter (when was the last time you checked a gmail spam folder).
The last generation of local spam filters were pretty good though. Anyone remember Eudora and Spamnix?
I just use bogofilter, and it worked almost perfectly from the start, just because I saved years upon years of SPAM and HAM. 10's of thousands of messages each.
It got slightly worse over years, because I incrementally only train it on new SPAM but not on new HAM, because of laziness.
People probably have HAM archives, but don't usually save their SPAM, to be able to start using Bayesian spam filters right away with great results.
Personally I find it much better than whatever Google uses. I don't even bother with SMTP level domain/IP blacklists, or reverse IP/domain checks anymore. All mail is just passed right to the mailbox and is then pre-filtered by a bogofilter to SPAM folder that I check once weekly, and barely find any HAM there. I receive about 500k mails a year.
What if the majority internet usage is non-interactive, from so-called "bots", what we may refer to as "automated use". Google and Facebook, among others, rely on the use of automation and "bots". The non-interactive clients ("bots") being used by these companies are not asked to solve captchas. (In turn, after collecting data from public sources, these websites attempt to prohibit the use of automation by their users wishing to access it. What is interesting is that neither company provides any definition of "automated" nor any clearly stated limits on the speed at which a user may access resources or the quantity of resources they may access in a stated time period. One might be apt to find such limits associated with an "API".)
In 2013 an Incapsula report suggested that the majority of internet usage is in fact automated and not "malicious"^1 -- what if public information sources on the internet catered to the use of automation rather than trying to limit such use, e.g., with speed bumps^2 like "captchas". What if servers treated all clients equally, instead of having data forcibly collected by a few large clients that receive preferential treatment, then siloed and protected from "automation". What effects would this have on "centralisation" and levelling the playing field.
"Do not ask for permission, ask for forgiveness." What does it really mean when applied to the internet. Perhaps it means there is an endemic lack of clarity about "the rules". Prohibiting "automation" is far too vague and in many cases it makes no sense. The growth of computers and the internet is the growth of automation. Both servers and clients may have concerns about resource utilisation. Websites do not ask for permission when they decide to use large amounts of the user's computer resources.
Consider that a Google could not exist without being "given permission" to use automation. Does the GoogleBot have to solve captchas. No automation means no company such as this could exist. How useful would the web be without anyone being able to use automation to create an index. Based on the HN comments about web search I have read over the years, I would guess that for many commenters, it means the usefulness of the web would be dramatically reduced.
Imagine an automation-friendly internet. The truth is, I think (the data shows) we already have one, except we are in denial that "the rules" actually allow it. An early metaphor for internet and web use was "surfing". It may be that those who are constantly fighting against automation are fighting against the waves instead of riding them. Time will tell. It stands to reason, IMO, that every internet user, whether a server or a client, should be expected to use automation.
1. https://www.incapsula.com/blog/bot-traffic-report-2013.html
2. An early metaphor for the internet was a "superhighway". Speed bumps would seem out of place on a superhighway.
If you show a captcha after a failed password, you need to show a one after a correct password as well. Otherwise you leak information. You can have other solutions, e.g. in a login flow that splits the username and password entry, it's advantageous to put the captcha between those two steps. But even in those solutions the display of the captcha must be independent of password correctness.
Presumably, if the person has entered the right username and password they're going to get access to the service at which point they'll know they entered the right one. What information exactly is leaked here?
The information the attacker is looking for is the validity of the password. If you want to use a captcha to protect against this, the outcome must be the same whether the password is valid or not. Because if you only show the captcha for failed logins, the attacker can find out that the password was incorrect without solving a captctha, which by symmetry means they can also find out if it's correct without solving one.
> An incorrect login says so without presenting a captcha.
I inspected the source code of Google's reCaptcha offering and was disgusted at how many bits of information they were collecting. They also seem to be fingerprinting users so they can't keep registering new accounts on a platform, locking out anonymous users who are usually the best types of users on the platform, as IMHO anonymous voices are (usually) the best voices, or at least the more interesting of voices.
Google's reCaptcha code seemed to be very keen on knowing my 'cadence' or the way I used my mouse and how quickly (or how slow) I completed the captcha. It also looked at things like timezone, screen resolution, battery charge level etc So they could determine if it was 'you' who was using the captcha, soon after, in a separate session (even on a different device!)
I'd bet a good amount that they store that along with all the other personally identifying info they have on you (and google of course has a massive amount of that); which is basically why after a single reCAPTCHA solve, you wont see them prompt you again for ages - they know who you are.
I almost want to just add a "DeathByCaptcha" extension to handle these for me and pay a few cents for every page I visit, lol
It doesn't necessarily have to if Google supported privacy pass like hcaptcha does. The problem is that they don't.
If only. If the same site has reCaptcha across more than one page, within mere minutes of having to slog through multiple screens of one, I can guarantee I'll be doing it again.
And I'm never sure if Google has served me either a very long sequence of reCaptchas, or whether they've decided I'm not a person and are serving me an infinite reCaptcha.
A GPDR request naming an IP address should allow them to provide those scores.
If not, it’s easily demonstrable that they are storing and using information that they’re not including in a GPDR response, and they deserve their multi-billion dollar fine.
Also, ReCaptcha’s behavior is obviously anticompetitive, and also using Google’s dominant positions in some markets to establish dominance in unrelated markets.
This is anti-trust lawyer candy.
hCaptcha is not easy as is being claimed here. I have lost a lot of time and been blocked from much content due to hCaptcha.
It won’t solve the privacy issues but at least you’re not working on google’s training set anymore and captchas are automatically solved for you.
I forgot my password to one site and tried about 2 or 3 different passwords and in-between each it asked me to do about 7 or 8 of those labelling exercises. I finally just gave up and left the site.
Not only that, but the labelling exercises weren't clear. It wanted me to label a "公交車" which means more like a public city bus and there were also school buses which would normally not be called that in Chinese so I didn't label them but Google thought they were part of that class, and wouldn't let me proceed without me labelling them, and furthermore, punished me with more "hard" exercises like that. I guess they are trying to turn me into a stupid bot.
Good luck to whatever self driving car they are training using this data.
It just seems there lacks a way to make it happen beyond the hacker community.
For sure enable javascript to get the fancy stuff, but mostly we just want to read the text, view the picture and see the video.
I'm curious what kind of data may exist on the experience of switching for larger providers; do the users like it? how much more/less time do they spend solving? do they care, let alone even notice that it's not Google's ReCAPTCHA?
Regardless, as ReCAPTCHA is not only terribly annoying but also built for surveillance from the ground up, I still view this as a good improvement.
> Worth noting that this title is primarily due to Cloudflare having switched to them from ReCAPTCHA, and Cloudflare is... well, relatively popular, to say the least.
That's definitely a part of it, but we also have a number of other large sites and services that use hCaptcha to protect against bots, and more that get added every day because of our more advanced bot detection special sauce.
> I'm curious what kind of data may exist on the experience of switching for larger providers; do the users like it? how much more/less time do they spend solving? do they care, let alone even notice that it's not Google's ReCAPTCHA?
From what we've seen, the integration process is generally smooth, especially if you're a previous reCAPTCHA user, since we keep the interface and workflow largely the same.
Solving is roughly the same although we have a number of other protections that irritate bot maintainers and get activated when we detect them.
Not sure if the majority of people are aware of the change, I'm sure some technically savvy people pick up on it more than not.
> Regardless, as ReCAPTCHA is not only terribly annoying but also built for surveillance from the ground up, I still view this as a good improvement.
That's actually one the top reasons we've had a lot of customers come over to us; we put a heavy emphasis on user privacy / security, including adopting/supporting privacy-preserving protocols (PrivacyPass, Tor), and minimal retention of data (see our data privacy policy on our site).
Please do better. You're blocking off a non-trivial amount of the Internet to blind users. You will eventually be sued for this.
Most vision-impaired users have no issue in our testing, and it is a much more accessible option than audio challenges, which discriminate against those with auditory processing impairments.
(disclosure: work there.)
However, I checked secondary markets where you can pay a human to solve a captcha.
It takes a professional captcha solver 70 seconds to solve an hCaptcha but only 15-20 seconds to solve a reCaptcha. Is that typical? That seems horrible.
The market rate for a captcha solution is 1-3 cents, which is clearly worth it, until you think of the ethics of paying someone slave wages so you can browse the internet slowly, but at least without breaking concentration.
Have you considered a more ethical approach, like micropayments that go to charity or something?
This is completely anecdotal (and seems antithetical to the typical HN response to hCaptcha vs ReCAPTCHA), but I feel like I end up spending at least twice as much time trying to solve hCaptchas successfully because they have a lot less consistency in the objects you're searching for. I always have to zoom in to the modal and carefully search through each image, which invariably breaks whatever flow I'm in (moreso than other captchas).
For example, here's a screenshot from the hCaptcha website's "try it out" section [1] -- I barely recognized either boat in image #1 because it was so small. I missed image #3 because I didn't realize it was a huge cruise-esque boat (so big you can't even see any water) and I spent a good amount of time deliberating on #4 because, well, it looks like a car + windshield but... on the water? If it's a boat, I can't really tell, but I marked it as one solely because of the water in the background. Not sure if it was right or not.
It also seems to occasionally provide "find all the X" challenges without there actually being any X, which feels super cognitively weird ("am I just not seeing it?!").
I'd say ReCAPTCHA's main problem is deciding whether mostly-consistent objects being partially in-frame is enough to "count", whereas hCaptcha's main problem is actually recognizing the widely-varying objects in the frame. I think the former is a little more frustrating when you get something wrong, but the latter is mentally "harder" and takes more time on average, for me at least.
[1] https://i.imgur.com/uyqvs5u.png from https://www.hcaptcha.com/
... If there was an on-premise captcha implementation that actually worked, that would be great.
We are working through the IETF and directly with browser makers to support provably private options like Privacy Pass, and are currently the only CAPTCHA service to support this.
Similarly, on the enterprise side we offer various technical options to let our enterprise customers guarantee exactly what data we can and cannot see.
(disclaimer: work there, comments not official, etc.)
https://community.cloudflare.com/t/stop-using-hcaptcha/15896...
For a change that affects "15% of the internet" this seems like very little negative feedback in a period of 8 months.
> Most people do the convenience from Google CAPTCHA, although they sell some kind of info, but they won’t hurt you
I can't even...this is the Cloudflare forum wow.
I've personally had a few hiccups with hCaptcha quite some time back as I "wasn't sure what I was looking for" and consistently fail on VPNs. But in recent months these there's definitely been substantial improvement , and needless to say I hope to see hCaptcha be the majority provider
Although having said that, maybe I am hitting it and that I've been unaware and uninterrogated is high praise! Hm.
Absolutely. Having to solving only one captcha every few days beats solving 5 or 6 on each page visit. hcaptcha supports privacy pass but Recaptcha doesn't.
I really wish we could find something relatively foolproof that didn't rely heavily on tracking or really good vision.
I wish people would just face up to the reality that challenge-based CAPTHCA techniques have failed, and stop using them.
Also top-notch customer support. The CEO was personally in the slack channel helping us. Highly recommended.
Interesting didn't realize this was a thing hcaptcha did[0]. It's basically recaptcha in terms of tracking which sites you visit then, no?
Why they get on these IP lists is I think because it's a general consumer ISP and probably a lot of people get bot nets on there.
Majority of people complaining about captcha need to look at their system first. Of course any detection system has false positive, but the false positive rate is not double digit percentage in vast majority of cases.
Captchas filter out the 90% bulk of automated abuse.
Btw, web scraping is on the nearly harmless side of abuse.
Naturally if I was logged into my google account I wouldn't have much of an issue, because I would be feeding the surveillance machine.
It's possible that I just got unlucky (one of my recent experiences was a site that didn't let me in even after solving it, which really soured me), but I feel like the main reason it's hated less is because people haven't seen it as much yet.
Edit: TIL how offputting a single bad experience can be. From going through my HN history, I found out that this terrible experience was 6 months ago.
By comparison, I've been asked to complete recaptcha exactly zero times day to day (perhaps I'm lucky?). The last time must have been many months ago.
It's hard to speak to the specific UX qualities of the captcha itself, but I find recaptcha generally less difficult. But a captcha that I don't need to complete always wins out over one that I do.
I'm comfortable with the amount of data I'm exposing online. And where I'm at, hcaptcha is not better. And even if recaptcha prompted at the same rate, I'd still prefer recaptcha over hcaptcha.
Literally every time I'm in a situation where I'm required to use a captcha to access a site it is impossible to successfully solve the captcha in any sane amount of time.
This happens both with google and cloudflare.
Tbh. if they don't trust my connection can't they just tell me so instead of pretending to provide a "I'm not a robot" test which is practically (close to) unsolvable???
(Note that this post only refers to captchars guarden the access of an site if they somehow don't trust your connection, not "I'm not a robot captures" on forms or similar).
We collect the following categories of information:
Information that can be used to identify or contact an individual ("Personal Information"), such as name, email address, and country.... We may also verify the identity of our Integrators and Customers by comparing personal information against third party databases or official legal documents.
Information collected automatically as a result of an Integrator’s or Customer’s use of the our Sites or the Services ("Analytics Information"), such as IP addresses, browser type, Internet service provider, platform type, device type, operating system, date and time stamp of access, and other similar information. Some Analytics Information is collected on our behalf by third parties we engage for that purpose, and some Analytics Information is collected through a variety of tracking technologies, including cookies
In the preceding 12 months, we have shared the following categories of information with third parties for a business purpose:
Identifiers. A real name, unique personal identifier, online identifier, Internet Protocol address, email address, account name, or other similar identifiers. Shared with Service Providers
Personal information categories listed in the California Customer Records statute (Cal. Civ. Code § 1798.80(e)). A name, credit card number, debit card number, or any other financial information. Shared with Service Providers
Commercial information. Records of products or services purchased, obtained, or considered. Shared with Service Providers
Internet or other electronic network activity. Browsing history, information on a consumer's interaction with an internet website, application, or advertisement. Shared with Service Providers
Note: Fraud risk associated with an individual IP address may be shared with an Integrator upon request.
Can you link to it on their website?
"Not Applicable to Third Party Websites. Please note that this Privacy Policy does not apply to any website, offering, product or service of any third party, even if it links to our Site or incorporates the Service – please refer to the applicable privacy policies before deciding to provide any information to third parties."
End user data is governed by the Data Processing Agreement, linked here: https://www.hcaptcha.com/terms
We looked at hCaptcha, and the feedback we got was that their approach to accessibility is simply unacceptable.
If you can't solve the challenges, you have to sign up on their website, in advance, and provide them with your email address.
https://www.hcaptcha.com/accessibility
We just couldn't justify that sort of privacy imposition.
Personally, I think all Captcha needs to go.
Why can’t they do something like a reverse SSL where we have to authenticate ourselves as humans?
For example if I have an Apple account on my Apple devices, why can’t they figure out a way to authenticate me as a human from that information?
This doesn’t work for all scenarios (eg throwaway accounts), but it could work for the majority?
I run several long-lived (decades) sites with mostly static content but also some dynamic pages and a commenting mechanism. Here's what I do to prevent abuse: nothing. It's fine.
If you run a massively popular site or something politically controversial then you'll be targeted for abuse. If you're specifically targeted, I don't know how much captchas will help.
For the rest of the 99.9% of sites, just stop it. You don't need it.
This is very outdated intuition. Fresh IP addresses cost peanuts.
For example, your solution still allows an attacker to run a 50k item /login combolist against one of your users with $5 of botnet time, each IP address trying a single uname/pass combo.
Here you pay $18/GB to multiplex your abuse (cred stuffing being classic non-volumetric abuse example) across 72 million residential IP addresses. https://luminati.io/
If you're worried about password with between other sites that have leaked, then the real answer is to generate something like a username for each user that they won't be able to share between sites. In fact with the prevalence of password managers, generating users' passwords for them might just be the better approach these days. And just fall back to email auth every time if they don't want to store it.
Duct taping your broken system by throwing up an annoyance for every user who doesn't want to be tracked is not the way.
It's now supported in all major browsers but platform authenticators are likely not supported on all OS yet.
Your Apple account on your Apple device doesn't stop you from unwittingly being part of a botnet, for example.
Genuine question - I don’t know much about this field
I sincerely hope they, along with all other companies providing captcha services, go bankrupt.
I loathe captchas, especially Googles who seems to punish my use of privacy extensions (ie ublock origin, Ghostery Lite).
Captcha is a ok tool when you have valid reasons to assume the user is a bot (multiple failed logins, unusual traffic, password resets etc.). Used as a default it only antagonizes users.
>Presented Challenges:
>Comparison - Select all images that match query
>Bounding Box - Define bounding area for objects
>Categorization - Identify the corresponding labels
>..and other simple tasks.
No, hCaptcha, no way am I going to train your neural networks for free, so please join Google on your journey to hell.
If there's value in it, it sounds like a spammer could train an hCaptcha-defeating bot via hCaptcha.
And by a corollary, since they haven't started labeling different kind of data, it's clear that either they no longer need any kind of labels at all, or this is actually not a cost-effective way of doing it.
Can anyone talk about their experience running a (large) service with a spam problem where adding a captcha helped? How about still battling bots/spam/abuse despite having a captcha?
Supposedly these services have humans solving these. I remember hearing about a bypass in use where the attacker would pass-through captchas and present them to users on their own pirate/torrent/porn/etc sites, and then when the user solved it, they'd get at the content and the spammer would do whatever they were doing on the original site. I wonder if that technique is still in use, or if there are people specifically sitting there solving captchas all day being paid fractional pennies per solve?
It's a real shame that because of this "safety feature" the web is losing is programability. You can't even curl many pages these days let alone write some programs that connect to the web. This whole bloated scene of browser emulation had to be spawned and now instead of serving 1kb htmls to few friendly bots several megabytes of junk traffic and countless processing cycles are wasted on some menial tasks like retrieving sport match results from the internet.
The problem with these captcha services are _free_ which means people just throw them anywhere. Imagine a world where land mines are free for everyone at the tip of their fingers - you could hardly go outside! Well you can hardly go online now.
For spam/abuse, captcha is mainly about raising the cost of attack, not about eliminating completely, while still minimizing the collateral damage. Captcha is never meant to be a protection against any targeted narrow scoped attacks anyway.
If the attack is small enough that attackers can pay captcha solving service, it's not big enough to matter.
Captcha is here to stay. It is fundamentally a technical mean to deal with the tragedy of commons, and thus won't disappear anytime soon.
it's really annoying.
It used to be a minor annoyance and I understood the reason for them in those cases. Now they've moved way beyond the 'pain threshold' and it's moved on to a blind hatred of them for me.
A similar thing happened for ads.. At first they were OK and I was fine with some ads paying for a sites. Then they became more intrusive flash crap. I started getting annoyed. Then all the tracking came in and full-page motion and sound ads and they really started abusing their privileges. Now I hate them so much I will never turn off my adblocker again for any site. They've just overstayed their welcome too much.
The same thing is happening now with captchas. Started as a good cause, but totally took advantage of the users to solve a problem that's just as well solved on the back end in most cases.
Or introduce a law where they have to pay us for using up our 'brain time'.
Do the labels I provide belong to me or to Google? .. when I signed no job contract with them to provide that information.
edit: aha! https://www.hcaptcha.com/labeling - that's why it wasn't mentioned. One more labelling service that I won't like.
I feel like Cloudflare might have something to say about that, given that Cloudflare is an independent cybersecurity service and uses hCaptcha.
"enter your name and your favourite vegetable" is to lure bots into responding?
When "I am human", I am asked to select boats and fail because I didn't tick the images with ships.
I would not like to add a Captcha that appears to be smart Alec, has a tendency to trigger the same in the respondents or, worse, makes customers feel stupid. reCaptcha somehow seems better at avoiding this.
Like, hire me for more!
I'm curious how can they train ML models while preserving privacy. Where does the corpus come from?
Because, to me, it equates to someone bragging that 15% of the people they've slept with now have syphilis.
Congrats, albeit I have to say I had less problems with recaptcha captchas. I experienced a couple of cases of hcaptcha just not working correctly and being unable to access something despite the captcha success, which never happened with recaptcha (in my experience).
hCaptcha seems to pretty consistently only make me do 2 rounds, so I certainly prefer it.
I always try to miss some of the obvious items or make mistakes and I (almost) always get through. There's only one service that uses a Google captcha that I continue to use, so it's not really a huge issue for me anyways, and I have decided to stop using it!
It's not too difficult to host your own captcha, I don't see why this can't be an open-source effort.[1]
For me, all captchas are a stain on the web - in most cases, shifting (and multiplying) the wasted human hours from the company collecting data (eg the owner of the contact form) to the user (the person completing the contact form).
The company is saved from filtering through contact form responses from bots (spam and injection attempts) but simply shifts the work to the user who they hope to pay for their service, losing countless enquiries from frustrated users.
In my opinion, the only acceptable use for captchas is when you’re making a useful, free, no-login-required service available to the public and even then should only be brought in after bursting reasonable rate limits.
I run a contact form for a small business. Explaining to them why their tiny website has tens of thousands of spammy requests filled with porn keywords is not easy. Of course webmasters add reCAPTCHA, because they're vastly outnumbered by bots and users.
In my opinion, back end filters are the solution to the problem you’re describing. Not making the genuine users jump through hoops.
Doesn't make the best impression...
What I didn’t like: when I looked for pricing information, I saw that there’s a free tier and there’s a “Contact Sales” tier (for enterprise). There is no intermediate level if you want finer control and just want to know how much that could cost. If anyone from hCaptcha is reading this, I’d strongly recommend adding one or two more tiers or expanding the feature set of the current free tier, at least for some level of granular control.
Question for dang -- why does HN continue to use recaptcha? It's impossible to signup via Tor, and Google is a user hostile company.
One thing I will note, is that hcaptcha seems to be more loose in what answers it accepts. Sometimes I click random images and it still lets me pass.
https://developer.apple.com/documentation/sign_in_with_apple...
PS: Anyone have the corresponding Google link?
It's not exactly what you were asking for, but it can make their disease less painful
I wish Apple would offer a way for sites and services to verify that a client is indeed human via Touch ID/Face ID.
That being said it does make it harder to spam if you don't have a budget to start with.
The ambiguity might be because my phone is set to Japanese. I got asked to identify 電車 (Densha, electric train), but there were also images of diesel train. Is it translation error, or is it really asking me to identify electric train? (The correct term for train in general in Japanese is 列車 ressha)
We're not "remembering" in the same way, but have good enough instantaneous scoring to correctly guess whether or not a challenge is required most of the time.
Some customers may disable that option to meet their requirements. Not much we can do about that :)
I didn't mind solving a reCaptcha once, I mind forcing myself through these every 10 minutes.
Also since only cloudflare uses it, and I dislike cloudflare, I have this irrational hatred of it.
Assuming you mean Debian Buster: Then get a newer Firefox version from backports. This is more related to Firefox ESR than Buster.
Edit: Nevermind... After reading other replies, I think eznzt referred to https://github.com/dessant/buster
What happened to OpenAI? Do they belong to Microsoft now or something?