Google’s new reCAPTCHA has a dark side
fastcompany.com
fastcompany.com
reCAPTCHA v2 is superseded by v3 because it presents a broader opportunity for Google to collect data, and do so with reduced legal risk.
Since reCAPTCHA v3 scripts must be loaded on every page of a site, you must send Google your browsing history and detailed data about how you interact with sites in order to access basic services on the internet, such as paying your bills, or accessing healthcare services.
It's needless to say that the kind of data that is collected by reCAPTCHA v3 is extremely sensitive. Those requests contain data about your motor skills, health issues, and your interests and desires based on how you interact with content. Everything about you that can be inferred or extracted from a website visit is collected and sent to Google.
If you'll refuse to transmit personal data to Google, websites will hinder or block your access.
It’s nonetheless a shame that it’s so universally misunderstood how ad-supported megacorps make their money that even highly sophisticated users of the web still talk about the value of personal data (source: I ran Facebook’s ads backend for years).
Much like the highest information-gain feature for the future price of a security is it’s most recent price: ad historical CTR and user historical CTR (called “clickiness” in the business) are basically the whole show when predicting user cross ad CTR. The big shops like to ham up their data advantage with one hand (to advertisers) while washing the other hand of it (to regulators).
As with so many things Hanlon’s Razor cuts deeply here: if your browsing history can juice CTR prediction then I’ve never seen it. I have seen careers premised on that idea, but I’ve never seen it work.
That may be the case for some people, but that is not my complaint, nor that of many folks I know.
I simply don't care how FB, Google and other surveillance outfits make money. I don't care about marketers' careers or their CTRs. I don't even care about putting a dollar value on my LTV to them.
I care about denying them visibility into my datastream. It is zero-sum. They have no right to it, and I have every right to try to limit their visibility.
Why? None of your business. Seriously - nobody is owed an explanation for not wanting robots watching.
But I will answer anyway. It is because of future risks. These professional panty sniffers already have the raw material for many thousands of lawsuits, divorces and less legal outcomes in their databases. Who knows what particular bits of information will leak in 10 years, or when FB goes bankrupt? I have no desire to be part of what I suspect will become a massive clusterfuck within our lifetimes.
If you're correct that this data has so little value, then it is more likely it will leak. FB and Google are the equivalent of Superfund sites waiting to happen, and storing that data should be considered criminal.
That's entirely fair! But also: You have no right to use my website, and I have every right to limit your access.
Recaptcha is simply part of this negotiation.
Do I think Facebook/Google/etc are abusing my data right now? Probably not.
But do I think that large scale collection of my data could be abused in the future? Most definitely. If the Cambridge Analytica scandal has taught us anything, it’s that having access to this data is rife for abuse and often it might happen in unexpected ways.
And do I owe an explanation for wanting some basic privacy? Absolutely not. If a random stranger stopped me in the street and asked me lots of personal questions there wouldn’t be an expectation that I have to respond. Yet the likes of Facebook and Google seem he’ll bent on turning the discussion around when it’s data collected online.
I would think, that in the EU, under GDPR, collecting, transmitting and storing that data is in fact criminal, or at least subject to heavy fines. And under GDPR it won't help to just note the data collection in the TOS or ask the user for permission (under threat to not allow access to the service). So I really wonder how google plans to run this in europe.
What use is "we are transparent with our users about the data we collect" when the user does not want you to collect the data in the first place? And they give you no option to opt out of such data collection? (And for what - just so that they can create a better ad network that can better exploit us with our own data?)
(And don't get me started on Safari spying and all their "anonymous" cookie collection crap without giving the user any choice in the matter, essentially forcing everyone of their users to opt-in to be profiled through their browsing history).
I don’t want to single you out personally but there’s a broad trend on HN of bitter-sounding commentary on the surveillance powers of these companies by people who can easily defeat any tracking that it’s economical for them to even attempt let alone execute that reeks of sour grapes that a mediocre employee at one of these places makes 3-20x what anyone makes (as a rank and file employee) anywhere else.
Again, you’re not likely part of that group, but seriously who hangs out on HN and can’t configure a VPN?
The main topic we discuss is corporate surveillance. We are concerned about all the personal data that leaves our control. We are worried that evading this type of surveillance becomes increasingly difficult.
Some HN users may know how to mitigate these risks, but most people may not know how to defend themselves against corporate surveillance.
This is why me must speak up now, and not just for ourselves.
For the record I am inked all over with anti-equation group stuff: I agree that these companies are too big and powerful (and I would know).
I just don’t see a solution with the present judiciary. If anyone has a bright idea my email is in my profile.
I will thank you all in advance for not shooting the messenger.
For the record, many of your comments here have been thoughtful, and I've upvoted them. I've also downvoted many where instead of responding to other people's thoughtful comments, you just insult them instead. Those are also the ones that other people seem to be downvoting. I don't think anyone is shooting the messenger here.
You state yourself that Google/Facebook publicly claim to advertisers that personal data improves CTR prediction. So I have a hard time believing that personal data isn't useful.
Isn't demographic targeting exactly that, based on your browsing history? Will showing an ad for a car wash have the same CTR for people that liked car products as for people that did not like car products? Or is your point that it still has to be a human that inputs "this is about car things, please show it to people that like car things" and it's not a magic AI that optimizes it automatically? And in that case: isn't that just a matter of time? Build the profile today, build the tech that uses it tomorrow?
Can I believe you? ...even if you're telling the truth, big corps can hide their most malicious practices from most of their own employees.
To me, it doesn't matter how e.g. Facebook actually uses my data today, because even if they're telling the truth they could change their policies tomorrow, or get hacked, or some third party (incl. the gov't) could get hacked, etc. It's better as a user to try and prevent such data from ever existing in the first place.
That's great. So if my browsing history is useless then you won't mind not trying to snoop on it.
Why would anyone ever trust a goddamn thing you have to say about their data?
Unless they pay your salary and are asking you to give your expertise on hoarding and abusing user data, obviously.
If we allow users to harass and attack people who have genuine expertise for posting here, does that make HN better or worse? Obviously worse. Mob behaviors like this are incompatible with curiosity.
I have nothing to gain and everything to lose by shedding light on one of the most powerful entities in existence.
But TLDR it’s not as interesting as people like to think.
> If you'll refuse to transmit personal data to Google, websites will hinder or block your access.
I wonder how true this really is. 20% or so of web users have ad blockers, and most ad blockers block scripts like Google Analytics out of the box. It isn't hard to see that most of them will not make exceptions for a new Google tracking script. So any site that does any kind of testing at all is going to see that ~15% or so of their users drop off if they block users who don't have a reCaptcha v3 score. The only sane business decision in response to this is to go with some alternative.
(Of course, there will be some sites that continue to block users, it's just that they will mostly be the sites that already block users running ad blockers.)
Or were you referring to the risk that individuals would sue Google for getting blocked from random, potentially essential websites?
You do bring up a good point about the V3 being potential antitrust issue, but that has always been a potential problem even with earlier versions of recaptcha. With V3, it's also deferring the liability to the webmaster. The action that the website takes with the score is up to them - in the end it's just a number.
Also as a VPN user, I found out that migrating to more expensive, higher grade VPN, solved a lot of my problems.
In the end it is not privacy, not your VPN that matters from the service provider point of view. It matters that your IP address is spewing malicious garbage. I do not want to spend time sorting it out, as I can focus my activities to revenue generating tasks. Harming some cheap VPN users in the process is collateral damage, but I rather take it than build a form with a perfect attack mitigation and 10x cost.
I hope to see some alternative for reCAPTCHA that does not come with such a strong privacy oriented risks. hCAPTCHA https://www.hcaptcha.com/ seems to be interesting, also monetization point of view. But they are not yet well established company and I do not know what other risks their approach would bring.
- Your ISP is a source of a lot of malicious traffic
- You have some browser extension or other adjustments that makes it harder to analyse you as a genuine web browser
For example, using a browser automation like Selenium testing triggers "hard" reCAPTCHA. Not sure if this because of some automated API that Selenium exposes, or just because your browser profile looks virgin (no cookies) without any prior reCAPTCHA solves.
I don't believe this is true. You only need to include the JavaScript on pages which actively use the reCAPTCHA score. For example, you might only include it on the login and user registration pages.
> To make this risk-score system work accurately, website administrators are supposed to embed reCaptcha v3 code on all of the pages of their website, not just on forms or log-in pages.
So if the article stated that websites were required to put the code on multiple pages (as the comment I replied to did) then the article is factually incorrect.
Running headless chrome is trivial, so just having it sit on the one page where you need to check it won't help much. Collecting more data on the user's action on your site will provide a much clearer picture, much like a video from somebody walking through a store will help you make a decision about whether he's trying to steal something than a single picture of him standing at the check out.
The important difference is that unlike Google Analytics, reCAPTCHA v3 is inescapable. You cannot prevent the collection of your personal data, because then you would loose access to large portions of the web.
From a technical pov, how does one access a user's browsing history from client-side javascript. Isn't that something the browser should protect? or do you mean more that since the reCAPTCHA gets loaded on each page, Google can track what that IP is visiting by where reCAPTCHA gets loaded?
And if you use something to prevent tracking - in my case Brave - reCAPTCHA is a huge pain that often takes dozens of clicks to make it through - delayed by Google to wait out bots.
Some times I think reCAPTCHAs main goal is to bring back those opposing tracking back into the fold of Chrome with painful recaptchas.
please consider not using recaptcha.
Anyone got some URL's that I can block all captcha attempts or does it mean I have to also sinkhole www.google.com[1] ?
( I don't have a problem not being able to access captcha enabled sites. )
[1] quick check tells me I would have to banish this endpoint which sucks because I'd have to parse the URL on every request and can't do it in DNS: https://www.google.com/recaptcha/api.js
Recital 47: “The processing of personal data strictly necessary for the purposes of preventing fraud also constitutes a legitimate interest of the data controller concerned…”
Recital 71: “decision-making based on … profiling should be allowed where expressly authorised by … law … including for fraud or tax evasion monitoring and prevention purposes”
The fact that a third party server handles this is a problem. Because then the publisher has to have a data processing agreement in place with the third party.
This is what makes Google Analytics problematic too. The collection of analytics for improving the service can be a legitimate interest, however the data amendment for Google Analytics basically passes the blame on the publisher. I don't think many publishers read carefully Google's data processing amendment, otherwise they would drop usage of Google Analytics. Actually most publishers aren't even with GDPR for more serious reasons, like not anonymizing the user's IP or sharing data with Google for the purposes of ads targeting.
And there are many questions to be asked here.
Is that data private, for the use of the publisher in question, or is this a shared pool of knowledge between publishers?
If the later, then we have a problem, because even if there is a legitimate interest, it only applies to the publisher being visited. Can a user be blocked due to a profile that was built on another website? We are in murky waters.
---
Then there's always the question ... does the publisher really have a legitimate interest?
Claiming that you can have one under the law, doesn't mean you actually have it. There's a set of conditions that you have to comply with.
For example for the purposes of preventing fraud, at the very least you have to be able to show that fraud is possible. Just because you have a login form that's about managing the user's color preferences on the website doesn't mean that you can transmit the user's traffic to Google.
The requirements for legitimate interests are hard to comply with. And I have a hunch that in this case many websites won't comply.
I do all my mobile browsing on FF yet when I try to use some websites I always get this Recaptcha failed error(1) while it works flawlessly on chrome though I never use it often. Try it, maybe it will happen for you too.
Same happens on most sites which show you that "checking your browser" page via cloudflare too.
The web is very unusable unless you're using chrome because of such antics.
(1) https://cdn3.imggmi.com/uploads/2019/6/27/0dd96b25707ce6e236...
If only it was Google services alone. CloudFlare loves serving up a ReCAPTCHA for Tor users before they can even passively read site contents. That hugely expands the damage done.
The walled garden approach worked for a while for Microsoft, and it's working for now for Google, but eventually, it stops working. Once people leave, walled gardens keep them away.
You can't just opt out of using half the Internet because you value privacy, and nor should you have to. This requires legislation to stop.
This of course doesn’t help explain why Firefox is so heavily targeted by what’s supposed to be a neutral utility like Google Analytics...
One trick that seems to help fool that awful piece of tech: click slowly on the images, as if you were thinking a second or two before each click. Maybe click a wrong image and deselect it again. In other words, behave like a slow human, and it seems to work better than if I solve it as quickly as possible.
Again, being slower and more error prone seems to be rewarded.
If this reduces the world Google allows me to access, it doesn't diminish mine because of it.
> "If you have a Google account it’s more likely you are human"
So, in the future if we don't keep signed into our google account(and let google know every article we read and every website we browse), we'll be cut off from the half of the internet or even more. The amount of control a handful of companies have over the internet is suffocating to know!
I get .7 on my iPhone, I’m guessing that my liberal use of Firefox containers and the cookie auto-delete extension on my desktop will give me a much lower score and cause me to have to jump through extra hoops at websites that implement it, just like the reCaptcha V2 does.
Edit: I also got 0.7 on Firefox with strict content blocking (which is supposed to block fingerprinters), uBlock Origin, and Cookie AutoDelete. I get 0.9 from a container which is logged into Google.
To me, it feels like Google's entire strategy behind reCaptcha is to make it harder to protect your privacy. We've basically given up on the idea that there are tasks only humans can do, and to me V3 feels like Google openly saying, "You know how we can prove you're not a robot? Because we literally know exactly who you are." I don't even know if it should be called a captcha -- it feels like it's just identity verification.
I don't think this is an acceptable tradeoff. I know that when reCaptcha shows up on HN there's often a crowd that says, "but how else can we block bots?" I'm gonna draw a personal line in the sand and say that I think protecting privacy is more important than stopping bots. If your website can't stop bots without violating my privacy, then I'm starting to feel like I might be on the bots' side.
For the irony, I'm still logged into GMail and it still works perfectly, as basic HTML, even with google.com forbidden to run scripts. But it's the flippin' reCaptchas all over the place that make me temp-allow google.com, and then a reload later, temp-allow gstatic.com and reload again. Only then I get to use someone else's site normally, and I can disallow again... it's irritating. And then, this.
BTW that page plainly says the scores are samples and not related to reality. Refresh a few times and watch it change. 0.3, 0.7, and 0.9 seem to be my lucky numbers. I see everyone else getting those and 0.1.
Please stop reading things into it oh it's too late. Maybe they suddenly started seeing this page hundreds of times in the referrer and added that bit afterward, I don't know.
so why are they having you solve image puzzles if they know that they are going to fail you? even if they know that you are human...
The problem is that they aren't trying harder for users who aren't logged in.
That's on you, not Google.
Using Chrome, even incognito and with uBlock I get 0.7
(╯°□°)╯︵ ┻━┻. F you, Google, this is blatant bullying, technically unjustifyable abuse of your stranglehold over the whole web platform.
On FireFox with uBlock on and logged into my corporate gmail I get 0.9, switching to a private tab I get 0.7. This is with every privacy setting turned on in the FF options.
This is essentially going to let Google gatekeep the web if you aren't using their services.
I wonder how many people here are using a VPN or accessing from a non-western country -- I'd bet those are much bigger factors
FF incognito window not logged into Google account: 0.7
FF incognito window not logged into Google account through VPN: 0.3
FYI I have uBlock, pi-hole and a bunch of privacy widgets enabled
Come on, how is everyone in this chain so blind. It's literally in bold and the single largest block of content on the page:
NOTE:This is a sample implementation, the score returned here is not a reflection on your Google account or type of traffic. In production, refer to the distribution of scores shown in your admin interface and adjust your own threshold accordingly. Do not raise issues regarding the score you see here.
Please see the sibling comments (that were there before yours) where this is already being discussed, before being insulting.
>the score returned here is not a reflection on your Google account or type of traffic
I got random scores as well. It looks like this is just a sample of the data structure that the service returns, not the actual score.
I'm also logged into google and fb which also doesn't affect my score. Only shows how broken their algorithm is :(
edit: just tried it with chrome and my score jumped to 0.9! So definitely not my ISP. It's just my browser that Recaptcha doesn't like. If you put two and two together that's really evil shit, even for Google!
Almost unused Chrome installation, also without addons: 0.7
Privacy Badger and ABP on my work (less-locked-down) Mac.
// TODO: add impressive-looking math
if (signedin && trackedEverywhere) {
return 0.9
} else {
return 0.7
}
I think we give Google way too much credit for their talent. This is the same company that didn't feel like finishing their website for two decades and subsequently stole $75 million from their users even when Google knew [1].The same company that somehow still doesn't reconcile amounts owed and just keeps the money when they randomly-ban users and hide behind fake support emails, but they did feel like paying $11 million to keep that away from scrutiny [2].
[1] https://www.businessinsider.com/google-emails-adtrader-lawsu...
[2] https://www.searchenginejournal.com/adsense-lawsuit/248135/
Brave isn't particularly "unusual", and is even based on Chromium - surely this is Google blatantly punishing non-Chrome users?
I get a 0.7 on Chrome with no account logged in and uBlock Origin installed.
Same browser, same plugin but incognito it's 0.1.
Papa google needs my data to trust me. Makes complete sense but still interesting that you can affect your score by giving in.
> error-codes": ["score-threshold-not-met"]
Not sure if happy or not happy with that. I will conclude happy enough.
Linux, on VPN, Firefox. Not logged into any Google services. Cleared caches (still same IP), no difference.
Chrome: .9
Safari: .7
Firefox: .1
I have adblock running on all three, and I use containers on Firefox.
In incognito mode in chrome, I sometimes get 0.9 and sometimes 0.7 when I reload.
I guess this is a 0 for me then
Safari macOS with the same adblocker: 0.7
Firefox macOS with a lot of adblockers: 0.1
Then I remembered that I put this in my /etc/hosts a few weeks ago and forgot about it.
127.0.0.1 google.com
127.0.0.1 www.google.com
[Edit] So if nothing shows up for you on that page, check for that. Also I just generally recommend it. Google has some unethical practices and duckduckgo.com is pretty good.You need not to use hosts to block it, uMatrix could do it by itself.
Does the government realize the consequences of this? Both that it pushes users to use Chromium-based browsers, and that they're helping to solidify a company that already has a near monopoly in the browser space?
Further, this quote is very creepy:
> To make this risk-score system work accurately, website administrators are supposed to embed reCaptcha v3 code on all of the pages of their website, not just on forms or log-in pages.
With AMP, Google Ads, and reCAPTCHA, Google now has access to pretty much everything that people do on the web.
In which case those private companies should now be deemed an extension of the government and fall under all rules a government organization has to abide by. If they do not like it they can forbid the government from using their software/products and can sue if the government does not abide.
It will also log you in to google on the first page.
Additionally, the stations at the DMV all have tablets on stands, showing Google logins for some operations.
So I don't really see the amusement.
I think it comes from the sad fact that, generally, ambitious politicians create organizations to get themselves elected, rather than previously-existing and purposeful organizations presenting candidates that represent that organization's values to a larger audience.
I use Tor fairly regularly and it's a complete nightmare. I sometimes spend 5-15 minutes solving reCAPTCHA (since your Tor circuit changes every 10 minutes this can result in having to solve the reCAPTCHA several times).
That said, it's one of the most effective means of combatting automated spam and credential stuffing attacks. In a recent implementation I did, having 2FA active for your account bypasses the captcha requirement, but the vast majority of users are still too non-technical to use 2FA and are subject to the frustrations of reCAPTCHA.
A responsible spam protection system should allow every spam (and consequentially responsible user) from an ISP.
If a ISP shows sign of abuse, then show Captcha or other system that will block some spam while also blocking some valid users. This is a evil-for-the-greater-good solution. Do not fool yourself into thinking this is a solution (i.e. without caveats)
Impacted users can complain to both the service provider (you) and their ISP. And that failing, switching their ISP (i.e. voting with their wallet --how that happens in a monopoly is another discussion)
Bottom line, if you show captcha for all users (even for ISPs that are now showing signs of spam) you are intentionally blocking some users for no good reason. And you are part of the problem. Sadly, this includes the US government as they blanket censor all their forms (from visa request to DMV visits) behind Google(R) captcha(tm) at all times.
It still either lets me through, or doesn’t even manage to display images for me to click on because it doesn’t like my browser settings. (The latter is more common than the former these days...)
I wonder how much legitimate traffic bounces because of reCaptcha. Can sites even measure this?
> It’s great for security—but not so great for your privacy.
For individual users, security and privacy frequently go hand-in-hand. But for site operators, user privacy makes security a lot harder. The more you know about a user, the easier it is to figure out if they're an adversary.
Ultimately why does it matter if the user is a human or bot, as long as they are being a valuable user? What's wrong if a bot buys some of your inventory, pays for it and everything? What's wrong if an NLP bot responds to discussion threads with scientific facts and citations?
100% of the time, a bot buying things from a store is doing so to test a database of stolen credit cards the bot's owner has purchased/stolen. Accepting those sales means you'll get hit with chargebacks a few weeks later as the real owners of those cards see their statements. Then your store gets shut down for exceeding the maximum 1% chargeback ratio mandated by Visa and MasterCard. So preventing this scenario matters a lot, and when someone targets one of my stores for testing like this, enabling a CAPTCHA on the payment page is one of several, often-essential mitigations. Blocking IPs, blocking whole countries, including a nonce in the form, etc are on their own insufficient most of the time: the readily-available tools for this kind of attack already handle rotating IPs, retrieving a new form nonce on each try, spoofing the proper referrer, etc.
In the book "Spam Nation" (Brian Krebs), a group of students try to fight fake online pharmacies. To do this, they created an army of bots and placed thousands of fake orders every day. The goal was to create so many fake orders that the human processors (many fake pharma stores were not fully automated) had to spend a significant amount of time clearing out the fake orders before getting to the real orders.
Now imagine this on a legitimate website. Not every website is automated, not every organization has the same resources that Amazon does. CAPTCHA's are a great way to ensure orders are coming from real people. They could still be faked, but the bar is a bit higher.
[edit]
Also, NAT means that there could be hundreds or thousands of individual users on the same IP address (many dorms at smaller colleges are setup this way), so you don't want to rate limit by IP address either.
As a security consultant, it is not uncommon to recommend a CAPTCHA for things like successive, three failed login attempts from a single IP address within a certain time period. But I do agree that CAPTCHAs are used too frequently, and some security people recommend them for just about everything. As someone who blocks a lot of tracking and feels the pain of these tracking monsters (that's what CAPTCHAs are these days, more than the Completely Automatic Public Turing test they're supposed to be), I always think very carefully whether a CAPTCHA is the only option, and I'm sure to recommend CAPTCHAs that fit the situation but are less invasive than a third-party one.
Edit: oh, right, credit cards are common in many countries and banks set chargeback limits. It's still crazy to me that your 'public key' is also the only thing needed to withdraw money from your account, thereby necessitating a chargeback system. I guess for credit cards a CAPTCHA might be useful too.
I had problems with my spinner at first. However, this is one of the things that is really annoying about using captchas instead of passwords.
It is possible to create a secure account for something useless without using the system and you probably won't get spammers anymore.
What a dumb idea, you would want to implement it yourself, because the people working and maintaining the system(s) will all have some way of doing that already.
I don't know how much the government can take away from a site like this as well. But it's a bit like trying to ban a kid because the kid got an old friend on their facebook because his parents were "bad".
So, even though you have a big idea about voting systems, it would be a better option to require a system that only exists to be able to be used for good reasons.
Even something as simple as a question: "How many legs does a spider have?" ____
And then cycle through different types of free form questions of things that most people should know. Perhaps block the IP after {n} failed attempts for an hour.
And I am not speaking here about how Android and Android apps (which is allowed by Google) track users.
> an open decentralized protocol for human review that runs on the Ethereum blockchain.
Is this basically a JS crypto miner?
They've been doing it for a while now, Tech Altar even had a video about it the other day: https://www.youtube.com/watch?v=ELCq63652ig
Along with the censorship and privacy issues, I guess it's time for them to change their payoff, "don't be evil".
They dropped that years ago, literally and effectively.
So the downside here is that no one has a credible way to compete with Google? Maybe because their Google cookie actually is a pretty good indicator of humanity?
That's nonsense. Tons of people do. There's LOADS of great research on captcha that isn't implemented by any vendor. The roadblock is that NO ONE WANTS TO, because it's a thankless, unprofitable task that puts you dead in the crosshairs of a ton of very organized people who will devote huge resources to circumventing or breaking your offering.
"A land grab," sure. Of a nuclear wasteland covered in small arms battles.
Oh you sweet summer child. Would you by any chance be interested in buying a bridge?
Despite that, even assuming if it's true and we'll have a lovely accurate AI captcha system. The big down-side is that captcha is breaking programmable web. I maintain a lot of small software crawlers from simple notification applets to bigger analytic crawlers and the web in the past few years has been increasingly hostile towards this. API endpoints are disappearing and in general make distributing free apps very difficult. The web-crawlers are super simple to write and even maintain but services like cloudflare, distils, captcha break them and while there's always solutions to these systems they are very hard to distribute to users (you can't really pack in pupeteer, selenium or some other webengine automation stack with your app).
Public data should be public. These sort of idiotic measures are not compatible with web protocol. The web only know one thing - 1 IP address == 1 person and it should be encouraged not dismissed.
This is complete and utter bullying. Bullying on user privacy, bullying on Firefox.
Somebody please tell me where to go?
> For instance, if a user with a high risk score attempts to log in, the website can set rules to ask them to enter additional verification information through two-factor authentication.
Seems to me, this could easily flag genuine users who access the site through a non-standard flow - e.g. because they use assistive technologies. In the worst case, this could result in impaired users being forced to jump through additional hoops - or being blocked completely.
I would however use Google's system if the site is massive and there is the possibility that someone is using a script or some program to algorithmically bypass the (simple) captcha, and register accounts en-masse and trying to create a psyop[0], or disinformation campaign, or even a sockpuppet army.
[0] https://en.wikipedia.org/wiki/Psychological_Operations_(Unit...
That's not a viable reason. Anyone doing so is going to have a budget and human reCAPTCHA solving is less than $0.01 per CAPTCHA. It costs very little for mass account creation, reCAPTCHA or not.
They will get likely access of a small percentage of those accounts and you have to do damage control.
Bandcamp is an online music store and I'm prompted for a Google ReCAPTCHA every time I try to log in, which really causes me to do it less often than I normally would, as I must permit Google JavaScript for it to succeed.
I've wanted to send a complaint about this to Bandcamp, but their email is hosted by gmail and none of my messages get to them because I host my own email. Adding reverse DNS and SPF is enough for many email servers, but not Google.
I find it a bad situation that my experience with a business is worse, due to Google, and I can't even contact them to let them know, due to Google.
This is outright creepy.
For example, even Cloudflare, which has its own "checking your browser" protection, still uses reCaptcha in some other cases... Why doesn't Cloudflare offer a reCaptcha alternative to their customers? (a transparent one, more like reCaptcha v3 rather than the intrusive 5-second one...).
I try to never be logged into my google account as a matter of principle. Maybe I'm just fooling myself thinking this will make tracking me more difficult.
When I published a brand new site, I got thousands of bot sign-ups in the first couple weeks. reCAPTCHA apparently had no effect on stopping them. The bots signed-up with real user emails, causing my site to send unsolicited email to them, which affected my domain's email reputation significantly.
I rolled my own invisible CAPTCHA and immediately stopped ALL the bot traffic.
Rolling your own CAPTCHA is a fantastic option, because folks like 2captcha are never going to take the time to integrate against a one-off solution. When you do something like that you drastically increase the barrier by introducing the need for a reverse engineering skillset to bypass your unique solution... and that skillset is expensive lemme tell ya ;)
I spent all day clicking on sidewalks ans traffic lights! Or buses. No idea what they want the clicks on buses for.
I actually had problems upstream of the CAPTCHA. This was on a Wordpress site and I was patching up the Contact Form 7 implementation on there.
What shocked me was how naff Wordpress is. After however many years it does not come with a contact form built in. Comments yes, but a contact form, no. Then the fairly de-facto Contact Form 7 would not work with Google Captcha 3 and the latest V5 Wordpress. So there must be hundreds of sites out there with contact forms that do not work. Then there are people cussing ReCaptcha when there is this hideous mess of bloat going on.
The Contact Form 7 didn't even use HTML5 form validation and styling it was a nightmare.
I eventually went to CAPTCHA2 with the box you tick. Having the v3 box in the bottom right of the screen on every page was not what the client wanted. Plus it didn't work with this kludge known as Wordpress.
I think the issues raised in the article are not that big a deal. If you have logged in to Chrome and you are on your normal device and IP address then you can get a free pass. Why not?
I seriously advise anyone to test their implementations on a non-logged in PC, it is an eye opener. And a time consumer. But forms have to be made to work. You can't have people locked out.
There is a lot to be said for backend validation based on form data, I like to make forms unique with a hidden timestamp in it that is MD5 encoded. You can then see if someone has spent long enough on the form for it to be 'real'.
The platforms - web-browser or operating system - that run these apps - web app or native app - for the benefit of that user - should provide well-understood intuitive experience around what is allowed/possible to be collected and used from the user's device.
Now, technical mechanisms are one major part of the solution. In this regard, these mega corps should be held to a higher standards as they run the platforms as well as the biggest apps on those platforms.
But we also need legal protections that make both the application owners and the platforms owners responsible for any abuse of the user.
This particular case is eerily similar.
Credit card fraud prevention companies do the same thing - they say they need to know as much transaction data as possible in real-time for them to know which is a legitimate transaction and which is a fraud transaction. There is misdirection and fog around how they justify this with thinly veiled technical explanations about network effects and criticisms about monopolistic by design.
The reality is fraud can be prevent by designing the product differently in the first place - chip & pin - multi-factor authentication etc. technology is present to prevent theft and fraud without having to collect so much data centrally.
In this case, similarly, to prevent DDoS attacks, there are other anonymous non-data collection oriented solutions possible. More research and collaboration is needed to evolve the Internet architecture to react to DDoS attackers and other types of technical abusers of your app, catch them and prevent them from growing. Instead, we get these centralized monopolistic solutions.
This defeats the purpose of using 2FA: to require a second factor every time you login. This completely negates the benefits of 2FA if a hacker has gotten my username/password through a keylogger. It's easy enough to get a good score if you just disable tracking protections and login to a Google account, then a hacker can easily break into your account. I was thinking it had to be the author not understanding how 2FA worked, but Google is actually advocating this ([0]):
> login With low scores, require 2-factor-authentication or email verification to prevent credential stuffing attacks.
You would think the people behind recaptcha would understand how 2FA is supposed to work.
This is hilarious because this is the worst, most needlessly complicated solution to identity that one could ever imagine. It’s funny how it apparently takes a PhD to tell you that google isn’t analyzing your behavior on the website. They track you across every webpage you visit that has their code running. And it goes way beyond cookies. Look at the filter bubble research that duck duck go did — they have an idea of who you are regardless of what cookies you have or whatever else. And this data has informed captcha results before this latest iteration. It’s complicated, needless and also gives a bunch of sensitive data to a private company that shouldn’t have it. Nobody cares.
Identity services are the alternative to this. Have a company that has multiple ways of verifying identity including operating physical locations where you can show up and prove beyond any doubt that you are person X. Once identity is established, something like a yubikey or whatever can be used to authenticate various things like making an account on a website or what have you. If you get hacked then you can rectify by engaging in one of the identity verifications tasks, up to and including coming in person and being biometrically verified with absolute certainty. The company would make money with modest fees to users and charging websites to use their service to verify their users.
It should be that the government has all this in place, and you can use secondary identification numbers for websites. But I’m the United States identity is broken, based on a shitty Ssn where if it’s stolen you are basically fucked.
This makes sense to me. The presence of cookies is a strong indicator of normal human browsing, and Google would only be able to see their own cookie.
There’s a potentially bigger risk being overlooked. Google can execute first-party script. This means they get every user’s session credentials and can freely impersonate that user. So can all the other trackers in use.
I don’t understand how anyone thinks this is remotely okay.
hahah, wow. Put this tracking software suit on every page please - it's for your own good!
I used to kinda believe in Accelerationism[1] - philosophy where you encourage and accelerate flawed systems to promote a breaking point. However it turns out our society doesn't really have a breaking point.
I'm noticing that little blue bock in the bottom right corner of more sites now and I think I just want to block it entirely. If a site won't let me login because of it, guess I'll just stop using it. If it's something like my Bank .. well guess I gotta get a new bank.
Nothing really excuses this level of potential tracking.
Sure, we hate that, but it's there, and plus, Google probably already has a cookie on your computer anyway.
Given all that, this seems like a very minor privacy loss to whine about compared to all the other crap that Google/FB/etc does.
Please help me if I'm missing something unique about this particular issue.
Easy to say you'd bounce on a website you didn't care about. That's a bit of a tautology.
Doesn't that sound a bit dystopian?
Just vote with your feet and tell the site to drop it or lose you.
Just tried to register via Tor and got this: https://i.imgur.com/svjfLqo.png
Kind of toothless to say "I'll NEVER use ReCaptcha (except for websites I want to use)!" In fact, I'd go a step further and assert that, while you complain about ReCaptcha, you're actively benefiting from HN using ReCaptcha since you see less spam day-to-day because of it. ;)
https://news.ycombinator.com/item?id=20158386
Long history short, don't use a captcha if you don't need it. And most of the time your website don't need it.
Now remember that every interaction with the government must happen online (from requesting a US visa to going to the DMV), and all those forms are behind a Google(R) Captcha(TM) censorship system, which ranks users based on how well Google(R) can monetize the current user browser session. Let that sink in.
But Im not sure, since the browser sessions for some proxy users like TOR exit nodes are so short
... can we please get a serious antitrust investigation now?