Introducing the Invisible reCAPTCHA
google.com
google.com
"Whoever doesn't study history is doomed to repeat it.'
Obligatory xkcd about it https://xkcd.com/792/
It seems like you have a problem with people consenting to things you don't like, just like homophobes have problems with gays because they personally don't like consenting to same sex intercourse and think gays are ignorant of some facts (they're going to hell etc), maybe people don't care or believe in these facts and are just enjoying themselves, ever thought of that? What's the problems if other people enjoy things you don't? Live and let live.
“You are the product of t.v.(...) you are delivered to the advertiser who is the customer. He consumes you”
How exactly would you expect that "having a problem with it" manifests? By not watching TV and not reading any kind of newspaper?
But since you ask, it's very much the same, and people have very much had a problem with that, which raises a second fault with your argument: false premise.
Noam Chomsky. Pretty much his life's work.
Jerry Mander, Four Arguments for the Elimination of Television. (NB: Mander is a former ad executive himself).
Hamilton Holt, a magazine publisher himself, wrote of the fundamental problem in 1909, Commercialism and Journalism. https://archive.org/details/commercialismjou00holtuoft
I.F. Stone spoke of the problem in 1974 on "Day at Night". The video is now on YouTube: https://m.youtube.com/watch?v=qV3gO3zxQ1g
Robert R. Murrow's "Lights and Wires in a Box" criticised generally the television industry, in the 1950s. http://www.rtdna.org/content/edward_r_murrow_s_1958_wires_li...
A couple of Stanford researchers in the 1990s recognised the corrosive effects of advertising on incentives in the provision of online services. The prior awareness didn't save them from faling into the same pit: http://infolab.stanford.edu/~backrub/google.html
Niel Postman. Technopoly and Amusing Ourselves to Death.
Vance Packard, The Hidden Persuaders
Naomi Oreskes and Eric Conway, Merchants of Doubt looks at influencing on a much broader basis.
Banksy's art is, in many regards, a direct refutation of advertising. The letter attributed to him on the topic is originally by someone else however. It remains exceedingly good reading: http://thefoxisblack.com/2012/02/29/banksy-on-advertising/
There's a large literature on the subject.
Generally, psychological aspects of advertising: http://www.worldcat.org/search?q=su%3AAdvertising+Psychologi...
Criticisms of advertising: http://www.worldcat.org/search?qt=worldcat_org_all&q=critici...
At Wikipedia: https://en.m.wikipedia.org/wiki/Criticism_of_advertising
I was questioning someone feeling "hypocritical" for using other google services but disliking their new recaptcha tool.
As for "the contract": I'd wager that for most people, most contracts they see/sign are easier to get a clear understanding of, than Google's business model of "free" services.
Oh bull shit. The contracts you sign (or click I agree to) are thousands of legal definitions and words that most people don't even read, much less understand.
Google's contract is very simple: We show you ads (the more we understand your interests, the more targeted our ads) and you get to use our stuff without monetary cost.
And where do they spell that out?
Travel to/from the United States without some social media footprint is on the verge of becoming suspicious in and of itself.
Benjamin "Mako" Hill has long since observed "Google has most of my email because it has all of yours". https://mako.cc/copyrighteous/google-has-most-of-my-email-be...
Facebook creates shadow profiles of non-member identities.
It goes on.
I wouldn't have a problem if they'd open the data set, but I'm not exactly keen on being forced to work for free.
As mentioned elsewhere in this topic; I'd much rather spend some of my free time to help improve OpenStreetMap (and I do).
Going from throwing away the work to improving Google's image recognition is a Pareto improvement so I'm interested in the justification.
From one perspective it doesn't matter - your batteries are gone either way.
But from another perspective, inspectors who keep what they find have completely different motivations - making the same outcome much more sinister and corrupt.
If they want something in return, perhaps they should consider setting up a paywall.
If you, as a site owner, don't want people to share information your server sent to their computers, implement strong DRM or don't allow off-premise access altogether. You might loose many prospective users, but at least your information will remain "safe". Probably.
If you want to share information on a website but can't think of a reasonable way to make money of it, that's your problem, not mine.
Lastly, I fail to see how a discussion on copyright relates to recaptchas or even site-owners' failed business models. Unless you can get this thread back on topic, perhaps we should continue this (most likely frutless) discussion somewhere else.
This isn't information sharing, this is using a service which has costs (hardware, software, employee time, power, cooling, etc...). You are using the ReCAPTCHA service. If you don't pay them for that (either in time or in money) then in my opinion you are stealing the cost of that service.
Just like how if you got a taxi, then just walked away without paying. Yeah, you aren't stealing any physical goods, but you are costing the taxi company something and you aren't paying for it. And if you agreed to get a free taxi ride in exchange for washing their taxi car for them, and you just walked away at the end, that would be stealing in my eyes.
Using the reCAPTCHA service lets them use the data from it for machine learning. That's the reason they are providing it for free, and if you attempt to use the service without paying (or in your scenario maliciously provide false information, which not only doesn't contribute, but can mess up the results from those who do), I feel you are stealing it just the same.
I'm under no obligation (moral or otherwise) to make honest contributions. If we wanted to continue with your analogy, it'll be like being forced into a taxi, then driven one meter, then forced to pay a large fee because the landlord doesn't want me to walk on his lawn. I didn't choose to ride the taxi. The landlord did (ignoring saner options).
If google wants a revenue from their service, they should consider charging site owners.
If google wants to train their NNs, they'll have to pay people to classify data for them.
If google doesn't want to do either, and doesn't want people like me contributing to their system, don't let site owners allow us to use their site.
Otherwise, I think google should shutdown the service.
You do have the option of not using a service that uses recaptcha. Nobody is forcing you to fill them out, you are making the choice to do it because you want the service the site owner is providing, and the site owner wants recaptcha.
If a taxi company keeps getting screwed over by patrons of a given restaurant, the company should stop providing service for said patrons, seek compensation directly from the people who require their service (the restaurant), or seek legal action against either the patrons or the restaurant.
Demanding otherwise is just naive.
edit: Also worth noting that some of this "free work" that users do for Google is used to improve bot-detection overall. The comment to this blog post (the post itself, and its paper are great reads) is a nice example:
https://security.googleblog.com/2013/10/recaptcha-just-got-e...
Essentially, the user is complaining (and justifiably so) of being served the pre-Street-View versions of CAPTCHA; I had forgotten how bad they could get: http://i.imgur.com/01F2eES.png
Of course the site (and with it visitors) benefit in this exchange, so I agree with you that it is not entirely "free".
The original recaptcha was used to help clear up text from book digitization projects where OCR couldn't understand the data. The appeal of this was that these works were in the public domain, and thus proper digitization of these works is a societal-wide benefit.
Since Google purchased it, they've been using it to:
1.) Clean street view data for a proprietary product (Google maps)
2.) Build training sets for unknown ML purposes
These are activities that Google could very much pay a group of people to do. Instead, through recaptcha, they are getting that work from the end user for no payment. A case could be made that it's not free for the owner of the site that deploy recaptcha (because they get value out of the service, and Google gets data/ML services). However, the actual end user who has to fill out the recaptcha does not benefit in any significant way. Since a recaptcha is an inconvenience to the end user, that user pays both with the time to fill it out, and the data gathered by Google.
TL;DR Some people do not like that Google benefits from a transaction where Google is not a party, and where otherwise, Google could generate the benefit using their own resources.
According to Wikipedia and the New York Times, reCaptcha was not developed for public domain works. Its pilot project was to digitize the NYT archives, archives which were not released to the public domain nor are fully available without being a subscriber: http://www.nytimes.com/2011/03/29/science/29recaptcha.html
I'm not a machine learning expert but I'm going to laugh at your suggestion that Google could "pay a group of people to do". In the above referenced NYT article from 2011, recaptcha's creator says several million words were being processed by recaptcha per day.
And again, have to disagree that the end user "does not benefit in any significant way". We would not be discussing this if Google hadn't learned from massive user data to iterate their captcha from distorted word mush to what it is today. Captcha was a serious drain of user energy and patience, that's why recaptcha was invented in the first place. And the worst captchas were tolerated because automated usage was a financial threat to websites that end users use.
It was provided a free service[0] (search free). I'm pretty sure that no one in this conversation is saying that the cost to run the service is free, and arguing such is a strawman.
was not developed for public domain works.
pilot project was to digitize the NYT archives
While the original "recaptcha.net" website is no longer available online (even via archive.org). There are plenty of sources still available that bely this claim. that will help convert printed text into computer-readable letters on behalf of the Internet Archive[0]. The team is involved in digitising old books and manuscripts supplied by a non-profit organisation called the Internet Archive[1]. "There's still about 100 million books to be digitised, which at the current rate will take us about 400 years to complete - Luis von Ahn, Carnegie Mellon"[1]. The sources are from 2007, when recaptcha was first introduced. There are more, but I picked the first two by going back to the beginning of the wikipedia entry for recaptcha and looking at the supplied sources.
This is before Google bought it in 2009. While I can't speak for anyone else, this is what I mean when I say "original" recaptcha.
laugh at your suggestion that Google could "pay a group of people to do".
They got a group of people to do it for free, which implies (barring salaries) that they could also get a group of people to do it for not-free. Thus the argument that people don't want to work for Google for free.
Google hadn't learned from massive user data
And when Google has met it's ML goals and decides it gets no further benefit from recaptcha? Remember, it's not a free-to-run service. They only continue it as long as they benefit. If recaptcha shuts down, it would have been nicer to have the fruits of that work available through something like the Internet Archive, than something like Google Books.
[0] - http://www.cmu.edu/news/archive/2007/May/may24_recaptcha.sht...
This isn't true. The end user benefits in a very obvious way by being able to use a site that hasn't been crippled or spammed by bots. It is very possible that many sites either would not exist or would be of much worse quality without a recaptcha-like service.
Everyone profits from recaptcha.
In other words, if Google's service didn't exist, you'd be experiencing less of the site, or no site at all if they chose to use a paywall that you refuse to pay for.
Anyway, I am not completely against it, as I said above depending on the situation I make use of it.
The value Google gets from the thing they want, and the value you get from the thing you want, aren't particularly comparable.
IMO, this is strictly better.
Google is using ReCaptcha as measure to get an unfair edge over competitors and make competing even more impossible.
You must not be living in the U.S. Even as strong as the FOIA and public records laws are, there's a huge amount of information and records that are exempted from free access:
http://www.rcfp.org/browse-media-law-resources/digital-journ...
Does a sweet pastry count as a cake? Is this stretch of water a river, a lake or the sea? Are those trees I see in the distance on this photo of a mountain? Does the front of a restaurant or a veterinary surgery count as a "storefront"? As a British English speaker I had to guess at that last one, and some of the others seem culturally dependent too.
Or, my personal nemesis, the ones which ask you to select the squares containing road signs, and there's always a couple of squares containing a 3-pixel strip of the very edge of the road sign, and you don't know whether you're supposed to count those or not.
I think of the personal data they collect from me and these types of things as my payment for their "free" services.
Doesn't help that I only see it on fairly shady websites, same kind that use adf.ly, pop-unders and other obnoxious monetization schemes.
Expect to see more of this.
https://www.cylab.cmu.edu/partners/success-stories/recaptcha...
To be fair, if you use google maps, and don't have any ethical issues with helping googler, you do get a return on investment, in that you get improved ML capabilities and improved address recognition in google maps.
The original recaptcha was only text from book digitization projects.
Google car will replace drivers, Google VR + map will replace city guides. We're giving them knowledge and powers to replace us, for free. Professions will disappear, people will lose jobs, Google will get richer and richer, will not pay me or you for fixing their suggestions in Google Translate or fixing maps or improving computed route from A to B.
No one is being paid for that work, but someone is monetizing it.
It's just an opinion, but if the things Google has created and shared into the whole of human knowledge towards the betterment of future societies are still being used in a century for my grandkids (and the rest of humankind) to benefit from, then I think it's better than not having it because a few million people who couldn't have done it themselves said 'no' over a pittance.
I can't speak for you, but my dying thoughts won't be a sour reflection on all of opportunities to monetize my existence I might have missed.
Do they really share it, though? You can't download and reuse data from Google Maps or Street View the way you can from OpenStreetMap, say.
I'm not saying this is a fatal problem. Google Maps is a lot more popular than OSM so the free-with-ads closed source approach clearly has a lot going for it. But the data isn't public, they own it.
https://en.wikipedia.org/wiki/Lump_of_labour_fallacy
and some don't, so this point of view is arguable, to say the least.
Regardless, translation is an intellectual work, driving (depending on the point of view), isn't. Same goes for city guides; people doesn't necessarily prefer an electronic device to a human.
I've never personally hired a professional translator and I don't think I would if I was going to travel to a country where the majority of people don't speak English because I assume a professional translator is expensive. However if I know I have free access to google translate which will be useful enough for me to navigate by myself I would be much more likely to go on such a trip.
I'm sure some jobs will be lost but at least for middle class people who arnt able to afford translators the technology will be used to communicate better without affecting professional translators.
There are plenty of similar professions that exist but not utilized by the middle/lower class due to cost. I would love to hire a interior decorator, as I'm sure plenty of home owners would, but I haven't due to the cost and likely never will. If Google (or any other company) offered a service where I put photos of my home online and gave me a free layout with online links to purchase the furniture I would be thrilled and no interior decorator would be out of a job because I wasn't going to pay for one anyway. I think its all about the level of quality you want.
reCAPTCHA was a technology acquired by Google. Not developed in house.
Its a small rebellious act of mine (that is for moot because I'm sure they give the same captcha to other users for verification)
It was essentially "OK Google" for dumb phones.
For anyone not familiar, 411 (in the US at least) was the directory service number. You could call it to get things like what time a restaurant was open until. Directions to the airport. Etc.
Looks like they do little more than just check for a Google cookie [1].
[1]. https://www.blackhat.com/docs/asia-16/materials/asia-16-Siva...
edit: still, it's far better than the previous state of captchas. I'm glad they did this. But it's like for anything to be considered "advanced" or "good" in tech lately, it has to have been powered by "machine learning".
(That said, whenever I used that checkbox widget they had before this announcement, there was a noticeable framerate drop in the browser while the thing was doing its magic. So I suspect, they are at least doing some browser fingerprinting/benchmarking to see if the widget runs inside selenium or a stock browser.
I also remember rumors that they analyze keyboard/mouse input on the page and check if it looks "human", but I'm not sure if that's true.)
If your browser is standard (AKA no anti-fingerprinting plugins) and your advertising cookies are not blocked (privacy or adblocker plugins) you'll probably pass with no issues.
If either of those is not true, you have to solve a bunch of image captchas.
Mouse/keyboard input analysis was just marketing talk; at least when they first released the nocaptcha it wasn't even captured.
Really?
> It doesn't explain how it works nor a demo page
Imagine a web form without reCaptcha. Do you really need a demo of that?
> nor the reason behind why it went invisible.
Because Recaptcha had an annoying, bad, terrible UX.
How does Google determine if the captcha should be shown?
What are the "adaptive captchas" that are shown to suspect users? A demo would do a great job here.
How does an invisible captcha "create value by applying human bandwidth" if the premise is that humans never see the captcha?
No idea how the additional prompts (e.g. "select the parts of the image with a street sign") are shown nicely in this "invisible" UX though.
The next best thing a search company can do is have every website willingly track their users' mouse and key movements, and then willingly send all of that data to the company's inbox. In return, Google provides them with a binary classifier trained on all of the user click-stream/click-move data which determines whether or not the user is a bot!
It's an OK deal for the website owner; it's a great deal for Google. Not to mention, the user is now sending anonymous data to Google, at the expense of the website's Privacy Policy.
Google gets website owners to willingly install live-cameras on every corner of their website, and then willingly send over all of the footage, in exchange for "protection" from bots. Cough The Government gets citizens to willingly fill out lengthy tax forms, and then willingly send over a bunch of money, in exchange for "protection" from criminals. Cough
Say for instance I am signing up for a website, does the password I enter get sent to google servers to be analyzed now?
Oh buddy, I have bad news for you...
https://support.google.com/dfp_premium/answer/1716364?hl=en
Privacybadger does an ok job against tracking pixels. It uses machine learning to try to detect when a specific url or domain seems to be "following" you around the internet, which could indicate a tracking pixel among other things. It then blocks these entities and gives you the option to override.
And yes, I do think tracking pixels are super evil. It could be standard for websites to have a bar at the bottom with little logos[1] showing the companies that are tracking you, with the logos being the remote-loaded content. The fact that these companies feel compelled to make it 1x1 transparent pixels tells me very clearly that they know they are doing something people don't want them to do, because they've gone out of their way to hide it. It's a clear misuse of browser capabilities, yet they do it anyway. What a pack of cunts they are.
[1] or textual short names company ticker style if you want to minimize bandwidth
IIRC, GitHub does that. If you read the email of some notification, it won't show you that notification in your notifications. The solution is to block external image loading in your email client (I know that Thunderbird has that, and I know that Zoho's email client on Android has that).
They're why I block doubleclick.net, google-analytics.com, etc. in my hosts file rather than just blocking their JavaScript.
Honestly I feel mobile is easier to validate
Recaptcha does more than that though. It checks if you are logged into any other Google services, whether your browser user agent matches your actual browser, and I think one or two more things.
"NoBot is a control that attempts to provide CAPTCHA-like bot/spam prevention without requiring any user interaction. This approach is easier to bypass than an implementation that requires actual human intervention, but NoBot has the benefit of being completely invisible."
Works like a charm even now, 10 years later.
I know that they stopped the project, but did Google at least release the data (old public domain text of books)?
Seems like a bait-and-switch. Do free labor for a good cause (PD books), turns out you're just growing Goolge's library which can be taken down at a whim.
No, that was the Distributed Proofreaders project, which is unrelated (just used as an early example of crowdsourced OCR).
reCaptcha originally helped digitized the archive of the New York Times, but that was finished years ago.
The most frequent encounters with CAPTCHA's I see are rejected API requests over VPN.
Then again, it's allow-more-and-retry on many pages already so nothing new in that regard.
I couldn't load any more jobs, or view any more workers. I had to inspect the Network tab in Chrome, open the API request and then click a "I am not a bot" on their API page.
That was a poor implementation tbh.
There is a need for a decentralized Captcha that can't be circumvented by anyone.
It doesn't seem much different than the current "click here" one to me. They are just letting the page owner substitute their own button in lieu of the check box.
Edit: Yep. "Human users will be let through without seeing the "I'm not a robot" checkbox, while suspicious ones and bots still have to solve the challenges."
It really is just the current one, with the site owner's button instead of Google's checkbox.
I'm curious now, will have to give it a whirl when I get into the office :D
Suspect Google will rely on their vast knowledge of people's browsing habits based off IP/account/ad-tracking/browser-fingerprinting to skip the user input aspects. Although that said, a screen reader won't have the standard physical interaction clues client-side that a user is a real person, mouse tracking for ex. is probably a moot point. Not really sure how Google will handle those, or if blind users will get a degraded always-on captcha experience.
Either that or a headless screen reader becomes the scraping/botting tool of choice.
Someone who knows more about this than me can certainly school me on that last bit, however.
My company provides web automation in Brazil, not for spam, not for marketing purposes. Sometimes the client has a website (like an intranet) that she wants to automate through screen scraping. I know it sounds stupid, but that is pretty frequent.
The models are fast (most captchas solved under 100ms in AWS Lambda) and accurate (95 to 100%). I have about 40 different captchas being solved this way.
My team is still working on a solution to nocaptcha recaptcha. The problem for me is not solving the challenges, but submitting them like a human would do.
I'm progressing with chromium headless. As more and more sites add this captcha, I need to find a solution to automate it. And I'm confident I will find.
Google is betting that they can extract meaning from web pages better than bots, and they have had a lot of experience with that. On Web pages, each link does not have the same probability of being clicked by a human given the list of pages seen before. Knowing that probability requires the bot to understand what a human would see, and to perform actions that match a given goal which corresponds to the sequence of pages browsed.
And that goal mustn't always be the bot's goal. Bots have business incentives: they want to get people to do something by writing text that will be seen by humans. Humans, on the other hand, only do so once in a while.
You can view the captcha at https://mypost.io/ .. you cannot create a post without entering the captcha. I ended up creating my own because despite entering the correct answer (selecting the right images), Google Recaptcha would not recognize it.
[0]: http://haacked.com/archive/2007/09/11/honeypot-captcha.aspx/
Here's the problem Google needs to fix. If I'm logged into my account, years old, obviously not a complaint against it, I still end up getting this captcha nonsense. Doesn't matter the site.
Edit for clarification: I'm not even saying it's wrong, just plain ironic
Users need to be directly and visibly be informed, if and when data of theirs will be transmitted to third parties.
As Google's ReCaptcha is based on tracking what websites you visit, what search terms you enter, correlating this data, and comparing this tracking profile of yours, it's quite problematic that the user doesn't even see any captcha anymore. With the previous captchas, site owners could keep pushing the legal problems to Google, partially.
But this new solution doesn't fit with the EU Data Privacy Directive in neither intention nor letter of the law.
IANAL, this is not legal advice. You can not use this in court.
You actually have to make a separate, opt-in checkbox, directly informing the user what will happen.
Why resort to flimsy legal methods when you can just use uMatrix and block those requests directly? It's your computer, after all.
> You acknowledge and understand that the reCAPTCHA API works by collecting hardware and software information, such as device and application data and the results of integrity checks, and sending that data to Google for analysis. Pursuant to Section 3(d) of the Google APIs Terms of Service, you agree that if you use the APIs that it is your responsibility to provide any necessary notices or consents for the collection and sharing of this data with Google. For users in the European Union, you and your API Client(s) must comply with the EU User Consent Policy [...]
The "EU consent policy" from Google is here:
https://www.google.com/about/company/user-consent-policy.htm...