Google's Street View computer vision can beat reCAPTCHA with 99% accuracy
googleonlinesecurity.blogspot.com
googleonlinesecurity.blogspot.com
The street address/book scan approach that Google uses is interesting in that the exact solution is not known, so they presumably have to be somewhat forgiving in accepting answers (as their machine learning might have gotten it wrong). Perhaps this is what their "risk analysis" refers to--whether their response seems "human" enough according to their data, not necessarily whether it's correct.
I don't see a way around this problem for free services that still preserves privacy (so directly using some government-issued ID is off the table). Maybe some Persona-like digital signature system, where a person can go to a physical location with a government ID, and get a signature that says "Trusted authority X affirms that person Y is in fact a person". Obviously this still has problems, as you need to trust X, and it's a big pain in the ass.
There are parallels to the realm of passwords, which are also becoming obsolete (not that there's a good replacement...). Anything that a human can feasibly remember for a bunch of sites is becoming easier and easier for computers to guess.
So basically, computers are taking over the world, and we can't do anything to stop it. God help us all.
reCAPTCHA presents a known problem and an unknown problem to the user. If the user answers the known one correctly, it assumes that the answer to the unknown problem is correct too. I believe the text-based reCAPTCHAs will accept an answer within a character or two as correct.
On the other hand, why is that we insist on humanity? None of the systems can block a captcha-solving team based in Nigeria. What we need to block is bad behavior by users and not their humanity.
Of course, there's a quite relevant xkcd: https://xkcd.com/810/
What's to stop a bot writer from getting a signature and plugging it into his bot?
This said, i don't think "trusted authorities" are the way to go - we've seen the failure of that model with x.509, there is no need to repeat it elsewhere.
Person signs government ID number with public key, presents it to authority
Authority signs public key and asserts it belongs to a real human
Service presents random number at signup time
Person signs the random number with the same public key, and presents the key signed by the authorityAnd as an uncle (?) comment points out, there could possibly be blacklists of signatures of spammers shared between providers. That definitely would cause some new problems, of course, but who knows what the future may hold...
That seems pretty obvious to me, making me wonder what you're really getting at.
2. The signature is linked to a specific person, so if the private key is stolen, then you (a) revoke it; and (b) punish the thief - because obtaining that signature most likely be treated as a felony, unlike spamming comments or other uses for breaking captcha's.
As to your questions, I think the way it's going with these convolutional neural nets, the idea of being able to tell a human from a computer with something like captcha isn't going to work. I'm thinking a captcha that still would work would be a randomly generated natural language instruction that a user has to interpret with some top down reasoning (possibly with a natural text answer). We've still got a ways to go before we can do that really effectively with software.
I was going to propose that it's only returned on good behavior, but that gives sites an incentives to harshly police their new user population.
Ideally you'd have some sort of service that both parties trust that is holding the BTC.
However, it limits the user population to people with a debit card and $5 to spend. A bitcoin-based solution right now would be even more restrictive, since apart from buying bitcoin in person, you pretty much need a bank account to acquire them.
I don't disagree, but although this specific problem can be automated, the arms race will continue as newer problems (requiring uniquely human intelligence) are presented. We can keep moving the bar until the singularity.
For example, a captcha: "Pick the most unsafe environment", followed by pictures of a house, a park, a volcano, a bed. To solve this, AI would have to have a mastery of language, a mastery of object identification, and a mastery of the metadata of those objects as they apply to human safety.
The problem we would then experience is unintentional prejudice and Western-mindedness in our "are you human" queries, similar to a problem that SATs and other standardized testing methods have been criticized for.
Basically, my point is that there's no such thing as "uniquely human intelligence". There's lots of problems that fall into that category now, but with the continuing rise of computational power, it's getting harder and harder to find new problems of this sort without excluding a lot of real people.
I'm probably overgeneralizing, but it's an interesting framing.
I don't see how that follows. Many captchas are just text with a confusing grid overlayed - that's incredibly trivial to produce, and producing one tells you next to nothing about how to solve them. Or if that's too easy, throw in some random distortion - still easy to produce, even harder to read.
Also, being solvable for humans and not computers is not the only criteria for a captcha. For example it must be quick and relatively easy for humans, and provide a definitive consistent answer. The brilliance of recaptcha is that it was making the work that humans excelled at actually useful beyond just 'proof of humanity'.
Actually for a captcha to work, you have to have a problem that is easy for an AI to pose and verify the solution to, but hard for AI to solve. This could potentially rule out some classes of working turing tests.
My sense is that IBM's Watson could tackle this particular question. Also interesting to note that the question introduces external considerations, such as whether or not there's someone dangerous near the house, bed, park, etc. A human respondent will have to ignore the matter of whether there's additional context and second-guess the captcha author to get the problem right -- presumably the most unsafe environment is the volcano.
"Pick the most unsafe environment", followed by pictures of a house, a park, a volcano, a bed.
if you're a spammer is to choose randomly and make four times as many attempts. Which is why catchpas make you type six or more letters giving millions of combinations so you can't do it randomly. To get up to a similar resistance to random attempts with 'pick the unsafe environment' type questions you'd need of the order or eight of them (giving 4^8 or approx 65k combinations). Would you as a user want to fill in 8 of those things?
Watson played a simplified version of Jeopardy with no video or audio questions. This was a key compromise to allow a text based system to compete with humans. The AI problem of recognizing pictures is practically unsolved. It took Google's best researchers, a massive database of cat pictures, and a supercomputer just to train a computer to recognize when a cat was in a photo. So your sense of Watson's present abilities is a little skewed and unrealistic.
* "home and safe" "safe as houses" etc
* "safe in bed" "safely tucked up in bed" etc
* "volcano" ... no clear data. Certainly dangerous being a volcanologist, but that'd be great AI googlefu
* "pensioner attacked in park" .. "our parks are not safe" etc... Hmm, Parks are the most dangerous.
Because frequency of local newspaper reports is a bad measure for volcanos I guess.
But if you could have an AI generate those problems, it would be capable of solving them too!
Current captcha systems exploit the one-way nature of problems, mainly "character distortion". This is what permits easy computer-generation but not computer-solution, and does not seem to be present in the class of problems you describe: it's just as easy to go from "noun -> adjective" as "adjective -> noun".
This is the kind of problem that would still cause issues for some people, especially if you had to scale them to produce identical captcha's relatively rarely (if the same one shows up too many times it can easily be hard-coded, requiring the most basic detection to solve it reliably) in that as you come up with more of these types of questions the expected answer gets rapidly more nuanced.
Whereas volcanoes, and more specifically, active volcanoes can be unpredicable and spurt out hot lava and ash.
The question is worded fine and your attempt to poke holes into this is invalid.
Even if a question is worded well, it can always be interpreted incorrectly. The question should not be blamed for a person's inability to process information.
It's already happening.
It turns out humans come quite cheaply, and you can actually solve captchas using humans at very high QPS for very low cost. And people do!
Site owners will have to come up with alternate means of "bot" enforcement that does not rely on human vs computer detection.
Captchas approach the problem 1), by assuring only people can post and then hopefully increase the signal quality. It sucks because it requires mental effort, is hard on people with disabilities, blocks potentially useful bots (see reddit bots) and leads to an arms race.
What if we tried to solve the problem 2)? One possible solution is requiring a proof of work. Some computation that requires a few seconds, easily disguised while the user is typing the comment or filling the fields. It's an old idea but I've never seen an implementation in the wild. (yes, it would fail for mobile devices. fall back to captcha or something)
Also, regular users who visited your site without wouldn't install any software, so they'd be running the proof-of-work as slow javascript, while spammers could use GPU or even ASIC acceleration if proof-of-work became widespread.
You can be issued a secure key ('a passport') to make new ids ('HN registration') but nobody needs to know who ('the passport') generated the website specific ids once your key has been used. The passport only gives access to generate further keys.
In the original captcha algorithm there are two images shown, one where the answer is known and one where the answer is not known. The unknown - the street sign or that extract from an old newspaper - is guessed at by those filling in the captcha.
If 90 people say it looks like '11' and 5 people say it looks like '71' and another 5 say it looks like '77' then you can be 90% certain that it is '11'. The machine has therefore 'learned' what that number is, it is '11'. This can continue being used in the captcha however, now that it is known, people will have to type '11' for it.
That is how it fundamentally works and why the two snippets of information 'to get right'. You can try it for yourself, get the first one of the street sign deliberately wrong and see if it lets you through.
I suspect that google has been using techniques like this to validate their computer vision conclusions. Which makes their 99% assertion even more interesting, because it's likely 99% confirmed by a very large crowd sourced data set, not simply a staff member going through several hundred samples to come up with the success rate.
- the actual captcha and a word from a book (Google Books)
- the actual captcha and a house number from Street View (Google Maps)
Only the actual captcha has to be typed in correctly.I believe ReCaptcha presents the same unknown value to several users, and rejects outliers.
Recently they added special captchas for untrusted users. If they think you might be a bot, you'll get much harder captchas.
http://googleonlinesecurity.blogspot.com/2013/10/recaptcha-j...
I suspect the nonsense won't matter. You're drowned out.
http://www.google.com/recaptcha/learnmore
Only worked once for me, and only when I typed a word that was of similar length to the original, but I was only able to get it to work that one time. Leaving it blank or with a period or only four letters would be an extremely obvious thing to check and deny access to.
I still think the only reliable way to confirm identity (or humanity) online is an email or SMS verification. Recently, receiving a 2-factor SMS code took less time than the page refresh prompting me to enter it.
I think it would annoy many human users as well.
Yes, it will annoy some real users, so it's not a no-brainer. But it will put a major brake on many kinds of abuse.
And that's assuming that you just don't steal them outright. Suppose you boost somebody's Android game, repackage it, and add a little code that intercepts and replies to certain text messages. It'd basically be a SMS botnet.
SMS verification would make certain kinds of casual abuse harder, but I don't think it'd be a big barrier for well-organized assholes. E.g., the "Rachel from Cardholder Services" people that have been using the murkiness of the telephone network to sneak billions of illegal calls past the FTC.
"How was your day?"
- ntrtyLt (I'm mostly confident)
- tmincrw (I'm really confident)
- rrnmtht (I'm not at all confident of my answer, but I'm really confident that r's, n's, and m's should be prohibited next to each other in captchas like this!)
- MCruncy (I'm really confident of my interpretation, but are capital letters supposed to be reported as capitals on this captcha system, or is this one of the systems that we're only supposed to submit lower case characters even when the character that we're viewing is upper case?)
Making something bug-compatible with the human vision system might be much harder than pure recognition.
Oh, and I might be better at the imagination ones than the current crop, especially on my phone.
But that's how it already is! :) Eventually spam content will either blend in to be relevant and attractive (we can see hints of it in new youtube ads system, but that's matching to ads only, not generating) or it will be analyzed and scrapped (along with other comments) by other equally intelligent systems. So machines will judge both machines and humans equally I think.
If a computer algorithm can create a more informed relevant comment than a person, maybe i would prefer to see the artificial comment.
Which is generally my experience with captcha's these days, I only have about a 50% success rate.
CAPTCHA is a failed strategy, time to give it up.
Whatever replaces captcha, I really hope it's less frustrating for the actual humans
> Turns out that this new algorithm can also be used to read CAPTCHA puzzles—we found that it can decipher the hardest distorted text puzzles from reCAPTCHA with over 99% accuracy.
Am I missing something or could we improve CAPTCHAs by mimicking street numbers?
http://techcrunch.com/2012/03/29/google-now-using-recaptcha-...
Do the street numbers also have a lowered solve rate for humans?
As a result, reCaptcha & co tend to be more of an annoyance to honest visitors than to spammers.
One approach that I enjoyed seeing was the use of reverse captchas. Here you pose a problem that a computer can easily solve, but a human cannot. For instance, if you ask a simple question (1+1=?), but you place the question box off the screen so the user can't see it. A computer would be able to easily answer the question, but a human user would have no way of doing so.
What is expensive? Reputation. That's where credit money's value comes from.
I wrote a more comprehensive piece here, in case anyone's interested: https://news.ycombinator.com/item?id=7601690
Seriously though, I hope this does not mean there will be harder captchas, current ones are already stupidly hard
Now the race becomes who can write the better captcha solver, Google or the spammers? As spammers learn to identify things in the 1%, Google will hopefully improve faster and continue to narrow the "hard to solve" band.
Whenever I see these kind of captchas I switch to audio captchas. It is rather unethical for Google to use recaptchas in this way.
They new addition to the article is that now they have tested the same type of NN on reCAPTCHA, and (perhaps unsurprisingly) it works.
[1] - https://news.ycombinator.com/item?id=7015602 [2] - http://arxiv.org/abs/1312.6082v4.
I'd estimate my own accuracy rate to be 90% or less.
That's one order of magnitude higher.
For me this was quiet annoying to input street numbers of others. It's a privacy issue, it was like helping the NSA spying and one feels bad entering Google's captcha.
What is even more astouning is that Google does not even mention all the croud sourced "volunteers" that trained their OCR database. As Google use an open OCR software (former HP OCR app from '95) it would be a good choice to publish their data back to the community.
I removed Google captcha on my own sites and implemented my own traditional captcha (on the first sight of it about two years ago).
For example for street numbers they not only have picture of a number, they also have knowledge of all the other numbers on that street and guesses for those other numbers. Easy to guesstimate order of a number by checking neighbouring ones.
Same for book words, they have n-gram database. http://storage.googleapis.com/books/ngrams/books/datasetsv2....
Thats a lot of useful MAP/ML data.
But the example they give for the new captchas all look like random crap, "mhhfereeeem" and the like. Its like they are not interested in structure, just pure geometry of letters/numbers.
well, isnt that great? Because I, HUMAN, can maybe solve _one_ of those (lower right one).
I frickin HATE google Captchas and simply close the page if it wants me to solve one, they are too hard for me.