AI Model Fundamentally Cracks Captchas, Scientists Say
npr.org
npr.org
- A man is running. A dog is behind him barking and growling. What does the man think might happen?
- A man goes up the stairs to the roof. He walks to the very edge of the building. He takes one more step. What is the man trying to do?
The correct answer should be pretty easy to parse out. And I'd expect a better success rate for humans than some of the captchas today that increasingly are looking more like magic eye puzzles than character recognition. But of course the big question is generation. Can these sort of implication based stories be generated in a way such that the final text can not trivially be reversed to the answer (without even considering the 'meaning' of the question)? And for that matter can these even be realistically generated in mass?
* its back
* its belly
* Leon
(in case the mistakes were not for comedic effect)
Or maybe later Android versions know they should turn the Tortoise back on its feet.
* Mr.
* Professor
(in case you want some appreciation for your nitpicking)
Nitpicking: looking for small or unimportant errors or faults, especially in order to criticize unnecessarily.
See what I did there? We can keep going, but it adds no value. It's tortises all the way down.
If you want someone to learn, help them. If you're not willing to do that leave it be. If you can't leave it be, you probably only want to nitpick.
So what's the correct answer here?
* That's a mean dog
* I hope it's on a leash
* I hope it's not going to start running after me
* I hope I don't get bitten
* Where can I hide?
* OMG, I will get rabies!
[1]: http://science.sciencemag.org/content/early/2017/10/26/scien...
Are there companies relying only only selling captcha for revenues?
http://www.slate.com/blogs/future_tense/2013/10/28/captcha_c...
I thought the world moved on.
Sure, AI can break captcha, but it can be done at scale for far less than an AI research and GPU rig costs.
Google's approach to bot recognition is training their own bots incidentally, so even an adversarial network attempting to bypass it would give it a ton of training along the way to breaking in.
I don't believe it's a job. Isn't this the thing where captchas on target sites are simply mirrored on other sites like sketchy filehosts? Real human users are solving captchas to access some content hosted by this service, and the solution they enter is passed through to the target site.
If they are using filehosts, how would they verify captcha is correct? They can double check but it will lower their capacity and solve times.
I ask because IIRC I've never seen ReCaptcha be used on file hosting sites. (I don't think so, at least. There might be one...)
So people can actually understand the reference.
Dileep George, cofounder of Vicarious, is the former Numenta CTO, and claimed to use probabilistic graphical models as a basis for their tech.
What they mean by fundamentally cracked is that this method seems to be more robust against minor variations of spacing, font, etc. than CNN-based models.
Trust me this is nothing but simple and will take you on a wild ride.
So, given the financial incentives on both sides, I'd like to believe that continually creating and overcoming these tasks is a possible route to AGI.
It's right there in the name: "Completely Automated Public Turing test to tell Computers and Humans Apart"
Computers try to figure out who is human and who is not. In Turing test humans try to who is human and who is not.
At a certain point I just give up and refuse to use the worst sites that use this junk.
Perhaps this is an old news as this technique has been out for a while but I find that it is still relevant in the many cases I have encountered.
Furthermore, in my experience, I attribute Google's failure to improve reCAPTCHA's "I am not a robot" visual appeal as one of the key factors why many organisations are simply not using it.
We can cause the network to misclassify an image by applying a certain imperceptible perturbation, which is found by maximizing the network's prediction error. In addition, the specific nature of these perturbations is not a random artifact of learning: the same perturbation can cause a different network, that was trained on a different subset of the dataset, to misclassify the same input.
[1] https://arxiv.org/abs/1312.6199
EDIT: Also interesting: "Universal Adversarial Perturbations" https://arxiv.org/abs/1610.08401
Other problems include:
* Social security numbers are explicitly not ID numbers
* Not all Americans have social security numbers
* Even if they did, Americans are about 5% of world population
* Not everyone on the internet gets post (say hi to Nairobi, where addresses of relatively rich locals may be “$Name, third on the left behind the petrol station on the highway”, while poorer people have homes that don’t officially exist on roads that don’t officially exist. Some of these people rely on mobile phones for payments, as banks don’t care.
* Not everyone has an ID certificate or card or passport or driving license to provide, and if they did, why not just skip the middle man and ask for that directly like AirBnB (and every international airline I’ve used) does? Or your bank account, like PayPal does — after all, a bank will certify your ID*
(* they probably don’t all, and even if they did, not all people have bank accounts)
Basically, ID is not a perfectly certifiable thing, so you need to design systems to accommodate failure — even if we actually solved it in theory, all it takes is one criminal finding one implementation flaw and a system that assumed perfection would allow grand exploitation.
And if your goal is just “is this a human or a machine?”, well, let me introduce you to the idea of identity theft, and why people stole all that data from Equifax.
The only way to tell humans and bots apart is some form of automated Turing test, hence Completely Automated Public Turing test to tell Computers and Humans Apart.
But of course this isn't available to be used by some random kitten video trading site.
Also why would I want to give up my real identity to a lot of the sites that use a captcha?
Now I knew the theory already, since I spend a lot of time online, but I asked him to explain it to other co workers, most of them very liberal as they say in U. S, they all thought he is joking.
Americans are funny sometimes :)