Beating Google ReCaptcha and the funCaptcha using AWS Rekognition
bitbucket.org
bitbucket.org
1) ip addresses he uses have probably really good reputation, because even with rather bad image recognition ability it still lets him in.
2) surprisingly enough google's captcha apparently doesn't consider mouse movements at all - the bot selects images rapidly one immediately after another and always clicks in the same place of image.
Yep, it's Google cookie based, not behavior on the page. You can only use the tab and enter key to successfully complete the captcha if you have enough Google cookies.
But what if the device is mobile or tablet (or touch screen laptop) where captchas are solved without mouse movement? I guess that could be the reason why Google has not implemented it yet.
Well, I wouldn't really do that if I were Google. But I think it could be happening.
reCAPTCHA v2 and the badly named Invisible reCAPTCHA¹ assess the user and either let them in, or gate them on puzzle solving. So, it does pass or fail you, where fail means trap you in purgatory indefinitely (though some hold it’s actually hell rather than purgatory).
reCAPTCHA v3 never presents a CAPTCHA for you to solve, but decides a score (in practice, I’ve only seen 0.1, 0.3, 0.7 and 0.9) where higher means it’s feeling more friendly towards you, and it’s up to the site operator to decide what to do with it. (You should provide alternative means, e.g. if you’re doing fraud prevention for signup, fall back to SMS verification. Too many sites provide no recourse for a low score, which is illegal to do in various countries on accessibility grounds. With reCAPTCHA v2 and Invisible reCAPTCHA the site owners can at least blame Google, not that that gets them off the hook.)
Given the use of Rekognition here, I presume this project is breaking reCAPTCHA v2 and possibly Invisible reCAPTCHA, and not reCAPTCHA v3.
———
¹ I call Invisible reCAPTCHA badly named because it’s only invisible on the happy path—all it’s doing is hiding the “I’m not a robot” widget, effectively having the code “click” it when submitting. And given that reCAPTCHA v3 is then invisible, never showing anything to the user… yeah, it’s confusing. Arguably reCAPTCHA v3 isn’t even a CAPTCHA any more either, so… yeah, the names are all a bit of a mess.
but the linked source code is pretty clearly interpreting scores and referencing recaptcha v3, so i'm not sure what images they're running rekognition on. (unless there's a bunch of stuff going on in this repo and this file is unrelated to the title)
https://bitbucket.org/Pirates-of-Silicon-Hills/voightkampff/...
Although I admit my idea looks less probable this way. It's still less ridiculous to me that not using keyboard or mouse data.
But on the other hand they may be afraid of people screaming "Google is sending your keystrokes to their servers" etc.
Step 2: Sell improved AI.
Step 3: Add edge cases to the Captcha that your AI can't handle yet. Go to step 1.
When AlphaZero plays against itself it just gets better. Once that spread is so narrow, every deviation is in the range of possible error. And it can’t explain WHY a certain move is better or why the picture is a squirrel. And sometimes it can be wrong, but only on hilarious edge cases, in the rest we just “trust it” because we dont know one way or the other!!
Once deepfakes have covered every possible thing that could give away the deepfake, what remains is dubious arguments as to WHY something is or is not a deepfake. It could be wrong or not, we would just “trust it”?
I am saying that if there a range past the uncanny valley for adversarial AI then once the generative network is there, it’s game over. They can quickly generate any amount of speech said by anyone for example, and we will not know whether they said it or not. All audio and video evidence would be inadmissible without watermarks. And then we would have to trust whoever made the watermark.
In short - mutually distrusting byzantine consensus or signing will be required for any claim.
Now someone wants to frame you, perhaps because you are a celebrity or a person with money to extort from. So they claim that they saw you murder this person.
Normally, this wouldn't hold any water and no jury would convict you, we are not in the medieval ages anymore. There would be no evidence of you around the crime scene, because of course, you weren't there.
Now imagine that they are able to produce audio recordings of a conversation you had with the victim, screaming at them and saying "you are gonna pay for this". And then they also happen to have video evidence of a smart phone camera that happened to capture you stabbing the victim.
Now tell me, if the jury sees all this and then it gets dismissed, because of "deepfake" claims or whatever, how will it affect the outcome of this trial?
If you are truly innocent, this might still work out alright, unless you are not white. But imagine there is even the SLIGHTEST connection to the real world. It doesn't have to be completely fake. Maybe you had some fight with that victim earlier, or there are witnesses who testify under oath that you had several heated arguments with them, etc.
Deepfakes can change the entire outcome of trials, by biasing the jury and proceedings against you. That's why any audio and video evidence essentially needs to be rejected without a very very thorough forensic analysis of its authenticity. And even then you would still not know for sure.
Also you should see the type of video evidence used in court. It’s grainy 5 FPS surveillance footage that doesn’t even show their face. Deepfakes are not necessary if you want to frame someone.
In a society where anything digital can be manipulated, the veracity of a speech or event depends on a network of trust and people who witnessed the event in real life. If someone made a speech in front of hundreds of live, physically present spectators, these spectators can then verify the authenticity of any recording, and as long as there people in that group that you trust (directly, through your trusted network, or perhaps as journalists of good standing) you can trust that the recorded speech is what was said.
Conversely, a recording of a world leader saying they recommend Coca-Cola to prevent tooth decay wouldn't have any credibility at all if no one credible was actually there to witness it.
- Humans can lie
- Human memory is too fallible to serve as this verification, particularly for every detail
- Small details with big impact can be changed in a verified speech without some people noticing
Of course if someone reliable records the actual speech, any copy can be checked against that. I'm not sure why human memory would be considered superior, as humans have poor memories and can lie.
Ideally anyways. Realistically, we probably just argue on the internet about which talks were real (in this case generated by the person pictured in them).
As an anecdote, I know real humans who do truly believe deepfaked political videos. They were genuinely surprised to know that Pelosi can talk like a normal human.
“A lie can go halfway around the world before the truth has a chance to get its pants on”
Aren't they? I can come up with examples off the top of my head, let alone with a search. This is my favourite[1]:
> The fact-checker for the New Yorker who mistook a Marine veteran’s tattoo for a Nazi symbol has resigned, saying the “small mistake” has ruined her life.
"mistook" is being kind. Take a look at her Twitter feed[2] and tell me if you think she's less biased than the stories that would've passed across her desk, because her feed is incredibly partisan. Do you think she managed to remain impartial while she was still employed as a fact checker? I wonder if she intended or wanted to be impartial, let alone managed it. I also wonder how that kind of person comes to be part of a team of fact checkers, which brings into question their impartiality. How was she able to make this “small mistake” and no one else on that team caught it? Perhaps they lack editorial oversight like Snopes[3], the Snopes that employs people who are obviously biased:
> Of particular interest, when pressed about claims by the Daily Mail that at least one Snopes employee has actually run for political office and that this presents at the very least the appearance of potential bias in Snopes’ fact checks, David responded “It's pretty much a given that anyone who has ever run for (or held) a political office did so under some form of party affiliation and said something critical about their opponent(s) and/or other politicians at some point. Does that mean anyone who has ever run for office is manifestly unsuited to be associated with a fact-checking endeavor, in any capacity?”
I'd say it was unbelievable but it's really not[4].
> Snopes appears to be actively engaged in an effort to discredit and deplatform us.
Dropping one's scepticism because someone calls themselves an impartial fact checker strikes me as being in the same realm as following the orders of a complete stranger because they have a uniform and claim authority over you or a situation, i.e. unwise.
[1] https://nypost.com/2018/06/26/new-yorker-staffer-resigns-aft...
[2] https://twitter.com/chick_in_kiev
[3] https://www.forbes.com/sites/kalevleetaru/2016/12/22/the-dai...
[4] https://mailchi.mp/babylonbee/reparations-for-everyone-83635
Won't be perfect, but it would be a hurdle. Also, for things like audio and video if you know the location on the planet then there are certain artifacts that are hard to fake with mere AI. The way sound bounces off walls in a room or shared environmental noise that other devices in the area would detect, like a plane flying by.
If X is CNN, then half the population will not believe it. If X is Fox, the other half won't believe. Bottomline, the world will be truly living in virtual alternative universes and it doesn't even matter what truth is as long as everyone in that world believe that.
That literally accomplishes nothing. It would be a matter of time before someone just removes the light sensor (CCD? or something) and has a custom made encoder that can write the input data as though the device actually records it live.
Imagine two court cases. In one, there is a video signed with a key from a phone that shows no signs of compromise. In the other, there is a video with no signatures at all. Which is more trustworthy? It's a continuum, not a binary.
Edit:
And also, we can scale trust. There should be more webs-of-trust in the world. I should be able to tap my phone next to my buddies and it should create a link between us. We should have the concept of trust online when we need it.
With deepfakes, in general I agree but it is a bit like computer security. It may be no different from reality but it may still have marks of the network that generated it. Ways to detect it may come out slowly, like computer bugs, with some entities having incentives to research them heavily and not release them.
EDIT: just thought of something else, for videos and pictures increasing the resolution dramatically will make it harder to generate deep fakes, it might be interesting to see how much better cameras get and I’m sure that the amount of compute to fake a higher resolution image will be drastically higher. There are a lot of factors at play.
We then release the tech to the wild without any proposed way to address the problems, instead saying “the problems this causes are outside the domain of CS, so I won’t bother considering them – besides, if I didn’t create these problems someone else would have”
You're definitely right about how little we attend to the consequences of our advancements, especially in the field of AI/ML. This area is fraught with ethical and moral issues that have taken a backseat to the ideal of progress and I think we are making a mistake there.
But otherwise delaying progress what if in the end a more powerful A will eventually make any B and C irrelevant? All you’re doing is kicking the can down the road maybe?
Except this was time consuming and couldn't be done automatically in an undetectable way.
> human societies have survived without video evidence since forever
Stupid argument, we've also survived war since forever, except now a war between nuclear powers would be devastating in a way war historically wasn't. It will absolutely not be the same as it was before photos and videos.
https://www.vice.com/en_us/article/7kzj8y/why-joe-bidens-vir...
Unfortunately that's exactly what is happening with ReCaptcha checking for Google cookies.
Now I can be denied access because of an opaque score, with no way to get around it.
Yay, progress!
How would it make that determination?
What percentage of websites that you use is that, exactly?
That's fair. I think your solution would work for those.
https://continuations.com/post/180985743645/world-after-capi...
It also has the ability to identify visual cues, which this doesn't seem to do.
if we want to defeat time-wasting and privacy-invading captcha and the murky ethics aren't a concern, then we should go to every e-commerce site that employs them, with full crap-blocking privacy mode on, load up on products/services, and then abandon our carts at the captcha.
maybe we can have botnet operators take this idea and run with it, despite the murky ethics.
>we should go to every e-commerce site that employs them, with full crap-blocking privacy mode on, load up on products/services, and then abandon our carts at the captcha.
I've been doing that for years [well, not the shopping trolley bit] on websites which throw fullscreen modal announcements in my face. The second it happens, I click off that website. I dream that the webmasters in question will have some advanced stats setup that tells them that, as soon as they digitally shat in my face I abandoned their sites. But, unfortunately I know it's an empty gesture.I thought the credit card number was the answer to the captcha on an e-commerce site.
[1] https://stackoverflow.com/questions/23314528/how-to-reduce-r...
In my (limited) experience they serve you "hard" ones, and then continue serving you "hard" ones until you give up and go do something else.
All they have to do is give out impossible captcha's.
I suppose there are a bunch of privacy and corner cases that would mess that up (for instance, composing a reply in a text editor and pasting it would be indistinguishable from a bot).
Has anyone tried? Are there write-ups I could read about?
Go to the channel for the other 3.