How to break a Captcha system in 15 minutes with Machine Learning
medium.com
medium.com
This is the key to doing this in the easiest way possible. The training set is created almost automatically!
/s NNAlphabetPopulationSentimentBOT 0.1alpha-3zulu-beta-5-mark20
[1] We are getting better all the time at tracking activity in private/incognito tabs, but there are still some gaps particularly when users manually enable extensions like ublock in incognito mode. This hurts the user's experience on the web, and we suggest only enabling google extensions in incognito. It's incognito even without privacy-enhancing extensions. We promise. Nevermind that we know enough about you to let you bypass our captchas.
I have notice more recently at I am asked to identify object(signs, roads etc.)using multiple iterations. It's not unusual to be asked to identify these on 3 separate iterations. This is despite identifying all cells correctly on both the first and second passes.
I don't believe it was always this way. What is the reason for this? Is there some heuristic in the captcha program that decides to ask for further identification or is this randomly generated?
My guess is that either (1) your first attempt at validation didn't match what their automated system identified so it didn't have high enough confidence to proceed or (2) the first image is one you are classifying from scratch but the second/third image is the one that you are only validating against their guess.
In either case, it's amazing to me that Google is able to get the whole world to do unpaid menial work for them by offering a free product to website owners.
Its an example of the https://en.wikipedia.org/wiki/Principal%E2%80%93agent_proble....
In fact I want to know what other people choose to do for those one. I'd like to guess that Hacker News types would disproportionately choose the latter option.
- Do the posts or supports of the signs count as part of the signs?
- What if it's a sign but not a street sign?
- If the camera is looking across a road, but the pavement isn't visible, does that image contain a road? I suspect humans are answering that differently, but it's not as bad as the sign one.
- Storefronts. Sometimes I can't tell what I'm looking at. There's a building with lettering on it, but I can't tell whether it's a store much less a storefront.
Google needs to add a tutorial for humans on how to answer those captchas, because there's not enough context for us to figure out what google really wants.
If they'd dogfood this and require their employees to solve a captcha to login to their workstations, this problem would have been solved yesterday.
I'm surprised, since google bought youtube in 2006, that google doesn't let people add translation subtitles to videos on youtube, which is something people have historically been actively willing to do for free, compare multiple submissions against each other to gauge accuracy, then map that to the audio of the video to train and machine based language translation datasets. Always seemed like a killer way to leverage what people would do for free to advance the state of the art.
The round trip time of those services can be quite long, narrowly within the CAPTCHA timeout window. By forcing some users to solve three iterations of the puzzle (the exact instructions are “stop when there are no more images containing <cars>”), Google forces multiple round trips to a CAPTCHA solving service. Also, the solving service and its client need to account for the expectation that the final image will have no <cars>, increasing complexity.
This increases cost 3x, slows solving rate 3x, etc. Ultimately google can never stop the problem of outsourcing CAPTCHA solving to a human, so their best option is to increase the cost and complexity of doing so.
One co-op term I wrote a simple VBA script to open up Google and search for some addresses. It was a small 2-3 day research project. Well unfortunately for me, I subsequently had to solve captchas every day for the next month and a half, and although about half of them would be instantly solved, when I had to redo CAPTCHAs I could feel my sanity slipping away.
They don't allow even mild automated searching, the price per thousand searches is prohibitive and you can't get past the 1000'th result so you need to use combinations of words and other imperfect means to go deeper.
I would have liked to do deep web searches in order to build datasets for ML, but unfortunately there's no search engine (or is too expensive) for bots.
It's 2017, years after the Snowden leaks, and we still use Google and don't have a f*#ng search engine for creators, we only have search engine for 'customers'. We need our own search engine for many things - privacy, control over filtering, mass search, deep search, for creating novel projects, to be sure Russians haven't tampered with it, and so on.
I'd say Google's refusal to accept bots is akin to a lack of net neutrality and places limits on what kind search based open source projects can be created - I know it sounds entitled, and maybe I should just build my own crawler and index - but you don't solve a more complex problem as a preamble to solving a simpler problem and we don't want thousands or millions of private search engines pulling the same content and duplicating indexes - we still need a community solution.
If they love net neutrality, they should support search neutrality as well, for the maker community. It's the same argument as for net neutrality: to protect innovation from incumbents. Search is like the last mile and Google is like a nasty pipe throttling ISP.
FWIW Bing has an API.
Can you elaborate on what you mean by price per 1K searches and a 1K search result limit? Is this a B2B Google offering for search? Is this offering posted somewhere?
The labor cost can obviously be outsourced(or not) and it makes sense to charge as a service to sites.
However,this makes me ask one more question - what if popularity in ML creates meanial jobs for humans similar to my theoretical solution? Everyone looks at the benefits of ML,but all good technologies get abused and it's hard to imagine even more ML countering abused ML.
It would need 1) Not easily searchable 2) Not math-based 3) Easy 4) Internationalizable
I would see "spot the grammar mistake" as one that furfills 1-3, but can't figure out oone that does all 4.
In college I wrote a term paper on breaking Microsoft's captcha (which is a little harder but not by much) twice: first with a simple template-based classification method and then a CNN approach.
https://www.dropbox.com/s/jfp5xbv3eh589f6/6_857_CAPTCHA.pdf?...
At the end, we go over approaches that would help captchas fight attacks. I think the quick flickering approach would work best (split the image into uneven parts, flicker them quickly so the human eye can read the aggregate image but any single slice doesn't show the full picture, and the superimposed image is incorrect)
One of the challenges here (which I'm sure you are very aware of) is that perception tricks that fool computers like flickering images also can block out users with different types of visual impairments. Sometimes users with even minor or infrequently-symptomatic visual impairments won't be able to read an image[1] that uses a special "trick" like this.
For example, consider the risk of triggering an epileptic seizure with flickering. At a certain point it becomes an accessibility/legal issue.
[1] The animated example from nytf3's paper - please note that in contains strong flickering: http://people.csail.mit.edu/recasens/images/captcha.gif
Yep, it generates 4-letter CAPTCHAs using a random mix of four different fonts.
And we can see that it never uses “O” or “I” in the codes
to avoid user confusion.
this seems to be a very simple CAPTCHA system to beat from a ML problem perspective, right?I have to go through several attempts to verify myself because Google didn't like that I missed one box with a car in it and have to start all over looking for street signs.
But funny enough, I believe that data is used to help with machine learning to identify the objects, so we are CAPTCHA'ng ourselves into harder and harder CAPTCHA's.
The "I am not a robot" types are interesting because they don't just test you, they monitor your behavior. "If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck."
Are Captcha's a paid service and my examples are sites buying into only the lowest level of ID?
Or is this something akin to a decision made vs my access frequency/behaviour where the algo decides it's reasonably sure I'm a human so no need to waste bandwidth on sending and confirming images?
Dunno what to call them, but I consider them educational interactions with AI [0]
In fact, one of the most effective attacks against Google's ReCAPTCHA is more simple than anything in this article - you just request the audio version of the captcha, feed it to Google's own Speech Recognition API (or a competitor), and give Google back it's own result [1]
Telling someone they didn't need to do it with that tool is presuming that a) you know why they did it, b) that you are able to judge the suitability for that of that tool for that purpose and under those constraints, and c) that there is no one else who has ever needed to use that exact tool to solve that exact problem under those same constraints.
The highest tier (again, which you are referring to) includes 800+ pages, detailed experiment journals on how to reproduce the state-of-the-art publications (ResNet, SqueezeNet, VGG, etc.) on ImageNet (which is 1.2 million images). I demonstrate how to implement each model from scratch and then train them, detailing which parameters to change and when. The highest tier is for people looking to train really large networks on massive datasets where you could be spending thousands of dollars in the cloud for GPU costs (you can't train these networks without a GPU, or ideally multiple GPUs). I've also included the pre-trained models as well if people want to get started with them and skip training. This tier is really for researchers/practitioners who need to save time and finances by starting with experiment journals that detail how to replicate the results.
The lower tiers are for people just (1) getting started with deep learning in context of computer vision and/or (2) looking to apply best practices. Each book also includes video tutorials/lectures once I have finished putting them together. Realistically I should rebrand the book as a course as it's much more in line with something you would get from Udacity (only with more theory and more detailed code and implementations).
If anyone has any questions about the book do feel free to ask.
The way I like to buy a book is go to Amazon, look at table of content, read few pages and most importantly read some reviews. Your book currently doesn't even appear in Amazon search (or even Google search). Despite myself being quite active in the field, I had never heard about your book before (I know of at least other dozen books on the subject). I wonder this is why you might have relatively much lower volume and such a high price to make up for your revenue target. I would think putting your book on Amazon would increase your volume by an order of a magnitude (or two) and help reduce price to may be 1/7th or 1/8th without requiring tiered pricing (which again is a huge turnoff) while increasing your net revenue actually more than before (probably by an order of magnitude). You might want to look in the theory and economics of price-volume curves.
The biggest problem with your book website is that you as an author comes out as hard-selling hard-charging marketer who wants to maximize profits and make a sell like an old car salesman to anyone who is walking by rather than experienced calm expert for who learning, teaching, academic honesty and integerity is more important than making money. Again, not saying this is who you are, it just feels that way from the style and content of the book website. Hope this helps.
Setting aside issues of style and what other sites are doing it's very clear what the book is about, what's included and what makes that valuable.
The prices certainly are steep compared a "book", but it really is quite a bit more (more like a self study course in ML) and really targeted towards businesses that need to make this happen within their organizations.
If there was a way to tone it down for the referrals coming through Medium from HN (tough to track through 2 sites) or using some is-logged-in-to-HN hack (which would probably piss off people even more if detected) maybe it could be dialed back a hair to catch people like you, but it's probably not worth the effort.
The reply from the author seems to be perfect: https://news.ycombinator.com/item?id=15914307
PS: Typo in the #release_bar: "has been offically released!"
---
[1] http://i0.kym-cdn.com/photos/images/original/000/572/078/d6d...
I expect these sort of systems will move to 'social proofs' - do you have a long-standing account etc.
Simply integrate the code together and generate on the fly. Much faster and simpler!
Unfortunately, I think those mentioned in the post, will be a thing of the past.
Sometimes I think it might be easier to just run a ML algorithm to complete them for me...
1) it's not hackable
2) with every click you contribute to the driverless cars vision improvement.
Interesting I hadn't heard that. Supporting links:
--neural nets captcha analysis--
https://spectrum.ieee.org/tech-talk/robotics/artificial-inte...
--obligatory XKCD--
I am pretty sure he was fast enough to create it in 10 seconds or less.