Using WolframAlpha to Hack Text CAPTCHA
joelvanhorn.com
joelvanhorn.com
They provide an API, but I think this is a case of a project being a "service" to keep the database of questions from being free. There's no technical reason for this to be a service, and it's not a terribly complicated product that would be difficult to scale. It's a static database!
Might be neat to create an open-source bank of these CAPTCHA questions. Maybe I'll throw something together this weekend.
A few years back I was hired by a third party to build a system to break the CAPTCHA on a popular site for various evil deeds. Morals set aside, the money was good and I had a wedding to pay for. A CAPTCHA system becomes quite breakable when it becomes predictable. The system in question used an image based CAPTCHA that used the same (albeit annoying) font for each image, as well as a static distortion overlay and a second set of random distortion. By extracting a thousand sample images I was able to build a system in Perl that could determine the text with an estimated 98% success rate - and when it failed you would just request a new CAPTCHA.
My solution would be to mix up images with logic. I.E.
In the following list of images, which image number contains the green animal: {pic of zebra}, {pic of frog}, {pic of giraffe}
This would require image recognition as well as logic.
The amount of effort for a human to go through the list of answers and come up with answers may be non-trivial, but once completed it's applicable to every single use of the plain-text captcha system. That's bad.
It might be interesting to come up with a methodology for question structure that is harder for algorithms to interpret...?
Just tell me if you need the source code ;P
"2nd fruit in bear apple goat orange" would result in apple because it is looking for second in a list and neglects context of fruit.
"7th digit in abc123def456ghi789 " would result in d when it should be 7. Again not understanding context and merely looking at logical construction.
It barfs on these ones, but not like you predict (it's actually worse). "The 2nd colour in purple, belly, yellow, arm, white and blue"[3] gives back yellow, though, so it's not that stupid.
[1] http://www.wolframalpha.com/input/?i=The+2nd+fruit+in+bear+a...
[2] http://www.wolframalpha.com/input/?i=7th+digit+in+abc123def4...
[3] http://www.wolframalpha.com/input/?i=The+2nd+colour+in+purpl...
Query: The 2nd colour in purple, belly, yellow, arm, white and blue
Answer: yellow
Query: The 3rd colour in purple, belly, yellow, arm, white and blue
Answer: yellow
Query: The 7th colour in purple, belly, yellow, arm, white and blue
Answer: yellow
Query: The bluest colour in purple, belly, yellow, arm, white and blue
Answer: yellow
:-)
http://www.wolframalpha.com/input/?i=what+is+your+favorite+c...
http://www.wolframalpha.com/input/?i=The+2nd+colour+in+blue,...
Q: "Which word contains “z” from the list: zoologist, midwifery, spiderweb, crimps?" A: "zoologist"
But what if you change up that list a bit?
Q: "Which word contains “z” from the list: action, jackson, midwifery, zoologist, spiderweb, crimps?" A: "jackson"
Alpha's sentence digestion has always left me going: huh?
To clarify, there is not a precomputed DB of an enormous number of questions although this would be possible to derive and does occur on a lower level for caching performance purposes. The total count comes from permutation/probability maths based on the question construction algorithm -- when you request a question it generates one which means I can extend the pool quite easily without re-generating a monster cache table.
It is impressive how good Wolfram is at decoding logic, I'll have to have to think about the question construction but I can't make the question too confusing for a real person to solve. As someone mentioned, maybe more abstract questions would be stronger but the difficultly of course lies in generating them. I certainly think logic questions are weaker than a decently obscured/randomised image captcha, but they come with other advantages and work in text-only contexts (e.g. IM-type challenges).
- the CAPTCHA usually wants a single word or number
- the desired word is usually the rarer or later one
So for," What is seven hundred and forty four as a number? ", the interpreted input is a NumberQ function taking the main part "seven hundred and forty four as a number" and evaluating whether it is a number or not. The real result is true.
The zoologist one has already been talked about. The rest other than the 7th digit question are all false.
There are many different choices for the inputs, for example with the colour question
The 2nd colour in purple, yellow, arm, white and blue is?
There seems to be some popularity going on. The first choice as input is yellow and the second choice is blue. To further test replacing yellow with black leads to blue as the first choice. Then again even if you were to use the interpreted inputs you would have to determine the syntax for wolfram which last time I checked is not available and is basically a guess the syntax game.
http://webapps.stackexchange.com/q/1322/40
If someone would care to enlighten me on how this could actually work I would greatly appreciate it, otherwise this method does not seem like it will work. Nice creativity though.
Using your file I've been building up a solver for these questions in Prolog (using DCGs for parsing and simple predicates for the common sense facts and answer calculation). It's nowhere near "done", but it does get 45% of the questions now after only a few hours and 450 lines of code.
The trick to factoring someone else's generative space is to spot the symmetries and build your self a little domain-specific language for explaining those symmetries to your program.
Here's some snippets:
tomorrow(wednesday,thursday).
food(butter).
body_part(arm).
plenty(arm).
above_waist(arm).
ordinal(2) --> lit('2nd').
question( tomorrow -> Answer) -->
p(['Tomorrow is ',token(Tomorrow),'. If this is true, what is today?']),
{ tomorrow(Answer,Tomorrow) }.
question(count_2(P) -> Answer) -->
p(['The list ',token_list(L),' contains how many ',pred(P),'?']).
{include(P,L,Goods), length(Goods,Answer)}.
pred_name('something each person has more than one of',plenty).