TextCaptcha: Simple Textual Captcha Challenges
textcaptcha.com
textcaptcha.com
This might sound like a nonexistant issue, but lets take the example of a country like India with a huge number of internet users. There are many who use the Internet without being completely literate in English or even their mother tongue. Use cases like checking exam results, train tickets etc require captchas to be filled and the normal deformed letter ones do well. I doubt text captchas will be accessible to them.
But then again, user base matters a lot too and that should be taken into account always!
I had a user struggle to find the exclamation symbol on the keyboard.
Because it seems to use a fixed set of question types. All of which a simple script can solve.
Example:
What is Mark's name?
Writing a script that extracts the name from this string is trivial. And the other question types are similarely easy.The next step would be to allow the site owner to put in domain specific questions. For example a chemistry website might ask "What is the formula of $moleclue?". And then provide ['Water'=>'H2O','Salt'=>'NaCl'...] as questions. Then the site would be protected from a general textcaptcha solver and bots would have to be customized for each domain.
This is in contrast to ReCaptcha, which is used my millions of sites, so using crowdsourcing and systems to make them easy to use is worth it there.
For a long time I used the fixed captcha, "What is 1+1?", but written in my local language, and managed to filter out all the spam...
Not suggesting it's broadly useful, but cool nonetheless.
Why bother? All the hashes I tried were able to be instantly cracked online by CrackStation.
If the hashes are kept server-side, then why not use plaintext?
Keeping it in plaintext would have made doing this being dumb more obvious. Heck, two separate files would have even been better. Also you can do something like damaru levenshtein [1] on plaintext so you wouldn't have to be a stickler for exact matches
[1] https://en.m.wikipedia.org/wiki/Damerau%E2%80%93Levenshtein_...
If you generate the questions based on fixed rules, it is only a matter of time before an algorithm or a human figures out all the necessary patterns.
I personally feel like these text puzzles aren't as energy-draining to solve as reCAPTCHA, so a human could reasonably solve many hundreds in an hour of free time or so. That doesn't even include having a TTS engine dictate the question to Google Assistant.
Captchas in general are a bit of an illusion really. The best Captcha you can have is a home-made one as the attacker has to go through actual manual effort rather than enabling --solve-recaptcha flag on the bot script.
All you "protect" yourself from is casual script users and script kiddies which really can be solved by IP rate limit. If someone has access to thousands of IPs they can probably afford to drop $5 to solve the captchas too, right?
If there wasn’t, people wouldn’t use captchas
$5 per request is not a negligible amount of money. In practice it doesn't cost anywhere near that amount to call a MechanicalTurk API which will solve ReCaptcha for you. But it's still significant for any nontrivial number of requests, such as in the use case of scraping.
You should adjust your priors here. You're focused on the narrow case where a win condition is achieved by spending n dollars to solve a single instance of ReCaptcha. People who use ReCaptcha are (in my professional experience) overwhelmingly more focused on requiring ReCaptcha to be solved for every individual request of a given type.
I have been in the position you speak of, where I had a revolving set of IP addresses, requesting servers and user agents, and $5 per request would have immediately shut my operation down. As it was, the actual ~$0.15 per request to solve ReCaptcha was sufficiently significant that I couldn't curate enough data for what I needed, despite having all the other resources you mention.
It's dead easy to get around captchas unless you're just a casual scripter that wants to `wget` that one article - then fuck that guy, right?
The hash of the answer is just a string that concatenates the answers, and the challenges always mixes them in different orders. One possible example:
1. Does Red combine with Yellow to make Green or Orange? 2. If the answer to the last question was reversed, what letter would it start with? 3. If you added that latter to the end of these words, which word would be most edible? Mac, Nam, Pi, Snak
answers: orange,e,pi
Of course, you'd want to design the interface to support multiple choice selection.
I would also recommend against using MD5, since, even if the hash weren't known to the end user, which should be sufficient even in this case, the attack MD5 is most known for is the fact that it's trivial to generate text that could match any given hash, regardless of what was used to originally make it. It seems like a potential attack vector somehow, depending on the case, and it's not terribly harder to just use one of the many tried-and-true, known-not-broken cryptographic hashing algorithms. SHA-256 would be adequate.
I'm not 100% certain of all the logic behind this, there's always cases I'm not rigorous to consider, but I'd be interested in seeing how others might improve this approach in similar ways.
Source on this?
Wikipedia states what I have heard before which is that MD5 collision attacks are pretty trivial now, but carrying out a preimage attack as you describe remains theoretical at this time.
Actually I don't know why there is some hash used at all. According to the example, answers are stored in a session. CRC32 would do the job as well. Or no hashing at all. You would need some better hash in case when user downloads it. I can imagine different flow where you would need some better hash: I.e. you have some secret token, hash it together with captcha answer, send question with good hash to a browser and user sends back answer together with a hash he got. In such flow there would be no need to store values in a session.
I've posted this a couple months ago. Just before christmas, I've received an email from the mods asking me to repost, as they thought the story was interesting. Initially after reposting, it didn't get much traction, but, when I look at it now, it actually has upvotes. However, the number of upvotes it has is much bigger than the amound of Karma i received. I wonder how correlated this facts are.
However not everybody is very determined. There are for example bots that are using the comment forms to send spam. The expected payout for each posted message must be really, really low.
This could be useful for example in the comment form plugins created for various content management systems.
W3C has put together a comprehensive document about captcha types and their application: https://www.w3.org/TR/turingtest/
I'm curious how TextCaptcha would fare in terms of complexity compared to the other language understanding benchmarks.
That's a somewhat concerning choice of hash function. MD5's collision resistance is broken. It is well known to be broken, for more than a decade.
There's references to PHP on the page. 10 years ago [0] this message appeared in PHP's manual:
> The well known hash functions MD5 and SHA1 should be avoided in new applications. Collission attacks against MD5 are well documented in the cryptographics literature and have already been demonstrated in practice. Therefore, MD5 is no longer secure for certain applications.
So yes, I'd say that's a problem for this use case.
In this case, the only threat model might be brute forcing the answer, but that applies to SHA-2 as well, since both are designed to be fast so that you can hash gigabytes in reasonable well. For that, something memory-hard such as Argon2 should be used.
PHP is a red herring; this would apply for any language.
I used the warning from the PHP manual, because the site creator should be familiar with it.
You'll find similar warnings against MD5 everywhere.
Search Google for f6f7fec07f372b7bd5eb196bbca0f3f4, and you’ll see the answer is Friday. That stateless PHP example is rendered useless by this.
"How much is one + (image depicting a 6)?
It would certainly lose utility for visually-impaired users.
OCR is pretty well advanced by now; the image captchas that used to rely on it now use extremely distorted images.
> How much is one + https://i.imgur.com/VbyiURo.png
Examples:
- The man couldn't lift his son because he was so weak. Who was weak?
- The firemen arrived after the police because they were coming from so far away. Who came from far away?
- The drain is clogged with hair. It has to be [cleaned/removed]. What has to be [cleaned/removed]?
They are crafted in such a way that there are two variants of the challenge every time. Ex: The man couldn't lift his son because he was so heavy. Who was heavy?
I don't know if they are hard to generate automatically. They do require common sense to solve and it's still an open problem.