Breaking the Silk Road's Captcha
github.com
github.com
One observation: training a neural net to classify segmented characters is probably overkill. The author observed that the font never changed, but never ended up exploiting this fact. After the very effective preprocessing, thresholding, etc the characters are almost identical to the 'average' representations the author generated!
I bet it would be enough simply to classify an unknown character by the letter that it shows highest correlation with.
Also, given the successful character extraction, and the knowledge that (1) the font doesn't change, and (2) they're just translated and rotated, I think performing those operations on the individual characters could've yielded a pretty perfect success. Simply try a whole bunch of shifts and rotations on a given character until it matches a reference, almost exactly.
In retrospect, I'm sure there was a good deal of more focused, surgical solutions that'd save some overhead.
This is one of the unfortunate "math favors the bad guy" consequences in a lot of anti-abuse filtering tasks. (Anti-spam research has similar problems, which is why the main innovation wasn't making filters better but radically increasing the cost of getting caught, via burning the reputation of the offending IP. IP addresses are a lot more expensive to acquire in quantity than packets.)
I've created many similar programs to defeat captcha's. I would classify this as a medium severity bug, you would still need to brute force the passwords on a terribly slow and intermittent connection.
https://github.com/dawjan/Open_Me/tree/master/Captcha%20Crac...
Also /.. is php tor