OCRopus: high-quality open source OCR sponsored by Google and used by reCAPTCHA
code.google.com
code.google.com
http://googlecodesamples.com/docs/php/ocr.php
Now integrated into G Docs already.
http://recaptcha.net/digitizing.html
If so, I'm not impressed.
Red herring. I'm talking about the examples pictured.
Do you know _for a fact_ that this is the same software package?
I don't want to waste my time arguing about why it doesn't live up to my expectations, if it's not.
That page features the output of reCAPTCHA and compares it against an unnamed standard OCR.
The standard OCR does poorly, but it's on tricky documents selected to show the benefits of reCAPTCHA. It doesn't say it's this OCR code, nor does that page really tell you anything about how good it is, if it was.
The other thing you might be saying is that you think the reCAPTCHA output isn't very impressive either. As well as the human element, reCAPTCHA claims to use several standard OCRs to process their document and combine the output is some way. It's possible that the Google code is one of those that they use, but if so it's only part of the process.
I would definitely be interested in seeing evidence of it though.
Or maybe they've realized any human computer test based on text recognition is flawed, and so what better way to force the web to upgrade than to make OCR trivial? I rather like this shotgun approach to AI.
2. Captchas have limited output sets and special characteristics, which make using OCR for them both costly and ineffective in comparison to dedicated solutions. Specifically, you can generate as many perfect sample outputs from a captcha system as you want - and then analyse it in ways beyond the standard character recognition.