How Google Cracked House Number Identification in Street View
technologyreview.com
technologyreview.com
More interesting is the other change: reCaptcha now tries to detects real users and then only generates a simple number captcha for the classical part. Likely they are using your Google Account cookie or Google Analytics for that.
Brilliant, really.
Edit: thanks for the downvotes for apologizing for an errant downvote ... I guess ...
"reCAPTCHA is a free CAPTCHA service that helps to digitize books, newspapers and old time radio shows."
"Currently, we are helping to digitize old editions of the New York Times and books from Google Books."
They are lying and they are not benefiting the public interest - which is what reCAPTCHA relies on - they are only benefiting their company and their own shareholders.
This line gives it away. If they can remove that assumption I will be impressed, otherwise I would say they reinvented the wheel and probably could have used something off the shelf and got similar results.
"To start off with, Goodfellow and co place some limits on the task at hand to keep it as simple as possible. For example, they assume that the building number has already been spotted and the image cropped so that the number is at least one third the width of the resulting frame. They also assume that the number is no more than 5 digits long, a reasonable assumption in most parts of the world."
2. Spotting the building number, as the paper says, is taken care of by a different algorithm.
I'm not sure why this is also a big deal, since spotting the building number is "not the hard part" in most cases.
3. Your assertion that they could have used something off the shelf seems directly contradicted by the fact that the paper says nobody has ever published a multi-digit simultaneous recognition paper.
So i'm very curious what this "off the shelf" thing would be. Could you elaborate?
That said: The GP has a point, imo. I don't doubt that the paper describes something new and interesting (and all engines I work with during the day do segmentation/recognize character by character), but localizing a region of interest is usually the hard job for me. When I identify the right region and crop it/scale it/rotate it .. my job's "easy" and I can run a multitude of generally good OCR engines (off the shelf, if you will) and get decent results (maybe vote a bit, use engine A to segment and engine B and C to recognize the characters etc.)
So .. ignoring the 'we trained a neural network' part (which makes me nod thoughtfully and mumble 'whatever they did there..'), which I _understand_ is the interesting thing here!, they did more or less what I do all the time. The preceding algorithm is what I do quite a bit less often and which in my environment is more interesting and often challenging.
Then again, it can always be labeled as PEBKAC I assume :)
I'm trying to OCR license plates from random videos (stable, but essentially random camera location that sees cars going by and stopping - and I would like to read their plates). I've tried every commercial offering under $4000/license, and -- even when manually localized and a single photo selected, I'm getting less than 95% on a single letter/digit, which translates to ~70% for full plates. (Humans get >99% for full plate on those photos).
Where would you recommend I look next for a solution?
Regarding your particular problem: No idea about international license plates (or the one you are interested in). It helps a lot to restrict engines to character sets. German license plates are roughly [1] ([A-Z]{1,3})-([A-Z]{1,2})(\d{1,4}) which helps a lot.
Usually you try to combine 'dumb' OCR with datasets/fuzzy matches to rule out errors. Only you know if that is possible for your dataset.
Depending on whether I understood your problem correctly we might again have the localization issue (which I complained about above): Find the licence plate, crop and rotate it. Bonus points if your images might contain multiple license plates and a human operator would 'obviously' see the right one..
Recognition itself should be okay: Limited character set, a limited number of fonts (here: One only) and hopefully decent binarization opportunities (here: black on white, background reflective).
Feel free to shoot me a mail, details in my profile.
1: From memory, might be slighly inaccurate, sample only
When I get some breathing time, I'm going to try the latest Tesseract again (when I last tried, it was v2, and its performance wasn't good).
I'm reasonably sure that you're not in the market for the things we do, but if you like to chat/want to talk about the project you're doing: Again, feel free to drop me a line. I .. really wouldn't even think about selling things or something, I'm mostly curious about fellow HN users' projects and have an affinity for IL.
Good luck!
On buildings, building numbers tend to stand out dramatically from surroundings. Otherwise, they'd be useless as building numbers :)
- cars, with numbers (license plate? advertisement, a la "who do you call...")
- billboards, with numbers
- multiple numbers (some buildings are on street a, 1 and street b, 42)
- a lot of noise (garden obstructs most of the house, number isn' that special anymore)
Your next task is to find houses.. I don't disagree with you, but you make it sound more easy than it is, in my opinion. Again: Getting a decent start is probably no issue, but competing with human operators...?
You can watch this talk to get an idea for what they're about: http://www.youtube.com/watch?v=vShMxxqtDDs
Traditional OCR pipeline would be to use some heuristics to find line of text in an image, use some other heuristics to break line of text into candidate characters. Some candidate blobs may need to be merged to make a single character, so you use a separate character classifier pre-trained on correctly segmented characters to score the candidates, and then Viterbi/A* search on those scores to find the most likely interpretation of input.
Many problems with this -- how do you tune the heuristics? How do you recover from error in an earlier stage of pipeline? How do you get character level ground truth from image/text pairs?
With enough engineering time, you can solve those problems, but it's a lot of coding and tweaking. The point of the paper is that you can skip those steps and read OCR output directly off top layers of the network.
I thought there was more to it because the headline made me think that the general problem of taking a random StreetView location and finding the addresses in it was solved, which really set some high expectations.
This work is equally impressive and good stuff. Hope to see this headline again in the future soon.
(I wouldn't call it blogspam, because it looks like they interviewed the researchers, but the summary leaves something to be desired)