Looking at this guide, it's easy to understand the steps. But I don't get how one actually would apply them in code to a picture of a QR code. How?
Looking at this guide, it's easy to understand the steps. But I don't get how one actually would apply them in code to a picture of a QR code. How?
I built a similar thing for Telegram a while ago[3] for recognizing the text in the machine-readable zone of passports and ID cards. It does different things after the initial detection/transform/binarization steps, but those would be the same for a QR code reader.
[1] https://en.wikipedia.org/wiki/Hough_transform
[2] https://en.wikipedia.org/wiki/Edge_detection
[3] https://github.com/DrKLO/Telegram/blob/master/TMessagesProj/...
I think I have the general idea on how all that works, but... isn't there an official spec/algorithm on how one is supposed to do this? The devil is in the details.
Some more things to include:
- perspective correction may be needed (the surface with the QR code is not parallel to the image plane, so pixels further away get smaller)
- if the algorithm takes parameters, such as color filters / black-vs-white thresholds, try with different parameter sets in sequence. The user is pointing the camera at the code for a relatively long time, compared to the time it takes to process an image.
- pixely images should produce a recognizable pattern in the fourier transform of the image. This gives you the size AND rotation at once.
(edit: formatting)
Disclaimer: I an the author of STRICH (https://strich.io) and have dived relatively deep into the topic.
Basically the entire of subject area of electronics and signal processing exists to find clever ways to solving all the problems you mention. So for people that do for a living, there’s a huge set of standard algorithms and approaches you can apply to solve the problem of “detecting QR codes”, in the same way there’s standard approaches to sorting a list, or building a B-tree, or creating a query planner.
It’s realistically not possible to summarise the process into a HN comment, because that would require explaining decades of signal processing research in a handful of sentences.
To give a flavour for solving this problem, QR code alignment markers are designed to be easy to detect, regardless of angle or being partially obscured. Their large simple pattern means you can use crude algorithms to find all alternating back and white patterns in your image, then analyse those patterns to detect which ones are noise and which ones might be alignment markers. Stuff like the regular spacing of the alignment patterns, plus the timing strips between them, give you anchors to rapidly check if your looking at a set of actual alignment markers, or just things that happen to have the same shape as alignment markers. At which point you can start the expensive process of attempting to decode the QR data.
As for what algos to use to go from colour to greyscale, well that’s up to you to figure out. Building a crude QR code decoder is “easy” building a fast robust one is hard, and quite valuable. Nobody is gonna give you that kinda secret sauce for free, not when they can charge you for the hundreds of hours of R&D involved.
There’s an official spec for QR codes, it might even give you basic guidance on how build a very basic QR code decoder. But asking for a detailed spec on a fast robust QR decoder is like asking the IEEE for a detailed spec on how to build a 10Gbs Ethernet controller. It ain’t gonna happen, because building such a thing is hard, even once you know what signals you’re decoding.
Why is that the case?
Nobody expects a detailed tutorial, but it's hard to even come by a list of techniques that are known to work (or maybe I'm just trash at search).
But just like Reed-Solomon encoding is a very specific algorithm, with many different applications, of which QR codes is one. What you’re looking for is similar, it’s a huge set of very specific algorithms with huge set of applications, of which one is QR codes.
QR codes are simply too niche for anyone to have put together a pubic document on what exact image processing techniques you might use to decode them. If you want to teach people about 2D signal processing, then QR codes are probably not a good starting point.
I would also argue that Reed-Solomon encoding is not “deep” anything. It’s one of many different 1D signal error correction algorithms that exist. For people who work on signal processing as day job, Reed-Solomon is about as “deep” as quick-sort is to programmers.
As for why quick-sort is well documented with many open implementations, and QR codes aren’t. I would argue it’s simple due to the industry they developed in. QR codes started in manufacturing, being used for inventory management, almost certainly by embedded and electronic engineers. All of those industries pre-date the open source movement by decades, trade secrets are still important for them, so they’re not naturally inclined to share IP.
To determine orientation scanners look for “alignment markers”.