28 karma · joined August 9, 2016
Hope you enjoy this post, and keep tuned as we have other posts to be published in coming weeks about other part of our scanning feature.
If we were to use y=mx+b, then the hough transform image would look like many straight lines intercepting at a few points, which makes most intuitive sense. The issue with this form is it gets ill-formed when the line becomes near vertical (m goes to infinity).
The polar parametrization r=x·sinθ+y·cosθ solves this problem, and in the hough space, the axes will be r and θ. A point in image space maps to a sinusoid in hough space, which is why the transformed image looks like that.
Agree with that 3D deformation is a difficult open problem, and we haven't gotten into that yet. Currently we assumed the document is a flat rectangle, which maps to a quadrilateral in image space. A homography is then applied to rectify it, and it seems to work quite well if the paper is slightly curved or folded.
Yeah, Hough transform is definitely a time-tested algorithm that embodies both elegance and efficacy. I truly love that.
On the technical side, we wrote the detection library in C++, so that it can be easily ported cross-platformly. For iOS integration, we simply integrated with Obj-C++.