HNHacker News
TopNewBestAskShowJobs

yxiongdropbox

28 karma · joined August 9, 2016

Software Engineer at Dropbox.
submissionscomments
yxiongdropbox··on Fast and Accurate Document Detection for Scanning
The learning algorithm we used is not a neural network that got trained in end-to-end fashion. Instead, it is a local prediction model that takes an input image patch and produces a patch of the same dimension with probability for each pixel of belonging to a document boundary. Those per-patch predictions are then aggregated together to reduce variance, resulting in an edge map of the same dimension as the input image.
yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Hi everyone, this is Ying Xiong from Dropbox, and I'm the author of the blog post. Feel free to let me know if you have any question, comments or suggestions.

Hope you enjoy this post, and keep tuned as we have other posts to be published in coming weeks about other part of our scanning feature.

yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Very good question. As stated in the blog post (one line above that figure), we actually used a polar parametrization r=x·sinθ+y·cosθ than the slope-intercept version y=mx+b.

If we were to use y=mx+b, then the hough transform image would look like many straight lines intercepting at a few points, which makes most intuitive sense. The issue with this form is it gets ill-formed when the line becomes near vertical (m goes to infinity).

The polar parametrization r=x·sinθ+y·cosθ solves this problem, and in the hough space, the axes will be r and θ. A point in image space maps to a sinusoid in hough space, which is why the transformed image looks like that.

yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Yep, we turned RGB into LUV space before extracting edges, which helps a lot on contrast and keeps essential edge information that could've been lost if converted to grayscale.

Agree with that 3D deformation is a difficult open problem, and we haven't gotten into that yet. Currently we assumed the document is a flat rectangle, which maps to a quadrilateral in image space. A homography is then applied to rectify it, and it seems to work quite well if the paper is slightly curved or folded.

yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Indeed, this is a deceptively really hard problem that I think nobody perfectly solved yet. The main problem with the deep learning route is it being resource demanding (both computation and memory expensive). Hopefully these problems will automatically go away in a couple of years as the mobile devices become more powerful and the deep learning architecture gets more light-weighted.
yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Glad you liked the post!

Yeah, Hough transform is definitely a time-tested algorithm that embodies both elegance and efficacy. I truly love that.

On the technical side, we wrote the detection library in C++, so that it can be easily ported cross-platformly. For iOS integration, we simply integrated with Obj-C++.

yxiongdropbox··on Fast and Accurate Document Detection for Scanning
Yep, we do the entire document detection and other following steps (to be described in coming posts) on the mobile device.