Ask HN: How would you extract data from scans of hand-filled paper forms?
I've stumbled across a public data set that consists (largely) of scans of paper forms that were filled out by hand, often in idiosyncratic ways.
Besides the usual checkboxes and simple text fields, there are:
* illustrations of different options selected on some forms by being X'd, on others by being circled, and in others by scribbling an identifier for the option (eg "A")
* Supplementary data on separate pages (sometimes with handwritten tables) that are formatted differently for every submitting organization
* Various measurements expressed as both fractions and decimals
... and other complexities, such as the forms themselves changing over time.
How would you solve turning these scans into a clean dataset?