Some of the scanned forms have marginal notes (probably okay to ignore these), but I've noticed some cases where the form had prefilled data crossed out with handwritten data replacing it.
Please let me know if a system exists that can deal with this idiosyncratic mess and coaxed to produce structured data of some sort without basically using MTurk behind the scenes. I would be interested in re-releasing the data into the public domain.
Our requirement was that the the software would need to read the number of separating lines in the table and output accordingly.
Maybe it can do this, idk. I was an intern at the time.
We were using tesseract for our attempt too.
Or said another way, it can be fun and also effective to "Do Things That Don't Scale" [http://paulgraham.com/ds.html].
The bottom line was that we were able to achieve very good results consistently, but the margin of error was not acceptable for our specific use case.