If you want to do text extraction, look at things like Stroke Width Transform to extract regions of text before passing them to Tesseract.
If you want to do text extraction, look at things like Stroke Width Transform to extract regions of text before passing them to Tesseract.
There are samples here [1] and here [2] to get you started. The paper is here: [3]
---
[0] http://docs.opencv.org/3.0-beta/modules/text/doc/erfilter.ht...
[1] https://github.com/Itseez/opencv_contrib/blob/master/modules...
[2] https://github.com/Itseez/opencv_contrib/blob/master/modules...
No, a whole lot of pre-processing would be needed. It all depends on the exact layout - if your tolerances are tight you need much more logic than if you have, let's say, 2cm white space around one sentence you're after.
In summary, you need templates that map the field positions -> meaningful keys so that you can get back useful data as json/csv/xml. I have some tools that are still being polished that automate much of the template creation and do a lot of the pre-proc for you.
email is my username (at) gmail
We open-sourced the library that we use for exactly that purpose: https://github.com/creatale/node-fv