Show HN: I made a tool to convert images of tables to CSV
github.com
github.com
I don't know how OP uses it with images converted to PDFs though, as that would be just like a scan, and ImageMagick doesn't do OCR as far as I can tell.
"Update docs" isn't checked, and that's what I was going on.
As pointed out in this thread, right now it only works with text-based PDFs. But there's a PR[1] which will add OCR support (using EasyOCR) for image-based PDFs in some time.
Now, we've moved onto to ML-based approach to train generic models that can be applied to variety of documents for table structure recognition.
[0] - https://docsumo.com/free-tools/extract-tables-from-pdf-image...
https://docs.github.com/en/github/creating-cloning-and-archi...
Integrating an OCR library is something we always wanted to do.
Here is a gif of table detection for a scanned PDF doc (the first run is slower as it requires fetching the opencv is bundle): https://lh3.googleusercontent.com/-OobUBBtnydg/X6Vn_Ls3juI/A...
Here's a demo of the addon running outside of Google Docs: https://pdftableutil.possiblenull.com/app/
As for the handwriting, I think Tesseract can handle the recognition if the writing is good, but the table needs to fullfil the expected hypothesis. Also the pre-processing can't get rid of a lot of noise so it can be a problem too !
Disclaimer: I work for creator of said service