HNHacker News
TopNewBestAskShowJobs

vortex_ape

2,928 karma · joined May 20, 2016

Staff Engineer @ Clarify. Previously, Founding engineer @ June (YC W21). F20 @ Recurse Center. Working on open-source tools fleurmcp.com, camelot, present, excalibur, and many more.

site: https://vinayak.io

github: https://github.com/vinayak-mehta

submissionscomments
vortex_ape··on A Python Library to extract tabular data from PDFs
Hi plaidfuji! I did try HoughLinesP during experimentation. I vaguely remember (since this was almost 2 years back) getting the actual line segment as a combination of multiple smaller line segments in all cases (which could then be combined to form the actual segment using some heuristic). It came down to getting the actual table line segment out which a combination morphological transformations and cv2.findContours provided (without the need for another combining step).
vortex_ape··on A Python Library to extract tabular data from PDFs
Adobe released a way to attach data tables with PDFs. But I think it hasn't been adopted fully, since many organizations that release open data as PDFs don't tag the accompanying data tables. https://www.w3.org/TR/WCAG20-TECHS/PDF6.html
vortex_ape··on A Python Library to extract tabular data from PDFs
What does this pipeline do and what software have you used to implement it?

I have used Airflow in the past to create ETL pipelines, and plugged in Camelot in one of them to extract tables from PDFs. I also wrote a blog post about it in case you might be interested. https://hackernoon.com/how-to-create-a-workflow-in-apache-ai...

vortex_ape··on A Python Library to extract tabular data from PDFs
Hey squaresmile! Yes, right now table detection with Stream doesn't work nicely if the table is not present on the full page, for which you can use the table_area kwarg from [2].

You should use "pip install camelot-py[all]" to install Camelot (which will install opencv-python too). I had to take it out of the requirements since it wasn't available in any conda channels while I was creating the conda package. I'm looking to remove opencv as a requirement altogether by either vendorizing the opencv code that is being used inside Camelot or reimplementing the code using something lightweight like pillow.

Thanks for the catch in [2], I'll correct it!

vortex_ape··on A Python Library to extract tabular data from PDFs
Ah sorry I forgot about posting PII data online. Thanks for the link!
vortex_ape··on A Python Library to extract tabular data from PDFs
The extreme cases (where there is no/incorrect mapping between the glyph and the character they represent) are a real pain! This mapping is stored as a ToUnicode map inside the PDF. In the past I've used OCR to handle such cases but I'm planning to create an experimental interface where anyone can modify the ToUnicode map. The challenge would be to make the modifications automated/user friendly.
vortex_ape··on A Python Library to extract tabular data from PDFs
The API and docs (on which the blog post built upon) were inspired from pandas and requests!
vortex_ape··on A Python Library to extract tabular data from PDFs
Thanks for the suggestion sandGorgon! Can you also point me to an example of a PDF with signature data?
vortex_ape··on A Python Library to extract tabular data from PDFs
Hey danimolina! You can export the data into an excel by specifying it as the export format, Camelot comes with a command-line interface too! https://camelot-py.readthedocs.io/en/master/user/cli.html#cl...

You can simple do: camelot --output data.xlsx --format excel lattice input.pdf (lattice can be replaced with stream based on the type of tables in your PDF)

vortex_ape··on A Python Library to extract tabular data from PDFs
Note: I had to decrypt the second PDF using qpdf since the library I'm using to split a PDF into pages (PyPDF2) doesn't support the encryption type of that PDF.

Did this: qpdf --decrypt input.pdf output.pdf

vortex_ape··on A Python Library to extract tabular data from PDFs
Hey amelius! Though OCR would provide a generic solution, it would be an overkill for text-based PDFs. I'm working on getting a OCR solution up since there's still a lot of data that is trapped inside scanned PDFs and not text-based ones.

If you have any pointers in the OCR route, do suggest them here, or on this GitHub issue! https://github.com/socialcopsdev/camelot/issues/101

vortex_ape··on A Python Library to extract tabular data from PDFs
I assumed that you're talking about page 33 in the first PDF, since it has only 225 pages. I extracted Figure 6-23 from it and the table on page 45 in the second PDF. Here's a gist: https://gist.github.com/vinayak-mehta/cf30a5560f1b8ab4c0b25e...

Yes, Camelot takes care of cells spanning multiple columns! You can check out the Advanced Usage section for explanation on the keyword arguments I used in the gist! https://camelot-py.readthedocs.io/en/master/user/advanced.ht...

vortex_ape··on A Python Library to extract tabular data from PDFs
Hi berti! I wrote the library and the blog post. Can you point me to some PDFs which have these register maps?
vortex_ape··on Ask HN: What are some of the best documentaries you've seen?
Bigger Stronger Faster*

"... is a 2008 documentary film directed by Christopher Bell, about the use of anabolic steroids as performance-enhancing drugs in the United States and how this practice relates to the American Dream." -- Wikipedia

← PreviousPage 3 of 3