Demystifying Text Data with the Unstructured Python Library
saeedesmaili.com
saeedesmaili.com
However - if you can expect a certain format beforehand - then Python is better since you can extract higher-quality data (tables, lists) with the appropriate tool.
PDF has been hit or miss, but pypdf has improved in the last couple of years. Depending on the document you'll sometimes get random spaces or nospacesatall.
I realize now staring at this, that I might have broken API a little. You can't do "text = paragraph.text" anymore, but you can do "text = ''.join([run.text for run in paragraph.runs])" instead.
If you're curious at all why it breaks, it's because in the OOXML spec paragraphs are made up of a ordered list of runs or hyperlinks (and hyperlinks can then contain additional runs). The master branch just implements paragraphs as ordered list of runs (and ignores all hyperlinks).
Replace the `qn("w:ins")` in the example with `qn("w:hyperlink")` and that should hopefully work?
I can't answer your question about pandas, though.
Something that perl would excel at, although I used Python. Because Perl isn't as maintainable as Python
I was intrigued by this comment. A JVM solution would also be viable in my tech stack. Would Tika be easier than line processing compiled regexes in Python? I tried looking at the Usage examples but it wasn't clear.
However, I'm afraid it is not there yet. Other libraries like PDFMiner give higher quality outputs and specialized libraries like Camelot are still needed to extract tables as reasonably well formatted text. It also needs a lot of extra tooling for web scraping. Sure it can read plain HTML from a URL, but it cannot run JavaScript, or control things like User Agent. It could be argued that such features are not within the scope, but it is rather bothersome for a library that presents a magic `partition` function for most standard text sources.
I'm sure it will get there soon though. It shouldn't be hard to integrate with state-of-the-art parsers and tooling, and the simple API undoubtedly brings a lot of peace of mind.
pandoc --to json file.docx
or in Python: import json
from sh import pandoc
doc = json.loads( pandoc("file.docx", to="json").stdout )
Example output (reformatted slightly to reduce number of lines: {'pandoc-api-version': [1, 22, 2],
'meta': {'title': {'t': 'MetaInlines',
'c': [{'t': 'Str', 'c': 'The'}, {'t': 'Space'}, {'t': 'Str', 'c': 'Title'}]}},
'blocks': [{'t': 'Header',
'c': [1,
['first-chapter', [], []],
[{'t': 'Str', 'c': 'First'}, {'t': 'Space'}, {'t': 'Str', 'c': 'Chapter'}]]},
{'t': 'Para',
'c': [{'t': 'Str', 'c': 'I'}, {'t': 'Space'}, {'t': 'Str', 'c': 'like'}, {'t': 'Space'},
{'t': 'Emph', 'c': [{'t': 'Str', 'c': 'cursive'}]}, {'t': 'Space'}, {'t': 'Str', 'c': 'or'},
{'t': 'Space'}, {'t': 'Strong', 'c': [{'t': 'Str', 'c': 'bold'}]}, {'t': 'Space'},
{'t': 'Str', 'c': 'text.'}]},
{'t': 'Para',
'c': [{'t': 'Str', 'c': 'Here'}, {'t': 'Space'}, {'t': 'Str', 'c': 'is'}, {'t': 'Space'},
{'t': 'Str', 'c': 'a'}, {'t': 'Space'}, {'t': 'Link',
'c': [['', [], []], [{'t': 'Str', 'c': 'link'}], ['https://ix.de/', '']]},
{'t': 'Str', 'c': '.'}]},
{'t': 'BulletList',
'c': [[{'t': 'Para', 'c': [{'t': 'Str', 'c': 'Item'}, {'t': 'Space'}, 't': 'Str', 'c': '1'}]}],
[{'t': 'Para', 'c': [{'t': 'Str', 'c': 'Item'}, {'t': 'Space'}, {'t': 'Str', 'c': '2'}]}]]}]}
[1] https://pandoc.org/