Textract, a Python package for extracting text from any document
datascopeanalytics.com
datascopeanalytics.com
Edit: Looks like catdoc doesn't work with RTF files containing Japanese characters either. Might end up having to use libreoffice or something like that.
./autometa.py --author --verbose academic-paper.pdf
Author: "Edward Witten" Confidence: High (matches template "amslatex")
Not quite as simple a commandline interface as you suggest, but not too hard to set up, and pretty impressive. Now if only Google Scholar would open-source whatever they use...
out = ""
pdf = pyPdf.PdfFileReader(stream)
try:
if pdf.getIsEncrypted():
pdf.decrypt('')
for page in pdf.pages:
out += page.extractText()
except NotImplementedError:
# Yeah, this ain't happeningHere's a link to StructuRise's Textract product page: http://www.structurise.com/textract/
https://gist.github.com/djudd/1402751e2928cb8ac788
It tries either abiword or OpenOffice/LibreOffice for filetypes other than pdf, ps, and txt, which works pretty decently for doc, docx, ppt, etc.
One file type here that textract folks might want to add is Postscript.
I'm wondering why the authors wrote something from scratch ?
edit: this is answered by one author in the 2nd disqus comments of the link
Looking at the source, they didn't lie about “no muss, no fuss”. It just antiword .docs, cat .txt etc
"Its very similar to Apache Tika (which I didn't know about until yesterday), but I think it is different in at least two important ways.
"1. The intention of textract is to provide many possible ways to extract text from any document, provided words appear in the correct order in the text output. By being method agnostic, its possible to use different parsing techniques in different situations. Here's more on that philosophy http://textract.readthedocs.or... and, to be fair, I'm not sure that Tika's philosophy differs in any meaningful way on this.
"2. Another subtle difference is that textract is written in python, which is a language that is used by nearly all data people that I know. Since the intent is to be a preprocessing framework for natural language processing, I wanted it to be as maintainable by the community as possible."
On the same note, your pypi page is borked: https://pypi.python.org/pypi/textract
(look at Build status & co, there's a formatting error)
There go my hopes to see painless OCR library for Python…
http://textract.readthedocs.org/en/latest/#currently-support...
If you have any other (better?) ways of doing this, feel free to add some comments on the issue tracker.