I made a native Python module for MS Word docx
github.com
github.com
I was recently wandering around StackOverflow and PyPI and wondering why the only solutions to make MS Office documents from Python seemed to be based around either COM automation, OpenOffice automation, or calling Java or .net libraries.
So I made this module, which reads and writes Microsoft Office Word 2007 docx files. These are referred to as 'WordML', 'Office Open XML' and 'Open XML' by Microsoft.
They can be opened in Microsoft Office 2007, Microsoft Mac Office 2008, OpenOffice.org 2.2, and Apple iWork 08.
The docx module has the following features:
Making documents
• Paragraphs
• Bullets
• Numbered lists
• Multiple levels of headings
• Tables
Editing documents
Thanks to lxml module, we can:
• Search and replace
• Extract plain text of document
• Add and delete items anywhere within the document
• Run xpath queries against particular locations in the document - useful for retrieving data from user-completed templates.
It's only a couple of hundred lines - lxml does most of the heavy work - but it's incredibly simple to use, check out example.py in the link.
Hope you find it useful.
Mike
If more people used LXML and saw how easy it was, the whole idea of modifying serialized XML 'Cthulu style' would be seen as being as odd as modifying other data by manipulating their pickled forms.
Some plans I have to add to this module (if you don't plan too already):
Add support for:
* Pass a font-size, bold, italic, bold+italic and spacing to paragraph method to modify font
* Pass a table row header to table method
* Specify table column widths
* Change text alignment for table rows/headers, etc.
* Change text style in table cells
* Ability to change background color for every other table row
* Create method to insert a page/section break
* Create a method to add an image at a specific position in document page
On one other note:
I'm not 100% positive, but thumbnail.jpg/thumbnail.wmf may result in inadvertent disclosure of sensitive data if using this to generate report documents...
I'm already working on images and 100% nose coverage. I'm intending to do document properties after that.
Re: zebra striping for tables, we already have that via the inbuilt styles. But more options to control styles would be useful.
Coding style is:
* Functional
* Google style - http://code.google.com/p/soc/wiki/PythonStyleGuide
* Unit tests are handled with nose / coverage
I couldn't find anything about a license or copyright on your project's github page.
We've been using http://b2xtranslator.sourceforge.net/blog/ to convert legacy binary office formats to openXML, then using XSLT to transform the XML into our preferred schema for text extraction.
Does anyone know of a good way to read xlsx files with Python? I think I tried one library and it didn't quite work for this 20,000 row file, and I wound up using OpenOffice to convert it to csv. If not, hopefully this can be used as a starting point for developing a nice xlsx library.
http://github.com/brendano/tsvutils/blob/master/xlsx2tsv
I've never tried it on a 20,000 row file. I suspect it would work. A really large file might need to switch to a streaming XML parser, but probably Excel itself wouldn't handle that use case too well.
PS, .eml and .mht are actually the same format.
.mht uses MIME/Base64 to store all of the HTML and external files together, but IIRC it's not exactly the same as an email because the headers are different/missing (i.e. I wouldn't make sense to have a To:/Reply-To:/Cc:/etc header)
A while ago, I wrote a little Python script to read from xlsx Excel files. It's nice that these Microsoft Office XML files can be processed in pure Python.