Code to transform Hillary's emails from raw PDF documents to a SQLite database
github.com
github.com
https://github.com/wsjdata/clinton-email-cruncher
Kind of fun to compare methods...converting PDFs can be such drudge work sometimes. I'm interested in the data-cleaning/name-resolution aspect...though that is also another deadly boring programming task.
Nice code, BTW. How is pdftotext when compared to PyPDF2?
Bedarra's Text Analyzer[1] kinda floored me and I'd like to use something similar for various tasks, if there was something good and free.
Yes, tesseract[1] can do a pretty good job. Here[2] is a blog post which describes using it to perform OCR on PDF's.
As for searching the PDF contents, Solr[3] might be what you are looking for instead.
1 - https://github.com/tesseract-ocr/tesseract
2 - http://fransdejonge.com/2012/04/ocr-text-in-pdf-with-tessera...
3- http://stackoverflow.com/questions/6694327/indexing-pdf-with...
Not good form at all.
https://www.propublica.org/nerds/item/doc-dollars-guides-col...
And let's face it, if you are an end-user who is not a programmer, a PDF is very convenient. You can even do Ctrl-F to look for things...and if you aren't a programmer, that's all the search-power you need to do your compliance work. After the investigation was published, I fielded a lot of questions from auditors who were convinced I had built some amazing proprietary search algorithm...such people just don't know the gap between doing Ctrl-F versus "SELECT * FROM doctors where LAST_NAME like ..." :)
There were a couple of companies who arguably went out of their way to obfuscate their records. They had built websites that consisted of Flash widgets designed to look and act like a (crappy) HTML table...except with Flash, you can prevent the user from select/copy/etc. Unfortunately, the developers of those widgets never told the companies that their Flash containers operated by making XHR requests to external XML/JSON files...so these companies ended up being the easiest to gather the data from.
But in this case, if I understand you correctly, the original source documents were PDF? That they were distributed as PDF for all the reasons you state?
That doesn't apply in this case - the original documents were not PDF, and had no reason to be PDF except to make it more difficult.
The code posted here isn't doing any OCR, but whatever generated the PDFs (Acrobat?) might have.
http://www.nytimes.com/interactive/2014/12/09/world/cia-tort...
http://money.cnn.com/2015/03/11/technology/security/hillary-...
Edit: Hilarious part was where State says this is not a problem because "it would have to print Clinton's emails in the normal review process."
[1] Government officials have been burned before when using what they thought was a PDF redaction tool...i.e. using the box-drawing-tool and drawing a black box...was not actually redaction...Governor Blagojevich's case comes to mind: http://www.wbez.org/blog/update-blagojevich-lawyers-created-...
In this day and age it takes a lot more work to print, manually redact and scan than it does to digitally redact permanently. Your story about redacting PDFs is interesting these were not originally PDFs, they were emails.
I can't see any other reason for delivering documents as printed/scanned PDF than to make it more difficult to analyse considering they were originally not paper or PDF. Occams Razor applies here.
Because this isn't possible. How are you going to account for any possible misspellings, accidental spacing or other unknown unknowns? You can't. To properly disseminate things the documents have to run through actual people. Sure, people fuck up things too but they can be much more precise when it comes to processing real, human language than any computer right now.
This is the process the government takes. Convert to PDF, redact, release.
> I can't see any other reason for delivering documents as printed/scanned PDF than to make it more difficult to analyse considering they were originally not paper or PDF. Occams Razor applies here.
You are misusing the meaning of occam's razor. This is actually the easiest, most correct way to disseminate possibly sensitive information from the government that they have available. It sucks. If you can write something better then do it; you can make a ton of money and save the government a whole lot of time. But it better be really, really accurate.
How sure would you be of a tool that let you do this?
"Flattening" the document through a paper medium offers a benefit of reducing the digital trail to the tools used to convert from paper to PDF. Unintended data leaks are minimized.
Occam's Razor falls on the side of the print-to-paper-then-scan-to-PDFs. When you print an email or Word document to paper, you have effectively destroyed all non-visible metadata. When you use a Sharpie to rub out a name, you can be sure that mark is going to propagate to the electronic scan. You have none of those assurances if you work electronically...and imagine being someone who _isn't_ a programmer. Not too long ago the NSA was clusterfucked by a low-level employee who ran wget without anyone noticing. If you're the non-programming bureaucrat in charge of this bespoke redaction process, software seems very arcane and insecure...paper sounds pretty nice in comparison.
But yeah, I do wonder about the amount of human labor and late nights that go into this. But apparently, a lot of our legal system has revolved around legions of paralegals and lawyers sifting through uncountable volumes of paper...and a lot of the people doing the current redaction probably came from that field of work.
I get why the agency might print it out- but I don't get why the Clinton staffers would print them out. If the government is going to subpoena or raid me, they're just taking my computer. Sure, if it gets to public release, then they might print/redact. But they sure are not going to wait for me to print out stuff and hand it over.
Given the original 'two phones are too difficult' excuse for having the private emails, I can't see how printing out the paper and handing that over can be seen as anything but deliberate delaying. Unless I'm missing something here, it's not like the Clinton staff redacted the mails themselves - they just deleted the ones they didn't want to hand over.
If the agency staff are redacting for public release, then that is a different thing altogether, they could have done that from the original emails.
Compare with the 2013 IRS controversy, which could not be properly investigated because Lois Lerner's hard drive had been erased.
"Why do we always refer to female public figures by their first name, and male ones by their last?"
I heard/read that question somewhere, and it stuck with me, can't stop seeing it. And this is a fine example of that.
Google results:
"Clinton email scandal" 41M results.
"Hillary email scandal" 48M results.
Yes, I do realize hers is a special case, and it's necessary to use her first name to disambiguate wrt to Bill. But the phenomenon exists. Just some thing to keep in mind.
[Edit:
"In the tennis commentary, women athletes were called by only their first names 52.7% of the time, while men were referred to by only their first names 7.8% of the time."
http://www.la84.org/gender-stereotyping-in-televised-sports/
]
Co-workers who have known each other for longer than they have known their spouses address each other as Herr Schulze and Frau Schmidt.
I realize that's a weak justification, but it does make some difference... And 41M vs. 48M does not require a very strong justification.
Of course sometimes the issue is that people use first names in order to intentionally avoid showing respect, but that happens to "dubya" and "jeb" and many other public men as well.