How ProPublica Illinois Uses GNU Make to Load Data
propublica.org
propublica.org
If you focus on the headline, you'll miss the point. The point is they used open source technology to process public data for reporting once the government stopped updating its own tools.
I started ProPublica's Dollars for Docs [0], and the initial project involved <1M rows. Gathering the data involved writing scrapers for more than a dozen independent sources (i.e. drug companies) using whatever format they felt like publishing (HTML, PDF, Flash applet). This data had to be extracted and collated without error for the public-facing website, as mistakenly listing a doctor had a very high chance of legal action. Hardest part by far was distributing the data and coordinating the research among reporters and interns internally, and also with about 10 different newsrooms. I had to work with outside investigative journalists who, when emailed a CSV text file, thought I had attached a broken Excel file.
Today, the D4D has millions of records, and the government now its own website [1] for the official dissemination of the standardized data. I have a few shell scripts that can download the official raw data -- about ~30GB of text when unzipped -- and import it into a SQLite DB in about 20 minutes. The data for the first D4D investigation probably could've fit in a single Google Sheet, but it still took months to properly wrangle. But the computational bottleneck wasn't the size of data.
One of other tricky issues is that data management isn't easy in a newsroom. Devops is not only not a traditional priority, but anyone working with data has to do it fairly fast, and they have to move on almost immediately to another project/domain when done. There's not a lot of incentive or resources to get past a collection of hacky scripts, so it's really cool (to several-years-ago me) to see a guide about how to get things started in a more proper, maintainable way.
[0] https://projects.propublica.org/docdollars/
[1] https://www.cms.gov/openpayments/
edit: for a more technical detailed example of newsroom data issues, check out Adrian Holovaty's (creator of Django) 3-part essay, "Sane Data Updates Are Harder than You Think", which details the ETL process for Chicago crime data:
https://source.opennews.org/articles/sane-data-updates-are-h...
Here's a great write-up by Jeremy Merrill, who helped overhaul the D4D project after I left. Unlike me, Jeremy was a proper engineer:
https://www.propublica.org/nerds/heart-of-nerd-darkness-why-...
The trouble is, despite (or possibly because of) being cognitively difficult and requiring a certain discipline (for lack of a better word), this kind of work doesn't come across as very "sexy" anecdotally.
Even if it does get shared, the part that makes it hard gets overlooked.
I'm sure there's a spectrum, but my point was that the vast majority of what the companies we read about on this site ("tech") deal with is going to fall close to the consistent-format edge of the spectrum, hence the prejudice.
No it isn't. You are getting defensive for no reason. If propublica ceases to exist or never existed, it wouldn't matter a single bit to the world. You could even argue the world would be better off.
> Processing petty petabytes is not praiseworthy.
From a technical point and many other ways, it is.
I don't get why you are getting offended by people making a jab at the scant amount of data. Last I checked, hacker news is a technology oriented site. And from a technology point of view, what pro publica is doing is a joke. It's a toy amount of data.
Why not just say pro publica is not a technology company and hence people shouldn't expect technological feats of wonder?
> The point is they used open source technology to process public data for reporting once the government stopped updating its own tools.
Which is something I could have done on a lazy afternoon all by myself. It isn't anything to be impressed about. But good for them anyways.
Whether or not Propublica has produced something of value for society seems, at an absolute minimum, highly debatable.
Your comment is the one that seems defensive....
You’re being too literal. Yes if nothing exists it doesn’t matter
Look in the mirror and realize you’re just one of thousands that could do this in an afternoon
Given the “big picture” context, your personal computer skills aren’t much to brag about either. Literally good with computers. Get in line.
Were they being defensive? Or offering a context to consider the value from?
They and countless other journalism/civic orgs would likely be happy for you to show them up by whipping up usable ETL scripts relevant in their respective domains. Since it all involves public open data you don't have to wait for anyone's permission.
Covering an area of 1.25 million square kilometres, supporting 40,000 first responders helping to protect 8 million people.
Of course the databases are not the only important part of an emergency services network such as ours, but they are a critical component.
I would rather work on a project like this any day than working to prop up some faceless advertising/data collection behemoth such as Facebook or Google.
- Add hashes to their already-pinned requirements.txt deps: https://pip.pypa.io/en/stable/reference/pip_install/#hash-ch...
- Add a Makefile entry to run `[ -d your-environment ] || ( virtualenv your-environment && . your-environment/bin/activate && ./your-environment/bin/pip install --no-deps --require-hashes -r requirements.txt )`
It's better to use empty targets [1] to track when the file has last been loaded and re-run if the dependency has been changed.
[0] https://github.com/propublica/ilcampaigncash/blob/master/Mak...
[1] https://www.gnu.org/software/make/manual/html_node/Empty-Tar...
Tangential question: is it possible to use wget for ftp duties? Though may be additional FTP-specific functionality in `aria2c` of course:
https://serverfault.com/questions/25199/using-wget-to-recurs...
What do you folks use? Drake, "make for data" https://github.com/Factual/drake seems ok, but doesn't have "batch" jobs, (aka "pattern rules") where you can do every file in a directory matching a pattern.
Others have come up with different swiss army knives but nothing ever sticks for me, it usually ends up as a single Makefile with eg 3 targets that call a bunch of shell scripts.
The whole thing would be configurable to build from scratch, but not well set up to do incremental ETL on a per file basis, after I eg delete some extraneous rows in one file, clean up a column, redownload a folder, or add files to a dataset.
Here is a skeleton of what I came up with.
find $(pwd) -mindepth 1 -maxdepth 1 -type d -name ".zfs" -prune -o -type d -print0|xargs -0 -P 2 -I {} echo {}
where,$(pwd) indicates the starting point of the listing of directories
-mindepth 1 makes sure current directory is not listed once again.
-maxdepth 1 makes sure the list does not get recursive
-type d -name - only directories and list names
".zfs" -prune - makes it ignore .zfs (snapshot directories)
-print0 - makes sure to print results without newlines. just -print will print one result per line
xargs -0 will take care of processing out spaces or newlines in the input stream
-P 2 — run two processes at once in parallel
-I {} says that replace {} in teh subsequent command from stdin piped into xargs echo {} will be echo dir1 and then echo dir2 etc
That's just an example to show that we can do a lot with standard Unix tools before bringing in the external sophistication for data related tasks.
In my case, I was dealing with a FreeBSD server. I went the xargs route instead of installing something that is not available by default.
Not sure if Spotify still uses it but it is in their Github org.
http://pachyderm.io seems great but does require more engineering support (needs a kubernetes cluster)
I settled on it after originally using make, getting frustrated with the crazy work-arounds I needed to implement because it doesn't understand build steps with multiple outputs, switching to Ninja where you have to construct the dependency tree yourself, and finally ending up on Snakemake which does everything I need.
Is that using traditional (plaintext) FTP? Is it listening on port 21?
~ $ ftp ftp.elections.il.gov
Connected to ftp.elections.il.gov (163.191.231.32).
220-Microsoft FTP Service
220 SBE
Name (ftp.elections.il.gov): ^C
It looks like they are sending their password in plaintext. aria2 supports SFTP, so they should really talk to elections.il.gov about moving to SFTP or any other protocol that doesn't send the password in plaintext. 211-Extended features supported:
LANG EN*
UTF8
AUTH TLS;TLS-C;SSL;TLS-P;
PBSZ
PROT C;P;
CCC
HOST
SIZE
MDTM
REST STREAM
I didn't check if aria2 would actually use it, but I doubt it.Re alternatives to FTP: as far as only downloads are concerned, HTTPS should be way easier to set up than SFTP.
And the FEC has an API, but has long had the data hosted on public FTP: https://classic.fec.gov/finance/disclosure/ftp_download.shtm...
I spent some time trying to write my own processing system in Python before realizing this was a familiar task...
Why do I like make over shell scripting (sometimes) is that enforces structure. Shell scripts can turn into a real hairball.
When I did ruby I really enjoyed using rake.
There's a well-substantiated linguistic theory revolving around "maxims of conversation". Maxims of conversation are so strongly universal among the speakers of a given language that they become part of the implied meaning of a conversational act.
For example the maxim of cooperativitiy implies that when a person sitting in a cold room next to a window is spoken to by a person sitting further from the window and is being told "It's a bit chilly, isn't it", they can take it to mean "Please close the window".
https://en.wikipedia.org/wiki/Implicature#Conversational_imp...
Similarly, there are certain maxims of conversation which are part of the language game inherent in the formulation of the title of a blogpost. They are kind of assumed to be boasting about something. So when somebody says "We figured out a way to load a gigabyte's worth of data into a database in a single day" then the being-boastful-about-something maxim is violated. That's why it triggered so many people.
And pointing out that this is not something to be boastful about is a perfectly valid thing to do to keep certain facts straight.
...just saying.
But, by all means, if you get a thrill out of it, keep downvoting me.
Once you start abusing the ".PHONY" targets, the value starts to decrease.
That's very valuable for building things. If you change a file, you only need to re-build the files that could reasonably be affected by the change. And things happen in the right order without micro-managing.
Make has you describe a graph of outputs and how to produce them. It then traverses the graph to produce the requested output.
Bash is just a regular sequence of commands, with functions and loops if you wish.
If the pipeline you need to run can easily be turned into a dependency graph, I think make is a great fit. It's easy to use, comes with most of what you need built in and has some fun extras, like -jXXX, which allows you to parallelise things and built in caching so you don't regenerate the same asset twice if you don't need to.
You can do all that in bash but you'll have to write it yourself, which takes time you could spend on other things.
bash is usually the scripting language one uses inside of a Makefile.
It's the default, although one could use any scripting language. Point being, there's no "Make" language, beyond the syntax for describing those dependency relationships and variable assignments.
I'll grant that the distinction is important, though, in the face of the history of #!/bin/sh Linux scripts with bashisms breaking upon the Debian/Ubuntu switch to dash. Even if you're on a system where /bin/sh is bash, it's safest to set SHELL in your GNU makefiles to bash explicitly, if that's what you you're writing in.
In addition, expanding on @Pissompons's note -- make gives you job-level parallelism for free with constructs like:
make -j 24 transform
which will (if possible/allowed by the dependence structure in the Makefile) run 24 jobs at once to bring "transform" up to date.So for instance, if "transform" depends on a bunch of targets, one for each month across a decade, you get 24-way parallelism for free. It's kind of like gnu "xargs -P", but embedded within the make job-dispatcher.
IMO, this is a bad use case for make.