If you focus on the headline, you'll miss the point. The point is they used open source technology to process public data for reporting once the government stopped updating its own tools.
If you focus on the headline, you'll miss the point. The point is they used open source technology to process public data for reporting once the government stopped updating its own tools.
Covering an area of 1.25 million square kilometres, supporting 40,000 first responders helping to protect 8 million people.
Of course the databases are not the only important part of an emergency services network such as ours, but they are a critical component.
I would rather work on a project like this any day than working to prop up some faceless advertising/data collection behemoth such as Facebook or Google.
I started ProPublica's Dollars for Docs [0], and the initial project involved <1M rows. Gathering the data involved writing scrapers for more than a dozen independent sources (i.e. drug companies) using whatever format they felt like publishing (HTML, PDF, Flash applet). This data had to be extracted and collated without error for the public-facing website, as mistakenly listing a doctor had a very high chance of legal action. Hardest part by far was distributing the data and coordinating the research among reporters and interns internally, and also with about 10 different newsrooms. I had to work with outside investigative journalists who, when emailed a CSV text file, thought I had attached a broken Excel file.
Today, the D4D has millions of records, and the government now its own website [1] for the official dissemination of the standardized data. I have a few shell scripts that can download the official raw data -- about ~30GB of text when unzipped -- and import it into a SQLite DB in about 20 minutes. The data for the first D4D investigation probably could've fit in a single Google Sheet, but it still took months to properly wrangle. But the computational bottleneck wasn't the size of data.
One of other tricky issues is that data management isn't easy in a newsroom. Devops is not only not a traditional priority, but anyone working with data has to do it fairly fast, and they have to move on almost immediately to another project/domain when done. There's not a lot of incentive or resources to get past a collection of hacky scripts, so it's really cool (to several-years-ago me) to see a guide about how to get things started in a more proper, maintainable way.
[0] https://projects.propublica.org/docdollars/
[1] https://www.cms.gov/openpayments/
edit: for a more technical detailed example of newsroom data issues, check out Adrian Holovaty's (creator of Django) 3-part essay, "Sane Data Updates Are Harder than You Think", which details the ETL process for Chicago crime data:
https://source.opennews.org/articles/sane-data-updates-are-h...
Here's a great write-up by Jeremy Merrill, who helped overhaul the D4D project after I left. Unlike me, Jeremy was a proper engineer:
https://www.propublica.org/nerds/heart-of-nerd-darkness-why-...
The trouble is, despite (or possibly because of) being cognitively difficult and requiring a certain discipline (for lack of a better word), this kind of work doesn't come across as very "sexy" anecdotally.
Even if it does get shared, the part that makes it hard gets overlooked.
I'm sure there's a spectrum, but my point was that the vast majority of what the companies we read about on this site ("tech") deal with is going to fall close to the consistent-format edge of the spectrum, hence the prejudice.
No it isn't. You are getting defensive for no reason. If propublica ceases to exist or never existed, it wouldn't matter a single bit to the world. You could even argue the world would be better off.
> Processing petty petabytes is not praiseworthy.
From a technical point and many other ways, it is.
I don't get why you are getting offended by people making a jab at the scant amount of data. Last I checked, hacker news is a technology oriented site. And from a technology point of view, what pro publica is doing is a joke. It's a toy amount of data.
Why not just say pro publica is not a technology company and hence people shouldn't expect technological feats of wonder?
> The point is they used open source technology to process public data for reporting once the government stopped updating its own tools.
Which is something I could have done on a lazy afternoon all by myself. It isn't anything to be impressed about. But good for them anyways.
They and countless other journalism/civic orgs would likely be happy for you to show them up by whipping up usable ETL scripts relevant in their respective domains. Since it all involves public open data you don't have to wait for anyone's permission.
Whether or not Propublica has produced something of value for society seems, at an absolute minimum, highly debatable.
Your comment is the one that seems defensive....
You’re being too literal. Yes if nothing exists it doesn’t matter
Look in the mirror and realize you’re just one of thousands that could do this in an afternoon
Given the “big picture” context, your personal computer skills aren’t much to brag about either. Literally good with computers. Get in line.
Were they being defensive? Or offering a context to consider the value from?