At the moment, I post fairly infrequently. The whole site is written using emacs org mode. Most of the posts have to do with emacs and data stuff (often doing data stuff in emacs).
34 karma · joined August 10, 2021
At the moment, I post fairly infrequently. The whole site is written using emacs org mode. Most of the posts have to do with emacs and data stuff (often doing data stuff in emacs).
In March, we published an article on Stock Trades by members of congressional committees: https://innerjoin.bit.io/data-cant-tell-us-whether-congressi...
To conduct this research, we needed to know: (1) which members of congress made which stock trades, and (2) which members of congress belonged to which congressional committees. The data for (1) was available from the the senate/house stock watchers sites; the data for (2) came from the ProPublica Congress API. There was no primary key available for linking the two datasets: the best we had to work with were the names of the members of congress.
This would be fine, if the names were represented uniquely and consistently. This was not the case. You can't join "Mitch McConnell" to "A. Mitchell McConnell, Jr." without a bit of work.
Manually matching every single name from the first data source to every single name in the second would be tedious, time consuming, and error prone. Instead, we used the Levenshtein distance to compute a similarity metric between each name in the first dataset and each name in the second. Simply using the best match according to this metric correctly matched more than 95% of the names, and made it incredibly simple to review the list and manually fix the few incorrect matches.
There's also an accompanying Deepnote dashboard where you can compare string distances between pairs of strings of your choosing: https://deepnote.com/@dliden-bitdotio/Whats-in-a-Name-28418c...
This article compares the performance of different methods for writing a Pandas DataFrame to a PostgreSQL database using the to_sql method on DataFrames ranging from 100 rows to 10,000,000 rows.
One point I didn't go into is the fact that the labor force participation rate also dropped steeply in 2020 and hasn't recovered to pre-pandemic levels yet. So that could create labor shortages that are not necessarily represented in the quits rate.
I think the most interesting part is the decrease in layoffs that coincided with the increase in quits. People aren't leaving that much more than before, but when they leave, they're doing so on their own terms.
Pain points: data disappearing, moving, or being updated without notice and without indication of a change. Numbers from the same API endpoints or URL changing unexpectedly and without explanation can be an unwelcome surprise.
I use bit.io (https://bit.io -- I work there) to deal with these problems. It's an online PostgreSQL database; very easy to use with e.g. psycopg2/SQLalchemy in Python or DBI+dbplyr in R. Before any analysis, I copy the necessary data over to a repo/schema in bit.io, fill in the documentation with the dates on which I obtained the data, and use that as the source of "ground truth" for the analysis.
We recently wrote an article (https://l.bit.io/o-cop26) about methane emissions and the COP26 commitment to cut emissions. During the writing of that article, we found some serious inconsistencies in some of the data sources.
Discussions of data quality and validation in data science tend to end with recommendations for a few data validation checks, such as making sure data come from trusted sources; handling missing values; and investigating outliers. These sorts of checks are important, but they won't save an analysis from perfectly-formatted data from a trusted source that happens to be wrong for reasons that can't be found in the dataset itself. Even data of apparently good quality can lead to faulty conclusions.
This article delves into this question by exploring a case study. The U.N. publishes greenhouse gas emissions data supplied each year by parties to the UNFCCC (United Nations Framework Convention on Climate Change). The data are consistent, up-to-date, and well formatted, and the U.N. is a reliable source of official data. However, there is good reason to believe the data submitted by some countries is not accurate. There are other trusted data sources that show startlingly large differences from the U.N. data. In particular, we found that Russia's Methane emissions data were highly inconsistent with the World Resources Institute (WRI) Climate Analysis Indicators Tool (CAIT) data, even though these data were quite similar to the U.N. data for other countries.
A little bit of false enthusiasm now and then? Sure, it makes it easier to get through the day. In both a professional and personal context.
Interesting point, though: smaller states end up at the extremes of both over- and under-representation under the current system (though there does appear to be a systematic bias in favor of small states). I wrote about that in a previous article: https://l.bit.io/census-apportionment-bias. Larger states tend to be close to the national average constituent-to-population ratio while smaller states are more likely to be very over- or under-represented.