The Economist's excess deaths model
github.com
github.com
Even if it seems justified to shit on the code as is (bad style, lack of comments, no tests, whatever), all it does is discourage similar companies / researchers / folks in academia from posting their code as well. So then we end up with the same code, but now no one can see it - which is the worst of all worlds.
Let's support / work towards a culture of sharing our code before anything else. And if there's something you think can be improved, maybe consider opening a PR with improvements to really _show the value_ to all these parties of posting their code!
When we have the code, we can all work together to improve the model, add the tests, find the bugs.
And honestly, research shouldn't be considered valid if all of the code and data isn't available. Reproducibility is one of the pillars of science, and that's not possible without code and data.
I agree that code should be published alongside papers, and great that the Economist, being a newspaper, publishes their code too.
But when it is the most reputable sources (researchers, universities etc) publishing code of quality so low cannot possibly be correct (looking at you Imperial), it totally should be scrutinised and critiqued.
I think the root cause is, where research is code-based (I guess 90% of science), it should pass the most basic correctness tests before a paper is accepted, or indeed turned into national policy. There are many angles to bite this apple from, but what remains unacceptable is shit code being taken at face value.
If this is referring to the quality of the model developed by Neil Ferguson @ Imperial College, John Carmack audited the code and didn't have a big problem with it: https://twitter.com/ID_AA_Carmack/status/1254872368763277313
Tests can be helpful and informative, or they can be make-work boilerplate that will never reveal any flaws. And in this type of code base, it's hard to imagine the test that's written in code that will be informative.
When dealing with data, "testing" is most often done visually, with plotting.
The process of exploratory data analysis is not anything at all like writing code to an architected specification, and the same practices are not as effective for different types of work. Insisting that there should be automated tests for something like this reminds me of the cultural phenomenon of extreme table-phobia in the 2000s that happened as a reaction to using tables for layout. I had trouble getting people to believe that is was OK to use tables for tabular data, it was extremely frustrating. Or in the last 5-10 years, with design trends preferring spare layouts with lots of empty space, it was super hard to get information dense designs into production, because they were so averse to the idea that people may need information density for some things.
I'm also reminded of this video about the deficiencies of notebooks:
Yes, there are huge deficiencies to notebooks. But using them is an entirely different task than the type of code that one would write in an IDE, for code that gets put into a library and will be run by others. I have two work modes: 1) IDE (actually Emacs) where I write tests and the code is meant to be reusable, even if it's a one-off script, then 2) exploratory data analysis and processing, in notebooks. Though both of these use code, they do not and should not share much else in the way of practices.
We develop best practices in fields to work against our worst instincts. And once we establish best practices, we become strict about them to make sure we don't fall back into bad practices. But when approaching a different field it's good to reevaluate these best practices to see if they still fit.
f(4,5) == 6
is good enough. just, whatever you type in to test as you develop the code, write it down and save it for later.
I mean, yeah, you can go nuts with boilerplate. But, I always seem to regret not having written down a handful of cases. If nothing else, it helps me remember what the heck this function is for, and a few samples of expected input and output, along with source are more immediately useful (to me) than comments.Opinions differ. This isn't a hill I'll die on - as evidenced by my lack of tests from time to time. But I've never regretted the few minutes to provide them, and I have regretted not providing them
Suppose you want to validate that a certain column is in the file. You could place your assumptions around that explicitly in the read_tsv function call and have it generate the error for you on reading the file, or it can generate an error when you try to use that column. Either way, there's no "test" yet the intention is clearly stated in the code, and the code will simply fail out rather than proceed if it encounters input that can't be meaningfully used.
This code is in many ways a declarative use of other code, describing the data, doing some joins and filtering, and plugging it into well tested libraries of code. The data structures that store the data and the library functions that deal with the data are designed in a way that bugs often result in errors immediately.
In many ways it's like writing a SQL statement. One could make test data tables and make sure that the SQL statement does what one expects, but the SQL statement already embodies intent in the same way that a test of code would embody intent.
I would point out this guy - https://github.com/TheEconomist/covid-19-the-economist-globa...
that sort of has the look of something that took a pass or two to get right. And I think, would be nice to have a test case (sample call).
Again, not a hill I'm going to die on.
Overall, I don't think we have a big difference of opinion, yeah, it looks like gluing together libraries - as a non, native R speaker, I think I can make sense of the project, it's cool they put it out there, and I don't think there's anything _wrong_ with what they've got.
If you want to run a loess, it runs the loess function if not it runs a windowed average.
What exactly would you test it for?
If I were developing this, I'd probably be running it in a notebook, and inspecting output on a few rows from the data frame. It would be easy enough to capture both the input and output and throw it into a stopifnot to record the testing that was performed interactively.
> When dealing with data, "testing" is most often done visually, with plotting.
Even if your validation process is entirely manual, you still can benefit from automated testing. Snapshot tests are completely agnostic from whatever serialization format. If you eyeball check something and say “yup this is good”, that’s something you could snapshot and never test again.
Don't get me wrong— this is a good direction to be pushing in. But it's something that's been very much enabled by the internet, which is why I'm finding it weird to be blaming the prior status quo on internet culture attention span memes.
No, it has created a generation. of conspiracy-mongering freaks.
Or, more precisely, it has allowed that particularly seedy underbelly of society to become more noticeable.
The discussion should be on the model design and whether it’s sound.
Leave it to HN to start criticizing code whenever anything is posted.
I've done my share of implementing statistical models in code and have seen plenty of examples of incorrect code failing to implement the model correctly.
Simply clone the repo, download the data, and verify the results.
The discussion needs to not be on the code but the quality of the assumptions in their model.
Which you can read here.
https://www.economist.com/graphic-detail/2021/05/13/how-we-e...
Obviously the assumptions of the model are important. But I was responding to a claim that the quality of the code is not important, which is also obviously false. Neither discussion alone is sufficient.
When code is one shot, accomplishes its goal, and does not need to be extended or relied upon in the future, why do all these things?
Is there any goal above that is not also serviced, at least in significant part, by open-sourcing it?
Bravo to The Economist for releasing it. This is unabashedly a Good Thing.
There have been 7M-13M excess deaths worldwide during the pandemic - https://news.ycombinator.com/item?id=27177503 - May 2021 (457 comments)
as long as that is crystal clear, that "two old dead soon anyhow Alzheimer's patients" completely misrepresent the 500.000+ dead US citizens...
... such an anecdote is perfectly fine ...
[1] https://www.nature.com/articles/s41598-021-83040-3
[2] https://www.news-medical.net/news/20201021/More-than-25-mill...
[3] https://www.sciencedaily.com/releases/2020/09/200923124557.h...
I can't resist: males of america! you are the the victims, the "loss leaders" of the plague:
> We find that over 20.5 million years of life have been lost to COVID-19 globally. As of January 6, 2021, YLL in heavily affected countries are 2–9 times the average seasonal influenza; three quarters of the YLL result from deaths in ages below 75 and almost a third from deaths below 55; and men have lost 45% more life years than women.
let that sink in:
> men have lost 45% more life years than women.
mask up! vaccine up! and don't let them mn-killers tell you otherwise.
/scnr
Can you suggest me an advanced Bayesian Statistics book that focuses on application without sacrificing too much mathematical rigor?
I am graduating in MS in Stat. I've took a Bayesian Stat course that followed Statistical Rethinking by Richard McElreath. I liked this book because the author appeals to intuition instead of mathematical rigor. I took 2 semester long statistical inference course, so I am ready for some advance material.
So if you killed everyone on earth the mortality rate for the next eternity would be 0.
Please don’t tell the singularity
;)
We'll all be dead eventually, but most of us aren't OK with losing X years of it.
The number of cumulative covid deaths in france is 109k, of which 65k to the 1st of Jan 2021 [3]. So the excess deaths seem to be in line with covid deaths.
Now you could get a very different result for another country depending on how well covid deaths are reported. Also difficult to predict how this will look like in 2021. Covid deaths seem to be predominantly people near their end of life, but on the other hand the delay of lots of medical procedures as the result of the lockdown should have its own excess death, plus impact of change in crime, accidents, increase in poverty, etc. God knows what will be the net effect of all this.
[1] https://www.insee.fr/fr/statistiques/2383440
[2] https://zbpublic.blob.core.windows.net/public/excessdeathfr....
[3] https://www.statista.com/statistics/1103422/coronavirus-fran...
The point being, excess deaths might be better stated as premature deaths. Mind you, death is death. But someone further along in life already nearing death's door isn't the same as a 25 y/o.
Seems that it'll be a year or more before we have more data and can quantify how premature the premature deaths were.
There should be, due to the “harvesting effect”.
The real code quality sin here is that R, a vectorized-by-default language, is being used to employ for- and while- loops everywhere. It’s strange to see given the absolute plethora of R resources available aimed at avoiding just that.
But props to the economist for releasing the code, anyway. I’m a happy subscriber and will continue to be one.
> Until fairly recently, we were less comfortable with statistical software (like R) that allows more sophisticated visualisations.
https://medium.economist.com/mistakes-weve-drawn-a-few-8cdd8...
The US death rate didn't appreciably increase (~3 per 100k) in 2017 vs. 2016[1] and has generally been improving. That's about 10k extra deaths, not 400k, like the paper says (because the paper's usage of that term is very different than The Economist's!)
https://www.cdc.gov/nchs/data-visualization/mortality-trends...
But that's exactly what the paper shows. It's reasonable to ask the question why was the difference in deaths far greater in 2017 than in the surrounding years and why was the mean age of the persons excessively dying so relatively young? There most likely are structural differences causing excess US deaths, but the claim that those structural differences caused a huge relative spike in one year for no particular reason isn't very satisfying.
"Applying European age-specific death rates in 2017 to the US population, we then show that adverse mortality conditions in the United States resulted in 400,700 excess deaths that year."
To paraphrase: on average in 2017, US people died more than the equivalent people in some bit of Europe.
No great event happened.
Perhaps that was worldwide?
Here's a graph from before the pandemic, including a (now out of date) projection:
https://www.macrotrends.net/countries/USA/united-states/deat...
(Note that the slope looks exaggerated because the baseline isn't zero.)
In this context, excess deaths are variations from these tables - more people from a cohort dying in a certain time period than expected.
You'd probably want to gather life expectancy rates for individuals once they reached five years of age (since infant mortality necessarily lowers life expectancy but after a cliff people who make it tend to live longer than the average with a bunch of 0s, 1s, and 2s in it) for each country in question.
Then you'd need stats for the age of the dead in each country which I guess you don't really have in most cases since the deaths aren't aggregated anywhere since they haven't been attributed to COVID.
Thanks to the Economist for publishing their model. Would be neat to see more stuff like this. I haven't delved too deeply into the code (several 1000 line scripts which would take some sussing out) but it's nice people will have real world projects they can delve into.
It might also be interesting to measure the cost of the pandemic in terms of economic harm, how much the GDP dropped, how much production was lost, how many stores closed, etc.
Consider the average human worker. In their life, they will output around 10.5 human years worth of work. 50 million years represents the total output of about 4.76 million human lives.
So if you delete the production of 3 million people due to COVID deaths plus all the time spent in lockdowns, 50 million human years lost sounds about right, but probably could be higher. I'd say it's probably about 80 million human years lost.
We could count "years of productive work lost", but that feels like a very cynical way to look at it.
On way of contextualizing it that doesn't equate lives with economic output is to divide by life expectancy: if life expectancy is 70 years then 50 million years lost ~= 0.7 million human lives that never happened, or 0.7 million healthy newborns that died preventable deaths.
> From a purely utilitarian perspective, the worst age to die is 20-22,
that's easily adjustable: if 95% of newborns become healthy 20 yr olds, then just multiply the number of newborns by 0.95. The harder part is estimating total number of years lost, but it doesn't sound particularly out of the ordinary for a research problem in statistics/epidemiology.
Which "costs" more, 100k 20-year-olds dying today, or 100k babies dying today? The 20-year-olds have already had a lot of time, money, and opportunity poured into them, while the babies have not.
“oh, those slaves in [country] only cost me 0.03$/year, no big deal.”
As the hypothetical trolley controller, how many 50-year-olds would you kill to save one 16-year-old? And would you kill the 16-year-old to save a one-week-old baby?
https://en.wikipedia.org/wiki/Value_of_life
This does have this tidbit:
"Historically, children were valued little monetarily, but changes in cultural norms have resulted in a substantial increase as evinced by trends in damage compensation from wrongful death lawsuits.[38]"
Organ donation prioritization does this analysis all the time. They don't publish the exact rules they use for matching but the most important are urgency, compatibility, locality, and survival benefit. Survival benefit is almost literally expected quality life years (QALY) from an organ transplant. A 98-year old will not get an organ that could give a 50-year old another 20 years of life, but a 20-year old that can survive for another year will not get an organ that the 50-year old needs this week. Compatibility and locality are the biggest problems apart from basic availability (too few organ donors, too many patients needing transplants). Once we can print/grow organs the calculus will reduce to prioritization on urgency and then QALYs if supply is still constrained in some way.
That would capture not only years of life lost but also the quality of those years. Long-covid seems to be a serious issue for a lot of people and an accurate estimate of that impact would be very valuable.
“1.5 million potential years of life lost to COVID-19 in the UK, with each life cut short by 10 years on average”
https://www.health.org.uk/news-and-comment/news/1.5-million-...
For your other point about excess mortality being caused by lockdowns, you can view charts of excess mortality here: https://www.euromomo.eu/graphs-and-maps/
You can see that some countries, e.g. Norway had lockdowns but did not incur any severe excess mortality - but note that this only proves that you can have lockdowns with no excess, not that every lockdown is equal.
https://ottawa.ctvnews.ca/cheo-joins-other-children-s-hospit...
While that's not the same as saying actual deaths from suicide have increased, there is a serious problem, at least here in Canada.
edit: Attempts up 90% in Colorado as well:
https://coloradosun.com/2021/05/25/mental-health-emergency-c...
Suicide is a serious issue and shouldn’t be ignored. We should also apply the same logic to suicides as people are applying to covid - would these people have tried suicide anyway but covid/lockdown just brought it forward? Will we therefore have a drop in attempts now to mirror the spike up?
While a children's hospital in Colorado might have reported a 90% increase, across the whole of the US suicides were down 5%.
As far as I know, most countries are reporting an overall drop in suicides during the pandemic.
And US drug overdose went from 72k in 2019 to 81k in 2020. So less than 2% of the economist's 500k[1] might be attributable to that.
You might point at road deaths, which have increased quite strangely, but again, wrong order of magnitude, by 2 orders even, increase of 4k compared to 500k.
So to answer your question, no, it's not a questionable assumption.
Perhaps you might consider the simplest explanation, a deadly virus, is the best one?
[1] https://www.economist.com/graphic-detail/2021/04/05/deaths-i...
Theres a huge amount of data reshaping, mutating, and general wrangling going on here, how can one be confident without even a known input->output integration type test?
Testing is hard. The researchers who write this code typically don't have the necessary hands-on experience to write good tests, even if they had enough time in a day/week to actually do it.
Edit: Also the code tends to be very "high level", chaining lots of high-level API functions together, and even coming up with assertions to test tends to be a bit of a challenge. Testing such code turns out to be surprisingly difficult; you might end up just rewriting big chunks of your code in the test suite.
In my data science work, I've focused on writing tests for the complicated sections (e.g. lower-level string processing routines) and just trying to focus on the keep the other stuff very clean and readable.
Similarly, even such simple test harnesses help when yourself or other go to modify the code. Having flags for "Did I break something" is very important.
It isn't. It is the evidence that calculations in the code were done correctly.
It is like the problems in high school math class. The teacher controls the inputs so that the outputs are known. The teacher then runs the test problem past the student (aka the code). If the known correct answer isn't generated then we know the student didn't do it right.
Then, why don't you apply that way to the output? You don't need tests if you only need to do it once.
Isn't that what we are discussing, how to tell if a piece of code works? You could do a formal proof that the code works (which is long and tedious) or you can test it (less long and tedious but less rigorous).
>You don't need tests if you only need to do it once.
It doesn't matter how many times you plan on running it. What counts is how much we value the output. I wouldn't bet $20 that untested code works correctly.
In general (and in my opinion), a "good test" is one that asserts that an invariant is always so, or that a property that is expected to hold under certain conditions does indeed hold under those conditions. Defining such invariants and properties for "data science code" tends to be difficult, and even when you define them it might be difficult or impossible to test them in a straightforward fashion.
And that's even before you get to the probabilistic stuff.
Research code, because its requirements are not fixed and it is not frequently run, doesn't need to be as "stable" as a bonafide app. For a bonafide app, requirements do not change as often and the app is run virtually 24/7.
Once the research code becomes "productionized" however - i.e. it is deployed in an online system where uptime & accuracy matter - then I think absolutely it becomes more engineering-heavy and looks quite a bit less like this code.
Would be curious to hear others' thoughts on this distinction between research vs. production code however.
I suspect it will be a downvotably unpopular opinion, but a typical researcher has a significantly larger cognitive capacity than a typical software developer. Consequently, a piece of code of certain complexity will look simpler, and be easier to manipulate for a researcher than for an average software developer.
Also, if your job is writing code, I'd bet you'd have less difficulty manipulating spaghetti code.
My impression was ->
Software developers main focus is developing software so they spend a HUGE amount of time developing software, maintaining software, noticing patterns and bugs and pitfalls in software and thus they get pretty decent at writing software.
Data scientists main focus is developing models, so they spend a HUGE amount of time developing models, tweaking models, finding data, and cleaning data, and write basic software to achieve some of those goals.
I wouldn't expect Albert Einstein, Mozart, or Beyoncé to be some fantastic software developer just cause they're smart individuals. I'd expect people who spend a lot of time writing software to generally be the ones who write decent software.
Most of the time you're just testing the viability of things and just want to fail fast if they don't work. Getting too attached to some pipeline or abstractions is usually a bad idea.
I don't think this excuses bad code, but I kinda get it when it happens, since you can't just use best practices from the beginning like in software development
I'd argue it's more about startup time, or whatever the technical term is for the time between when you first start reading code and when you've built enough of the model in your working memory to be able to touch it without breaking everything.
That heroic developer doesn't need to concern themselves with other people's experiences, and they can just write whatever weird idioms suit their brain.
That said, having used a lot of research oriented software packages over the years, researchers typically do not produce high quality software. Their objective is to produce papers, not software that other people can use. If someone else can use it too, then great, but that's not the primary motivation.
I'm genuinely curious because I understand the need to move fast but is accuracy a necessary sacrifice? (or is there a trick I don't know about)
Here it is not as clear what the code should do. What should the amount of excess deaths be? In what ways would we change the logic such that the test case would break? If the input data set is static, isn't it more like a mock anyway?
I think for this reason you often see more sanity checks in research code because the should case is not as clearly defined.
With several fake known inputs and there associated outputs we should be able to determine if the calculation is right.
The result on the real world data is not known but when calculating a statistic you should be able to figure out if you are calculating the right statistic or returning 42 for all inputs.
The only way to make sure it's right is the same way you'd do full verification of other code, going through it line by line and making sure it's doing the right thing. Tests can not do that, they can only assert some forms of intent, and they are not good at catching the types of errors that result from ingesting varied data from lots of sources and getting it into modelable form. It's ETL plus a bunch of other stuff going on here.
However, "the only way to make sure it's right is to ... go through it line by line" is a bold claim. There are multiple named functions and some unnamed functions in this example that could be verified for programmer mistakes (such as typing 1000 vs 10000) and edge case handling (edge cases that often arise from messy ingested data).
But even if they were tested, I'll concede that mistakes can still be made.
I too know first hand that students and even academics who aren't properly trained and given the right tooling will write prototype code that's not production ready. Occasionally there's even a bona fide mistake, though usually it's just very un-generalizable. It's part of my job to make some of it more production ready. That's the best use of my time and training. Conversely, I'm shit at generating useful research ideas or writing papers. Just don't have the right combination of intuition/training/experience. Good thing my academic institution has the resources to employ them and me.
With tests it is often best practice to write the tests before writing the code that gets tested. But what if you don't know ahead of time, and can't possibly ever know ahead of time?
Writing "tests" for this would be about adding stuff afterwards, like "the max and min values of this column were X and Y". But that is expected to break, if anything changes, because it's not testing anything useful.
My question is: what is one concrete test here that would be useful and actually provide confidence that the code is doing what it should be doing, and how does that test provide better sanity checking than inspecting the tables at each step of the process?
It should be said that the statistics themselves do make for some 'guide rails' .. the stats are quite demanding, require a lot of upper-division training to use correctly, and visual feedback can course-correct as the results are iteratively found.
As said in other comments, the content is often associated with a research goal that is weighted more than code-quality. In contrast, a general purpose development language has very broad application, and probably can go wrong in very broad ways. The coder is getting feedback on results, but often lots of code review and expectation of professional results.
I picked out an R ETL sequence from an article last year, and I still pull it out once in a while, to see the extensive, clever and (to my eye) really hard to read R data ingestion and manipulation. Personally I think it is a fair tradeoff to say that the expertise to use this environment is a bit of a filter, and the rigor that the results (therefore intermediate results) demand, does move the expectations on code quality ..
As for tests, there is probably no defense on not having tests also.. it is probably going to be more common as the field of R and data analysis inevitably grows.
It's widely known on the "theory of testing" that data oriented systems can't be tested without immense amount of effort. If you are interested there are plenty of academic talks about the subject.
Good software practices don't only have to be for "production" or CI or large teams of coders. Testing for correctness could be seen as part of the work of delivering a high quality product/graph/prediction/model.