Grad Student Who Took Down Reinhart and Rogoff Explains Why They're Wrong
businessinsider.com
businessinsider.com
My co-authors and I chose Nature as the target for the paper because of its reputation and because I believe in going to 11 (http://en.wikipedia.org/wiki/Up_to_eleven).
If I brute-force detection of the first 10000 prime numbers (for some non-CS context) by testing their divisibility from n to 1 it's horribly inefficient and programmers may laugh, but it's accurate, simple, and easily reproducible. Given that a definition of a prime number as divisible only by itself and 1, it may be better to present a brute-force algorithm than to get into a sideline discussion about validating some new-fangled method like the Sieve of Erasthostenes, that upstart.
Science aims at proof rather than efficiency. What matters is the quality of the result, and its reproducibility. If you can get to the same result in a much more timely fashion that's awesome, but that's an engineering achievement rather than a scientific one. I don't think that anyone wants to put out code labeled as CRAP in an attempt to mollify engineers.
And that's to be expected, because of how most scientific code gets written, and what the incentives are. And releasing this code would still be strictly better than not releasing it. We just need to keep the expectations of code quality low, which the CRAPL license does.
Your comment reads to me like you want non-programmers to ritually humiliate themselves by labeling their stuff as CRAP before you will deign to parse it for correctness. Patronizing people is not a good way to get them on-side.
We just need to keep the expectations of code quality low, which the CRAPL license does.
The purpose of the CRAPL is to make code accessible, not to lower expectations. It's not about what you think of the code, it's about whether the code yields correct results.
Not to dredge up a functional vs imperative battle, but I feel FP has a lower impedance mismatch with mathematical concepts generally.
You might have better luck in a quasi-research public venue. For example, I work at a public agency that uses your tax dollars to forecast travel demand, and then use the results of that statistical modeling (plus a thick shmear of political wrangling) to decide how to spend more of your tax dollars.
Much of this model code is developed as part of an honest to goodness research process (yay!) by contract software developers that public agencies can afford (not so yay). In other words, things like revision control and unit tests are mostly dismissed as extravagances and unwarranted delays. Validation that is performed doesn't exactly inspire confidence.
I'm betting there are a million and one of these kind of projects. Some of us are trying to figure out internal and external issues so we can post this sort of thing on places like GitHub. Others are already there.
If you pick a topic you're interested in, and ask the right people, it'll probably be worth some authorship cred.
Does that mean to 'notify the Author about Your work', or 'notify the (Author of Your work)' (i.e. yourself)?
If this license is meant to be serious then I should probably notify its author about these concerns.
So did I... however, showing code and data should be mandatory for acceptance of papers. Errors are all too easy to miss and since papers draw conclusions from the data a paper without data should actually be simply refused for being incomplete.
Why publications would accept a paper for review and eventual publication without giving reviewers the ability to actually check the results is a mystery to me. The current system basically amounts to 'you're going to have to believe me'. I singled out CS papers because there the code is pretty much the essence of the paper but of course the same goes for every other field.
In addition, researchers are likely to discover mistakes while cleaning up the excel sheet, data sources and code before publishing it. Just like we find mistakes in our work while refactoring and cleaning up code before we push it to github.
So even when nobody ever looks at the data and code we can expect the quality of the research to improve significantly: just because the code is there to be looked at.
I think the case in favor of making data & code public for academic research is pretty overwhelming.
The more I learn about this particular incident, the more I feel the scandal is not (a) that famous harvard professors made a mistake, or (b) that microsoft excel is more error prone than alternatives (which seems like nonsense to me because the more complex the replacement the more error prone it will be)...
To me the scandal is that you can be a tenured professor in economics and produce work that amounts to simple averages of widely available data and call it a research paper, and that people take you seriously, and presidential candidates use you as a reference, and your department doesn't bat an eyelash.
The fact that they screwed up seems incidental - mistakes happen.
It seems like the kind of back-of-the envelope work that any old blogger would be capable of doing, we just don't have a way to take good ideas, no matter where they come from, seriously. no matter how much we profess to try - at heart, the consensus is still status-driven, and pedigrees matter.
Not necessarily everyone. Collecting good data is hard, long, and tedious, and the 'glory' part is the analysis. People get accolades for making analyses, not for good groundwork.
[1] http://www.aeaweb.org/aer/data.php [2] http://www.aeaweb.org/aer/pandp_faq.php
If data and code are open this is fantastic for science and progress because errors do not replicate and permute but it can be terrible for individual scientists and graduate students. As always, it comes down to money. If we are forced to correct or even worse retract a scientific publication, the community (both local and worldwide) bares down on you like a knife. The stakes are so high that not only is falsification lethal, in many cases honest mistakes can be lethal as well. Conversely one cannot take eternity flipping every stone to bulletproof every possible single problem and it can be very hard to identify which stones to check!
This is not a defense of closed work and closed data, but it's a realization that opening data and code is not a simple and straightforward process. There are severe and deeply embedded cultural problems that make post-publication sharing difficult.
An example from economics: Meese and Rogoff [1] has (by google scholar) about 3000 citations, which is massive for econ and is one of the key references in Intl Macro. They showed that all of the models in use at the time of publication sucked.
[1] Empirical exchange rate models of the seventies: do they fit out of sample? links to the paper here: http://scholar.google.com/scholar?cluster=170917352624567913...
That was clearly not the case with the flawed paper this OP is concerning itself with. So, not a valid excuse in this case.
At any rate, the team should learn to write papers as they build up the data set. That way they always have a little more data than anyone else working with it. Because without release of the data on which a given paper is based, there's no way to know if the paper is actually valid.
"It would be absurd to think that governments never have to worry about their level of indebtedness. The aim of our paper was much more narrowly focused. We show that, contrary to R&R, there is no definitive threshold for the public debt/GDP ratio, beyond which countries will invariably suffer a major decline in GDP growth."
Honestly, there are so many variables and dimensions in this system and so little data I'm not sure we should ever be drawing sweeping conclusions like this.
That doesn't mean they were asserting what the grad students claim, that the falloff in growth is unavoidable, which sounds like an extreme strawman.
But their own public pronouncements, and the way they responded to other's interpretations of the results, pretty much indicated that they though they'd found firm evidence that there was some big non-linearity at 90%.
"We show that, contrary to R&R, there is no definitive threshold for the public debt/GDP ratio, beyond which countries will invariably suffer a major decline in GDP growth."
RR's claim is entirely statistical, so of course they're not going to claim that exceeding the ratio always ("invariably") leads to a major decline in growth.
RR's being wrong doesn't mean it's okay to insinuate they hold ridiculous positions that they really don't. That's just compounding public misconceptions.
The former is what RR said, the latter is the strawman being attributed to them.
The response by Herndon doesn't insinuate that countries should disregard debt as a matter of concern. It doesn't even say that debt increase isn't correlated with GDP decrease. The main issue that the paper points at is that US and European governments have, based on this thesis, adopted drastic policies of austerity and budget cuts, when actually over history there isn't such a nonlinearity.
I'm not sure this is right. Rogoff said "back in 2008-9, there was a reasonable chance, maybe 20% that we’d end up in another Great Depression. Spending a trillion dollars is nothing to knock that off the table." I didn't see any policy recommendations in the paper.
The main issue that the paper points at is that US and European governments have, based on this thesis, adopted drastic policies of austerity
_Maybe_ for the US, but the timing doesn't match for Europe. For the UK, the usual candidate for self-imposed austerity, Cameron was arguing for austerity before the paper became well-known.
"Our finding is that when properly calculated, the average real GDP growth rate for countries carrying a public-debt-to-GDP ratio of over 90 percent is actually 2.2 percent, not -0.1 percent as published in Reinhart and Rogoff. That is, contrary to RR, average GDP growth at public debt/GDP ratios over 90 percent is not dramatically different than when debt/GDP ratios are lower."
Well, except #1 should be "not being able to run experiments in laboratory conditions". Statistical controls in an experiment are still controls, and it is still possible though (because of the point made briefly by #2, to run statistically-controlled experiments) sometimes of limited utility; the dearth of data (specifically, the small number of data points compared to the number of independent variables that need to be controlled for) is what limits the utility of many experiments with statistical controls in the field.
There are just so many variables. And the more interconnected the world becomes, the more variables there are. Pointing to any one correlation and you'll find a hundred statistical fingers pointing at correlations with totally different consequences.
> There are just so many variables.
Right. That's the problem with too few data points given the number of independent variables that need to be controlled for. I addressed that explicitly.
But when the goal is to create a result rather than report it, transparency is the enemy
I strongly believe that errors are a much more pervasive problem in science and related fields than malice is.
But we are talking about economics, and republican groups like fox news eat it up.
My favorite example is the 2011 chart distorting the display of the unemployment rate:
http://mediamatters.org/blog/2011/12/12/today-in-dishonest-f...
Wherever the line is drawn, the dishonesty of some random Fox News chart is not related to the honesty or dishonesty of actual scientific research.
The much bigger problem in science is error.
Econ and Psych really do lie in a gray area, and by avoid the demarcation debate you completely miss the relevance of the issue at hand.
If policy makers weren't using this study to justify more austerity, then we probably wouldn't have such a prolonged discussion.
True, but I haven't seen a liberal think tank manipulate charts in that specific manner.
I suspect there are plenty of similar issues in economics. People with money are much more likely to support researchers whose work benefits or protects people with money. E.g., I happened to read a paper from a U of Chicago prof arguing that insider trading is actually beneficial. Boy, I wonder who the big donors are there. Probably not Mother Jones Magazine.
In science generally, probably. In areas tightly connected to perennial areas of sharp ideological policy divides, like macroeconomics, I'm less convinced.
The problem is that economics is not a field related to science. The vast majority of public policy economics starts with a conclusion, and then creates facts to support that conclusion. This is not science, it is religion.
(insert disclamier about this being an overgeneralization)
From Science: All data necessary to understand, assess, and extend the conclusions of the manuscript must be available to any reader of Science. All computer codes involved in the creation or analysis of data must also be available to any reader of Science. After publication, all reasonable requests for data and materials must be fulfilled. [...]
http://www.sciencemag.org/site/feature/contribinfo/prep/gen_...
http://en.wikipedia.org/wiki/Presidential_Records_Act
http://en.wikipedia.org/wiki/Bush_White_House_e-mail_controv...
By all means, force researchers to publish all the tools necessary to reproduce their results. But you can't expect to set up surveillance in their head.
It could go as you suggest. Or it could be like a locker room: if everybody is naked, then nobody cares.
What made me add drafts to the lists is Daniel Dennett's energetic description in Consciousness Explained of how he repeatedly circulates drafts of papers to colleagues for comment. At least in philosophy, that's an important part of the process.
Having to show interim steps would make fraud much harder, and it's a zero-overhead thing if people are already backing up their work.
Your shower analogy doesn't work because there's no way to force people to post drafts. We already have the option posting of drafts. It's called personal websites and/or the arXiv.
According to the article, many people wasted a lot of time attempting to recreate their results. Further, the paper is highly-cited and it probably shaped, directly or indirectly, opinions, further research directions, and possibly even policy.
We are focusing on the Excel error, which was likely a mistake. I'm really struggling though to justify their choice of weighting/averaging. It's confounding that two highly regarded and experienced academics would choose something so particularly bad. It absolutely should have been noted in the text.
Jelte M. Wicherts, Rogier A. Kievit, Marjan Bakker and Denny Borsboom. Letting the daylight in: reviewing the reviewers and other ways to maximize transparency in science. Front. Comput. Neurosci., 03 April 2012 doi: 10.3389/fncom.2012.00020
http://www.frontiersin.org/Computational_Neuroscience/10.338...
on how to make the peer-review process in scientific publishing more reliable. Wicherts does a lot of research on this issue to try to reduce the number of dubious publications in his main discipline, the psychology of human intelligence. It appears that the discipline of economics research needs help with data openness too.
"With the emergence of online publishing, opportunities to maximize transparency of scientific research have grown considerably. However, these possibilities are still only marginally used. We argue for the implementation of (1) peer-reviewed peer review, (2) transparent editorial hierarchies, and (3) online data publication. First, peer-reviewed peer review entails a community-wide review system in which reviews are published online and rated by peers. This ensures accountability of reviewers, thereby increasing academic quality of reviews. Second, reviewers who write many highly regarded reviews may move to higher editorial positions. Third, online publication of data ensures the possibility of independent verification of inferential claims in published papers. This counters statistical errors and overly positive reporting of statistical results. We illustrate the benefits of these strategies by discussing an example in which the classical publication system has gone awry, namely controversial IQ research. We argue that this case would have likely been avoided using more transparent publication practices. We argue that the proposed system leads to better reviews, meritocratic editorial hierarchies, and a higher degree of replicability of statistical analyses."
I noticed more than one fishy looking 3rd party domain loading while the page downloaded, so went into Inspector to see what was up. There are one or more resources loaded from each of the following domains, in many cases including javascript...
2mdn.net, scorecardresearch.com, bizographics.com, tynt.com, optimizely.com, google.com, sail-horizon.com, facebook.com, 247realmedia.com,akamaihd.net, vizu.com, pubmatic.com, imrworldwide.com, advertising.com, googlesyndication.com, doubleclick.net, chartbeat.com, sharethrough.com, fbcdn.net, skimresources.com, gstatic.com, stumbleupon.com, tynt.com, adadvisor.net, youtube.com, shareth.ru, agkn.com, yimg.com
'shareth.ru' seemed particularly suspect, until I realized it was probably sharethrough.com trying to be cute.
There are so many domains being trusted here the drive-bys could have drive-bys.
This seems obvious, but established academics have a vested interest in not sharing. Sharing data and code opens their work up to impeachment (as is seen here), and it gives others a jumping off point to extend their work.
Both of these side effects of sharing code and data are bad for the careers of successful academics, but good for literally everyone else in the world (including, ironically, these academics). It should be a no brainer, but then again, who referees these papers? The very people who have the most to lose by such a change.
It will take a lot of clamoring from the outside to bring about a world where scientific work is considered not legitimate without full documentation of the experiment performed.
Perhaps the greater lesson is that we should trust proper journals for a reason.
http://johnbtaylorsblog.blogspot.com/2012/10/simple-proof-th...
(one of a host of commentary pieces on both sides... but essentially none in the middle)
I should program it into my business intelligence software, it's gotta be useful for something--like for cooking the books.
(Sarcastic, but it's real life)