GitHub issue calling for retraction of Imperial College study for codebase flaws
github.com
github.com
Yes, I wish the original, apparently C code, had been released but let's fix the bugs in the code.
The issue claims the tests are broken because they're checking for checksums instead of actual data -- but what they really do is to check for the sanity of the implementation to be actually implementing the mathematical models that are underlying to it.
That's the beauty of open-sourcing the code: people can help or verify. Needlessly shitting on other people's work without any proof is just disheartening to me.
With that said, if a research paper's main contribution is a model whose results were evaluated using a code with major bugs then retraction of the paper is only professional.
There's no proof of that at all in this case. The person complaining didn't find any scientifically valid problem, only that he personally for some not cleanly stated reason doesn't like the tests provided, which is clearly stupid.
Since when should any tests be such that some random guy on github must like them? If the code was used by the experts, and they maybe even used it for years, who says that these experts are in any way obliged to publish all their logs of their use of that code (which doesn't have to even exist in publishable form)?
Even if all that they possibly iteratively did for potentially years could be condensed to some tests, who says that it wouldn't take too long (as in, months, years) for that? It's practically just a matter of good will that the code is published at all, in any form at all.
The overall conclusion may be correct, who knows, but based on my experience, I do not believe, at all, any numerical predictions of this code.
That does lend itself to retracting the paper.
This I take issue with.
Yes, it is a political proclamation without any practical substance, and should be recognized as such.
The 116 is "if I use it this special way I would expect something else and not what I see" which can be simply answered with "well, don't do THAT."
The simulations by design are not expected to produce exactly the same results in different runs. That's why they are simulations. The reporter tried some "partial save" and expected something else.
The 30, if I understand correctly, is observed on Cray to behave in some minor detail not exactly the same as on the PC (which again doesn't have to even mean that the output is scientifically wrong). Seriously? If one is a Cray user, he can fix it, so what?
https://github.com/mrc-ide/covid-sim/issues/168
Looks linked to the awful style, though. When you have 700 LOC functions starting with
int i, j, k, l, m, i1, i2, j2, l2, m2, tn;
it does tend to happen that you reuse a variable that you shouldn't have reusedJust smells like FORTRAN to me.
Did this model provide clarity and insight for the decision making process? Or, did it instead induce a false sense of panic?
Here's some more info as to what this repo does and does not contain, and why what was released is not enough.
https://www.aier.org/article/imperial-college-model-applied-...
People in industry don't always do better, either. Look at the notebook code shared on Github by Kevin Systrom for the rt.live site. It's just as messy. But the point is that he shared it and that helps everyone see what is going on and lets anyone improve on the work.
Publicly sharing model code on Github is a recent development and is amazing for everyone. It gets the work out faster, allows for feedback on methodology and coding style and generally lets everyone get smarter over time. We shouldn't discourage this.
Posting retraction demands in Github issues based on coding style is not helpful and shows a lack of communication skills. What is helpful is providing useful, specific bug reports, submitting improvements, or sharing your own, improved model. If you want to demand a retraction, do it via alternate means.
Remember, this crisis appeared suddenly on the world stage and whoever had a model sitting around got thrust into the forefront of public discussion. But scientists don't make policy. Policy makers make policy. A github issue is not the place to blame someone for a policy you disagree with. That only discourages people from sharing their work which would be a loss for everyone.
Your reading is very charitable. The political debate is a complete gutter, and everyone is ready to jump at absolutely anything that can advance his argument in the slightest.
As you can see in this very thread, jMyles is hardly unbiased on the matter. He's just playing politics, code standards are an excuse.
So hard in fact, that I am about to invest time to fix that mess (because I like the project).
I think it's important to understand that we only require information to generate algorithms written in the paper. As such, assumptions and limitations should be listed enough to generate the material in the paper. One way to do so is code and data, it is certainly not the only way.
We don't expect universities to build their own laboratories and halls of residence from first principles - they will call on property developers, architects, and contractors. It would be good if there were more IT services, built by specialists, that academics could tap when required - so they would spend more time advancing their fields of expertise and less time writing mediocre code. Things like PythonAnywhere / Notebooks, Lambda etc are a step forward, but there is still ample space for advancement.
Even for real reproductions, that require work just like they do in any other scientific domain, the points in that issue seem pretty irrelevant. There's things like FAIR standards on reproducibility but as far as the status quo goes that repo doesn't look too bad. I could not care less if tests for some project are badly written, at least it's in written in a non-obscure language and shows a somewhat sane structure. What's next? Calling for redactions because somebody didn't follow the same tabs vs spaces paradigm?
There's nothing in that issue w.r.t. whatever scientific finding this was used for and this general phenomenon is really fascinating. Instead of a constructive discussion with the technical or scientific folk that put their work out there, or engagement with the politicians that drew conclusions based on those scientific findings, you get github issues and Twitter threads mixing a bunch of unrelated concerns.
Maybe a very dedicated reviewer might check a short simulation code that can be written in a few hours max, but even that is exceptionally rare in my experience. Reviewers usually look at the results and validation simulations in the manuscript, and if they look reasonable trust the code doesn’t have major bugs affecting accuracy.
Even if it is re-implemented and verified it won't be the same. There are a million possible conditions for "re-implementation". Where would you draw the line? How many "re-implementations" and "types of conditions" do you need to be sure?
You probably need to have expectation mismanagement from a scientific paper. A scientific article will only say "We tried this idea under conditions X,Y,Z and it works with A,B,C metrics and I, J, K assumptions". That is the crux of ANY empirical science - Computer Science or otherwise. That's all that they are paid (and incentivised) for. A computational scientific experiment is not a software product. If you want more, you need to pour in more funding and give them more resources explicitly for those purposes. Either that, or if you want to be 100% verified - you can choose to read theoretical papers where mathematical proofs are "verification".
> "We tried this idea under conditions X,Y,Z and it works with A,B,C metrics and I, J, K assumptions".
Is what I'm calling the assumption (the whole sentence). The implementation is essentially the methodology, and IMO therefore needs to be included.
Note I'm arguing against 'academic, paper-supporting, code does or should not need to be open sourced', I'm not saying it needs to be 'software product' quality, written in the same way, to the same expectations, packaged, or anything like that. Just available for someone to say 'wait a second, you didn't try it under conditions, X,Y,Z, because Y gets negated here', or whatever.
And obviously the 'risk' of that happening isn't a reason not to publish it - you don't not publish the paper for risk that peer review reveals an error!
How about you write tests that clearly prove that the results from this simulation are absolutely wrong? The code is right there! And you're a "software engineer"! Then start talking about retracting papers.
This reminds me of the horror mainstream code monkeys express when they see the source code for automotive firmware. Thing is automotive people functional test the shit out of everything. And they don't let jr developers refactor proven good code because they don't like the way it looks.
I wouldn't expect research scientists to write an elegant and well structured program that ticks all the coding practices boxes at the first time as they are not "software engineers" unless that's their area of research or expertise. They want to present their results and the "code quality" comes secondary to them which can be done later. Surely the Linux source code wasn't cleanly structured in its first open-source release.
However, a quick skim at the source, one may suggest that the authors were writing C++ in a style of a C programmer. Maybe one can run a clang-analyzer on the source to find all sorts of issues, I guess.
I specialise in legacy code. I've seen my fair share of abysmal code that works. This is pretty awful, but far from the worst I've ever seen in my career -- even in systems where public safety was critical. I've not looked at this code in any real detail, but I suspect John Carmack (who has doubtless seen a fair amount of complex C code in his life) has it right.
On the plus side, if anyone wants to practice refactoring gnarly C code: here's your opportunity.
He explicitly writes there:
"it turned out that it fared a lot better going through the gauntlet of code analysis tools I hit it with than a lot of more modern code. There is something to be said for straightforward C code. Bugs were found and fixed, but generally in paths that weren't enabled or hit."
"the performance scaling using OpenMP was already pretty good, and this was not the place for one of my dramatic system refactorings. Mostly, I was just a code janitor for a few weeks, but I was happy to be able to help a little."
and
"I can’t vouch for the actual algorithms, but the software engineering seems fine."
As shown, he even points that the believes "a lot of more modern code" would have been worse than that "straightforward C", which matches my experiences.
But as I said in another comment, let's blame whoever listened to that garbage fire of a model, not the model itself. We don't listen to Minecraft modders for structural engineering advice either, and we did, they would not be the ones to blame for collapsing buildings.
[1] https://www.theguardian.com/world/2020/mar/12/coronavirus-ir...
Also remember, all models are wrong, some are useful. Is this particular model useful? Probably.
There is the great (Carl Sagan?) quote "extraordinary claims require extraordinary evidence". If other parts of the puzzle are of the same quality, I want to see heads rolling. The PCR test for example.
However, it seems that the author of the GitHub issue has personal/political problems with the Imperial researchers because of the lockdown: https://pastebin.com/LadbM3E1.
@OP: So let's get this straight, you created that GitHub issue within one hour of seeing the repository, rallied whatever Discord server that is to join your cause and put it on HN for reach?
As an academic and software engineer, please point me to where I can file an issue for you to retract your issue.
You don’t do testing via unit and integration tests, you do testing through simulating known systems and building confidence in code correctness over time. And by reviewing the mathematics behind the models, and reviewing that the code matches the math.
That’s harder for outsiders who want to discredit you to dig into that and point fingers at it. But not at all impossible, however in the science community we don’t start pointing fingers until we have actual criticism to back our critique.
Doesn't this put the code far above most academic code by having some tests?
lol
This study was, of course, the basis for lockdowns around the world.
[citation needed]
Afaik this was a lone study that might have been relevant only to the UK debate (and very late at that). Italy had already been in full lockdown for weeks before it was published, quite a few other European countries had gone in shutdown or closed borders, and obviously Asian countries had already taken countermeasures for months.
Being so cavalier with the truth is, of course, why intelligent people don't take seriously a lot of anti-science attitudes.
But I don't really blame Ferguson. If President Macron listens to him, or Alex Jones, or any other crank or random guy with an opinion and a model, instead of his own experts, he is the one to blame. I think the culture of the "safety principle" and excessive caution is also to blame.
https://www.lemonde.fr/planete/article/2020/03/15/coronaviru...
[0] https://metro.co.uk/2020/02/25/towns-italy-lockdown-coronavi...
As I said, "around the world" they were already taking countermeasures.
There is also a "control" group of sorts in us Swedes, as the Swedish health authority choose to act differently, initially. Their recommendation as of now is essentially equivalent to quarantine, except for non-adults. It's important to understand that it is really to instigate anything like a curfew in Sweden, at least unless we get attacked military by another country. The laws as they stand don't really allow that. So the government is also somewhat limited in what they can do.
The situation here is that urgent care has been overwhelmed for about a month, and a lot more people have died or been gravely ill than during even really bad flu seasons. We're also probably far from done yet, as it's much easier to get people to stay indoors when it's cold and rainy outside, and as its been getting warmer more people have been disrespecting the recommendations.
Its model for China predictions seems to be quite fitting.