How a coding error fueled a dispute between two condensed matter theorists
physicstoday.scitation.org
physicstoday.scitation.org
Simulation code contains a lot of numerical tricks and hacks - and the papers never discuss them.[1] So reproducing the results of a paper were mostly pointless, and the general practice was "Let's see if we can reproduce their findings using our technique". If you got somewhat similar results, you'd publish. If you didn't, you'd sometimes publish (remember, there's a bias against negative results).
[1]In fact, one reviewer even insisted that those details be removed because the journal was about publishing interesting new findings in science - not interesting techniques to make the computation more amenable.
The real 'mistake' is allowing the drama to escalate to the point that it is toxic. People who are truly interested don't care about who is right and who is wrong.
When disagreements like this come up, having common and good test cases are probably the most important (and is indeed the way in which the problem, generation of non-Boltzman behavior, was found).
In the end ST2 itself is flawed, and Princeton admitted that the discussion of water was not significantly advanced through this drama. Is it worth it?
Understanding the argument and its importance should be the focus.
Sharing code is not
always the answer
Why let the perfect be the enemy of the good?Sharing code is not the full answer, especially in case of Monte-Carlo simulations in physics, because that kind of algorithm is hard to test: What is the testing oracle? What is the specification?
But sharing code is part of the answer.
Setting up a culture where it is unacceptable to submit a paper without open-sourcing the code and suitable testing (for simple edge ases), and suitable scripts that make reproducing the software simulations easy, is good scientific 'hygiene'. See for example [1, 2] for efforts towards reproducible software submissions in computer science.
Reproduciblility is the very essence of the scientific method.
You might want to be careful about using that "was it worth it?" argument - someone might reasonably ask, was it worth funding either team? If science doesn't show that it does good work, then its funding can be more easily questioned.
A) embarassement at having other people see how bad it is. Not necessarily wrong, but bad in other ways.
B) wanting to keep an advantage by continuing to publish on top of the work already done, whereas anyone else wanting to get in on the same idea would have to first re-build the original work.
Not pre-publishing a research thesis / methodology is bullshit.
Not publishing data is bullshit.
Not publishing exact code (and runtime, etc) is bullshit.
Not publishing when research is complete is bullshit[0].
Not publishing is bullshit. I don't care who your funder is, it should be illegal for funders to selectively publish research. It's obviously one sided propaganda. The only time not publishing makes sense is when there are national security or global security ramifications to the research or if funding needs to be lined up for patents.
[0] I've seen a PhD candidate delayed 4 years from publishing because her supervisor was trying to coordinate a team to publish with research along the same vein all at once to make it a groundbreaking set of discoveries. It's basically p-hacking by some other name.
Anyways, we were working on a fairly large company-wide effort. We were migrating from 32-bit to 64-bit, had 2 OS upgrades (new version of RHEL, migrating from Windows XP to Windows 7), 2 compiler upgrades (I forget the gcc versions - it came with the upgrade, and on Windows VS2003 to VS2008). Because of the coupling of the libs & services, both the Linux & Windows sides had to be released at the same time. We could update the Windows boxes at anytime and run the old software on it, but we basically couldn't develop the old software on Win7 as VS2003 wasn't compatible (it could be made to work, but the IDE or compiler or linker may fail randomly and hard). This is just to explain the scope of the effort we were making.
Back to why I mentioned the core files. To anyone that's done it, it's obvious the magnitude of effort doing this all at once is. There will be bugs. Specifically a metric-shit-ton of them. Those core files were needed. They stopped being generated because our shared core space filled up. The culprit? A PhD quant running some model in Python using some C/C++ extensions kept crashing Python, each one dropping a multi GB core file. This one quant was single-handedly using over half of the shared space. When confronted/inadvertently shamed in a development-wide email (we sent out usage breakdowns per user/project when a usage threshold was exceeded), his response was golden: "What are these core files, and how do I stop them from being generated?" Uniform response from everyone else: "Fix your damn code!" Mind, at the time, this was in front of ~600 developers.
The kicker: the reason this whole effort was being made? Because someone sent the CEO an XLSX spreadsheet in 2008 (the exact year may have been later, but immaterial to the story), and still being on WinXP & Office 2003, he couldn't open it. So...down the rabbit hole we went.
As for B, that harkens back to the pre-scientific attitudes of medieval alchemy.
Researchers aren't paid (and usually aren't trained) to write software that can be installed on more than one computer. I don't like that situation either, but that how it currently is. Research code also suffers from accelerated bit rot. The dependencies are often research code themselves and projects are abandoned when the results are published.
It took about a week before
I was happy with it.
And what made you thing that "I was happy with it" is enough from the point of view of the pursuit of scientific truth (e.g. how did you deal with the problem of confirmation bias)? More than half a century of experience with software engineering has told us the hard way that it's near impossible to write bug-free code. And where this happens (e.g. aviation software) it takes extreme dedication and resources.There is a reason that Alan Turing invented program verification in the 1940s [1].
Let me finish with a quote from computing pioneer M. Wilkes [2]: "I well remember when this realization first came on me with full force. The EDSAC was on the top floor of the building and the tape-punching and editing equipment one floor below. [...] It was on one of my journeys between the EDSAC room and the punching equipment that "hesitating at the angles of stairs" the realization came over me with full force that a good part of the remainder of my life was going to be spent in finding errors in my own programs."
[1] A. M. Turing, Checking a Large Routine. http://www.turingarchive.org/browse.php/b/8
[2] M. Wilkes, Memoirs of a Computer Pioneer, MIT Press, 1985, p. 145.
I don't think you can reasonably expect software that extracts features from photographs to be proven correct. What would the specification even look like? How many man years would you spend on verifying OpenCV and the two dozen other dependencies you rely on? It's not like all math papers provide machine checked proofs and that would probably be an easier endeavor.
Then when you edit any parts of the code and the test suite breaks - you can see that some of the invariants had been broken.
This is even more important for python, as at least having a simple test suite can already tell whether your environment is sane. And that's about 90% of the newcomer's time saved.
I would contend that if you have proven to yourself that your code works and are therefore capable of proving it to other folks (via e.g. solid testing), you should not be ashamed of the spit and glue.
It is research code, we all know what that looks like. But if, on the other hands you haven't proven to yourself it works, then it's definitely something to be ashamed of - scientifically speaking.
Most academics write code to just work. Not work well, or to be generalized, or to be efficient, just work. And while that’s absolutely fine, as your results being reproducible from the code is all that really matters, a lot of people don’t see it that way and will only see code slapped together haphazardly and dismiss you because of it.
Imagine if this reproducibility excuse were applied to experimental results and technique: we don't have be careful or explain in detail what we are doing, as reproducibility will take care of any errors. One consequence would be that, as the current state of knowledge became less certain, it would become less clear what to do next.
Publishing scientific papers that no one can re-implement does, and hugely so.
Therefore, not a very adequate comparison.
Especially in fields like physics there are many limits that you can derive analytically. Some times those are highly non-trivial.
I do appreciate papers that give you either a reference implementation or at least enough details on how it was done. I spent too many hours on papers that didn't trying to reproduce their claimed results...
That's what this license is for: http://matt.might.net/articles/crapl/
In todays rockstar-obsessed recruiting hell, I would not want any quick-and-dirty hacks published under my name.
There's a lot to be said for the IETF's policy of "rough consensus and running code".
Having reviewed papers on multiple occasions, I wouldn't spend much if any time looking at the actual source code if it is available. Every paper submission is rushed because it needs to be ahead of what anyone else is doing (at least in CS and related fields).
(A deeply cynical and suspicious part of me suspects this is because the prime motivator for "peer reviewed journals" is to make Elsiver richer, and academic knowledge sharing and reputation building is just a side effect of that...)
Every peer review process is different from the next, it varies by journal, by conference, and by month of the year. I probably can't emphasize enough how highly subjective every peer review is. Anyone that has a paper rejected should consider the validity of reviewer comments (which may be total BS on occasion), then edit and resubmit the paper to another venue.