It reminds me of Kerckhoffs's principle in cryptography, which states: A cryptosystem should be secure even if everything about the system, except the key, is public knowledge.
It reminds me of Kerckhoffs's principle in cryptography, which states: A cryptosystem should be secure even if everything about the system, except the key, is public knowledge.
In science, code is not an end in-and-of-itself. It is a tool for simulation, data reduction, calculation, etc. It is a way to test scientific ideas.
> how do you expect anyone with the right expertise to assess your findings
I would expect other experts in the field to write their own implementation of the scientific ideas expressed in a paper. If the idea has any merit, their implementations should produce similar results. Which is exactly what they would do if it were a physical experiment.
And if you're a map maker, it's a bit rich to start claiming that the accuracy of your maps is unimportant. If code is "a way to test scientific ideas", then it kinda needs to work if you want meaningful results. Would you run an experiment with thermometers that were accurate to +-30° and reactants from a source known for contamination?
Of course, it is a difference whether you make a clinical study on drugs, and use a pocket calculator to compute a mean, or whether you research in numerical analysis, or are presenting a paper in how to use Coq to more efficiently prove the four-color theorem or Fermat's last theorem.
In short, much of science is not computer science, and for it, computation is just a tool.
If I'm given bad information and I act on that information, then problems can occur.
Similarly, if the software is giving the scientist bad information, problems can occur.
How many more stories do we have to read about some research getting published in a journal only to have to retract it down the road because they had a bug in the software before we start asking if maybe there needs to be more rigor in the software portion of the research as well?
There was a story on HN a while back about a professor who had written software, had come to some conclusions, and even had a Ph.D. student working on research based on that work. Only to find out that a software flaw meant the conclusions weren't useful to anyone and that student ended up wasting years of their life.
---
This stuff matters. This isn't a model of reality, it's an exploration of reality. It would be like telling a hiker that terrain doesn't matter. They would, rightfully, disagree with you.
We will always hear stories like that, as we will always hear stories about major bugs in stable software releases. Asking a scientist to do better than whole teams of software engineers makes little sense to me.
Of course, a bug that was introduced or kept with the counscious intention of fooling the reviewers and the readers is another story.
This is not what is being asked, shame on you for the strawman.
Your entire post can be summed up with the following sentence: "if we can't be perfect then we may as well not try to be better".
The thing is that it has little to do about rigor -- or if I may sin again, it is equivalent to say that software developers lack rigor: sure, some of them do (as some scientists do), but even among the most significant and severe bugs of the history of software, it is seldom the case that we can tell "right, definitely the guy who wrote that lacked rigor and seriousness".
Of course this is not a blank forgiveness for every bad scientist out there. Of course we should aim at getting better. But we should make the difference between the ideal science process and the science as performed by a human, prone to errors, misunderstandings and mistakes, and realize that these things will always happen, however many times we call for "more rigor because bugs have consequences".
And stop comparing scientists to software developers, it's a hidden argument by authority, and it isn't needed.
How this "rigor" you are calling for should manifest, then? Put bluntly, my point was that every software has bug, so how "more rigor" would help? What should we do, what should we ask for in _practical_ terms?
Also, please do not rephrase this last sentence as "oh so since every software has bugs, then you obviously say that we shouldn't fix bugs, anyway other bugs will remain!".
That's exactly what I'm going to do. Point out that we can demand better even in the face of a lack of perfection.
There are two problems here with your stance.
1. The assumption that all bugs are created equal, and 2. The assumption that the truth isn't the overriding concern of science.
It's real easy to define the set of bugs that are unacceptable in science. Any bug that would render the results inaccurate is unacceptable.
The fact that some jackass web developer wrote a bug that deleted an entire database in no way obviates that responsibility of the scientists.
I think that, while we could use a bit more training in software engineering best-practices in the science, the thesis is still that science is hard and we need real replication of everything before reaching important conclusions, and over-focusing on one specific type of errors isn't all that helpful.
It's not clear to me why you think I would argue that inaccuracies should be avoided in software but accept that they're ok for electrical systems.
Sure, some people might point out spurious bugs and "design issues" or whatever, boo hoo. But others might actually find flaws in the code that meaningfully affect science itself: true bugs.
Sure, they could do this by doing a full replication in a lab and then custom coding everything from scratch. But even then, all you have is two conflicting results, with no good way yet to determine which one is more right or why they disagree. Technically, you can use the scientific progress to eventually find bugs in the scientific process, but why waste so much time when publishing the code will allow for reviews to find bugs so much faster. Its pure benefit to science to not obscure its proofs and rigor.
How many actually try to reproduce the results by writing corresponding code themselves? Apparently lot of papers with slightly wrong findings because code errors have passed the peer review (all of us in the SWE bubble know how often bugs occur), at least in less prestigious journals.
There is nothing wrong with mandating the code to be supplied with the paper. Because, many time code is somewhere between the experimental setup and proof / result.
I myself had very bad experience with extending the undocumented Fortran 77 code (lots of gotos and common blocks) of my supervisor. Finally, I decided to rewrite the whole thing including my new results instead of just somehow embedding my results into the old code for two reasons: (1) I'm presumably faster in rewriting the whole thing including my new research rather than struggling with the old code and (2) I simply would not trust in the numerical results/phenomenology produced by the code. After all, I'm wasting 2 months of my PhD for the marriage of my own results with known results which -in principle- could have been done within one day if the code base would allow for it.
So yes, If it's a one-man-show I would not give too much on code quality (though unit tests and git can safe quite a lot of time during development) but if there is a chance that someone else is going to touch the code in near future it will save time to your colleagues and improve the overall (scientific) productivity.
PS: quite excited about my first post here
This makes me a little uneasy, as I'm not too worried about code quality can easily translate into Yes I know my code is full of undefined behaviour, and I don't care.
> PS: quite excited about my first post here
Welcome to HN! reddit has more cats, Slashdot has more jokes about sharks and laserbeams, but somehow we get by.
The latter isn't great practice, but if your environment handles behavior deterministically, and you publish the version of the compiler you're using, it doesn't seem to be a problem for this type of code.
'Undefined behaviour' is a term-of-art in C/C++ programming, there's no ambiguity.
> if your environment handles behavior deterministically, and you publish the version of the compiler you're using, it doesn't seem to be a problem for this type of code.
Code should be correct by construction, not correct by coincidence. Results from such code shouldn't be considered publishable. Mathematicians don't get credit for invalid proofs that happen to reach a conclusion which is correct.
Again, this isn't some theoretical quibble. There are plenty of sneaky ways undefined behaviour can manifest and cause trouble. [0][1][2]
In the domain of safety-critical software development in C, extreme measures are taken to ensure the absence of undefined behaviour. If scientists adopt a sloppier attitude toward code quality, they should expect to end up publishing invalid results. Frankly, this isn't news, and I'm surprised the standards seem to be so low.
Also, of all the languages out there, C and C++ are among the most unforgiving of minor bugs, and are a bad choice of language for writing poor-quality code. Ada and Java, for instance, won't give you undefined behaviour for writing int i; int j = i;.
[0] https://devblogs.microsoft.com/oldnewthing/20140627-00/?p=63...
[1] https://blog.regehr.org/archives/213
[2] https://cryptoservices.github.io/fde/2018/11/30/undefined-be...
See also my longer ramble on this topic at https://news.ycombinator.com/item?id=24264376
Let the scientists publish UB code, and even the artifacts produced, the executables. Then, if such problems are found in the code by professionals, they can investigate it fully and find if it leads to a tangible flaw that invalidates the research or not.
You would drive yourself mad pointing out places in math proofs where some steps, even seemingly important ones, were skipped. But the papers are not retracted unless such a gap actually holds a flaw that invalidates the rest of thr proof.
Let thdm publish their gross, awful, and even buggy code. Sometimes the bugs don't effect the outcomes.
Granted, it's not a guarantee that the results are wrong, but it's a serious issue with the experiment. I agree it wouldn't generally make sense to retract a publication unless it can be determined that the results are invalid. It should be possible to independently investigate this, if the source-code and input data are published, as they should be.
(It isn't universally true that reproduction of the experiment should be practical given that the source and data are published, as it may be difficult to reproduce supercomputer-powered experiments. iirc, training AlphaGo cost several million dollars of compute time, for instance.)
> this mindset is what keeps people from publishing the code in the first place
As I explained in [0], this attitude makes no sense at all. It has no place in modern science, and it's unfortunate the publication norms haven't caught up.
Scientific publication is meant to enable critical independent review of work, not to shield scientists from criticism from their peers, which is the exact opposite.
> Let the scientists publish UB code, and even the artifacts produced, the executables. Then, if such problems are found in the code by professionals, they can investigate it fully and find if it leads to a tangible flaw that invalidates the research or not.
I'm not sure what to make of 'professionals', but otherwise I agree, go ahead and publish the binaries too, as much as applicable. Could be a valuable addition. (In some cases it might not be possible/practical to publish machine-code binaries, such as when working with GPUs, or Java. These platforms tend to be JIT based, and hostile to dumping and restoring exact binaries.)
I agree with your final two paragraphs.
Glad we agree, if you're aware of how your compiler handles these things, you can construct it to be correct in this way.
It won't be portable at all (even to the next patch version of the compiler), I would never let it pass a code review, but that doesn't sound like an issue that's relevant here.
I presume we agree but I'll do my usual rant against UB: Deliberately introducing undefined behaviour into your code is playing with fire, and trying to outsmart the compiler is generally a bad idea. Unless the compiler documentation officially commits to a certain behaviour (rollover arithmetic for signed types, say), then you should take steps to avoid undefined behaviour. Otherwise, you're just going with guesswork, and if the compiler generates insane code, the standards documents define it to be your fault.
It might be reasonable to make carefully disciplined and justified exceptions, but that should be done very cautiously. JIT relies on undefined behaviour, for instance, as ultimately you're treating an array as a function pointer.
> It won't be portable at all (even to the next patch version of the compiler)
Right, doing this kind of thing is extremely fragile. Does it ever crop up in real-life? I've never had cause to rely on this kind of thing.
It would be possible to use a static assertion to ensure my code only compiles on the desired compiler, preventing unpleasant surprises elsewhere, but I've never seen a situation where it's helpful.
This isn't the same thing as relying on 'ordinary' compiler-specific functionality, such as GCC's fixed-point functionality. Such code will simply refuse to compile on other compilers.
> I would never let it pass a code review, but that doesn't sound like an issue that's relevant here.
Disagree. It should be possible to independently reproduce the experiment. Robust code helps with this. Code shouldn't depend on an exact compiler version, there's no good reason code should.
Sounds like it is quite good science to do that, because it puts the computation on a pair of independent feet.
Otherwise, it could just be that the code you are using as a bug and nobody notes until it is too late.
The cross-checking is anyways good scientific practise, not only because of bugs in the code (that's actually a sub-leading problem imho), but because of the degree of difficulty of the problems and the complexity of their solutions (and their reproducibility). In that sense, cross-checking should discover both, scientific "bugs" and programming-bugs. The "debugging" is partly also done at the community level - at least in our field of research.
However, it is also a matter of efficiency. I -and many others too- need to re-implement not because of bug-hunting/cross-checking but simply because we do not understand the "ugly" code of our colleagues and instead of taking the risk to break existing code we simply write new one which is extremely inefficient (others may take the risk and then waste months on debugging and reverse-engineering which is also inefficient). So my point on writing "good code" is not so much about avoiding bugs but about being kind to you colleagues, saving them nerves and time (which they can then spend on actual science) and thus also saving taxpayers money...
Yes, that should be this way.
Also all cases where some company research team goes to a scientific conference and presents a nifty solution for problem X without telling how it was purportedly done, it should be absolutely required to publish code and data for this.
*And that's also something which is broken about software patents - patents are about open knowledge, software which uses such patents is not open - this combination should not be allowed at all).