I can understand it may be part of a meaningful personal journey for you, and I appreciate that. But if no one else can validate your research they're correct to discredit it and you.
So what is the optimal outcome here? Should we hold you to a standard of reproducibility even if it is as minimal as, "actually describe your algorithms correctly and don't misrepresent a piece of code and a paper?" Or should everyone just decide you can find your own research funding if it's not going to help anyone?
This is the actually important definition of "reproducible".
"has an install.sh" is really nice, but it's far more important that an informed reader can recreate the artifact for themselves from the written description.
The "must be able to apt install or it's not real science" is a particularly dangerous path to go down. It's how you end up with e.g. shit loads of well-engineered LISP or FORTRAN code with no actual scientific insights or knowledge transfer. Which has actually happened in the past. A lot.
Isn’t the whole point of a paper, to obviate needing to talk to the authors? Like, so that science is a ratchet that doesn’t slip backward the moment the authors die?
Word2vec is a great example of a piece of research that conveys a great idea with plausible results, where in fact, the numbers are less useful than the story behind them. Of course, google is a big corp that can just put out useful interesting stuff without worrying about how to play the science game.
Now, that said, reproducibility is terrible in many fields. CS has an opportunity to act as a trailblazer here, but it should be noted that this would be holding themselves to a higher standard than their peers in other fields. As a result, there's going to be a learning process for everyone as they figure out how to make this all work. :)
And some pretty good science got done before computer scientists were gifted to the world.
I'm genuinely skeptical that modern software engineering practices are a good way of thinking about reproduction in science. Even in computer science. There's a lot that scientists can learn from software engineering (and in fact I've helped run workshops in the past on exactly this topic), but science is not engineering.
I'm happy to talk about this if you want. One of the most important aspects of this work was that people like Dijkstra started using notions that approached what real computers could read while remaining human-readable. This is some measure of classical "reproducibility". And work like McCarthy's was revolutionary in part because it was a definition of reproducibility as a result!
I can give examples of shockingly good papers that are struggling to see the light of day in their industry because they're written in ways that make them hard not only to understand, but to reproduce.
So don't presume to lecture me about this. Part of the reason the word2vec paper stands out is precisely because this is such a deviation from the norm to have a paper misrepresent its most fundamental component: the algorithm.
But similarly, if you say, "Here are our statistical models, they're in a notation for a private, custom MCMC system you can't use" that's probably failing the standard even if the work is good.
I am one of those people who actually did the extra weeks/months to properly test/review/document/release my code and data sets (you can apt-get install my "research artifacts").
In retrospect, it was a poor use of my time and a poor use of my sponsoring institution's time. "apt-get install" is NOT what we mean by "reproducible" in science.
High-quality or easy to install code is not necessary for a result to be reproducible. True reproduction would mean coding the algorithm from scratch by following the written description, and that's what it literally means in most other fields of science.
You can't download & install a Large Hadron Collider in an afternoon. Does that mean the LHC experiments are "not reproducible"? Of course not.
But that's not even the important point. The really important point is that, in most cases, high-quality code is not even sufficient for a result to be reproducible! See: the article we're discussing.
IMO, the "blindly rerun the code" definition of reproduction is actually a HUGE barrier to creating a true culture of reproducability in computer science. It results in super lazy reviewing where "public source code that's easy to install and puts the correct-sounding shit into STDOUT" becomes a stand-in for "paper actually describes a novel idea in enough detail that it can be truly reproduced".
> So what is the optimal outcome here?
An optimal allocation of scientists' time and effort.
As a scientist who has actually done that leg work, I don't think packaging code so that it runs with a single click is the best use of public money in science in 99.9999% of cases. That time is much better spent on writing and other dissemination explaining the ideas that make the code work (in some cases well-documented source code is the best description but in other cases prose is much more effective and illuminating). Or on coming up with new ideas that are even better than the old ones.
Which I guess is just another way of saying that scientists should spend their time on science, not engineering.
P.S. When shitting on "scientists" for not being good enough software engineers, please remember who's going to be doing the actual work you're demanding. It's mostly phd students who make $30K/yr. And they have to do this work in their free time because their 60 hr/wk day job is fully allocated to doing the actual science. I.e., treat scientists who maintain their code as you would treat FOSS contributors who are making 5x-10x+ less than you while working longer hours. Because maintaining high-quality code is something they are almost certainly doing in their free time.
Which failed here. However, you're going to have a difficult time convincing many people that the attributes you described would be bad properties, just that they might not justify their expense.
> You can't download & install a Large Hadron Collider in an afternoon. Does that mean the LHC experiments are "not reproducible"? Of course not.
By the same token though, the LHC repeats experiments and solicits feedback on how to improve their methods, which they go to great lengths to publish and simulate, because they're aware of this problem.
> IMO, the "blindly rerun the code" definition of reproduction is actually a HUGE barrier to creating a true culture of reproducability in computer science. It results in super lazy reviewing where "public source code that's easy to install and puts the correct-sounding shit into STDOUT" becomes a stand-in for "paper actually describes a novel idea in enough detail that it can be truly reproduced".
Ah, yes. Yes. "If this code is TOO reproducible then people might reproduce it, and handwave handwave the quality of papers would decline.
That's certainly NOT the case in pure CS papers, which have only improved since the days when folks felt that "Lenses, Bananas and Barbed Wire" was how folks should go about writing papers.
Now, physics might be different. But there is surely a middle ground between, "I've shipped you a LHC just plug in in lol" and "This paper doesn't even remotely describe how we achived the results."
If you believe that wasting the time of scientists is bad, then surely you're for clear papers with accurate descriptions of the methods so that those who go and reproduce your work are not sent on wild goose chases?
> As a scientist who has actually done that leg work, I don't think packaging code so that it runs with a single click is the best use of public money in science in 99.9999% of cases.
No, we got that part. But surely someone does and maybe you can design your work to leverage that rather than reproducing and discarding scaffolding. My big concern here is that a lot of scientists (like you claim to be) are underqualified and unpracticed at software, and thus are surely seeing at least some aspect of their work distorted by software and hardware issues.
> Which I guess is just another way of saying that scientists should spend their time on science, not engineering.
Scientists are not going to be able to escape engineering. No one else is going to build what they need besides them.
> please remember who's going to be doing the actual work you're demanding. It's mostly phd students who make $30K/yr. And they have to do this work in their free time because their 60 hr/wk day job is fully allocated to doing the actual science.
Yeah, I'm aware. I suspect their lot would be better if your attitude wasn't that their work is disposable and unimportant.
Related to most ML code, it's not written in C but in Python frameworks that implement whole layers as building blocks. So it's not as hard to read/change. You can take something complex like the transformer and understand it in a few minutes. The most puzzling part with framework code is figuring out the shape of tensors and, sometimes, what each dimension means.
Regarding reproducibility - it's hard to achieve on account of parallelisation. You see, neural nets use floats, and floats are not real numbers. Half the float numbers lie between [-1, 1] and the rest outside. If you do 1e-10+1e10 you get 1e10. So if you add 1000 times 1e-10 to 1e+10 you still get 1e+10, but if you add them up together first, the result is different. So float summation is not commutative. Depending on race conditions the order in which summation occurs could change and the result change as well, even if you use the same random seeds. And if we give up parallelism then we can't run the experiments any more.
That's if float rounding issues are even a big problem in the first place. If your results are within .1% over a few runs it's reproducible enough for most purposes.
How wrong should the original work be before it ceases to be a case of 'clean your own pipettes'?
Could you PLEASE consider reading the article we're all discussing before you roll into the comments section of an article about it with your strong-but-loosely-held opinions? You're arguing against a point that almost no one is putting forward (that software should be one click).
> I have to fix bugs before the code can run. I'm very curious how did the author run that code with the bugs.
Pro tip: at paper submission time, md5sum your code and also git tag it in your private repo. When you release the code, if you have to / want to release with a clean history, make submission-time your initial commit and then make the current state of the repo your second commit. I've never encountered an institution that won't allow that level of history in the code release, even places that are pretty hard core about wiping pre-release history.
The software types love to sht on researchers for their poor quality code. Yet a few threads over, there’s always a discussion about how more tests don’t make it onto the kanban board (or whatever you guys are using these days) when crunch time approaches. It’s not just us.
Sorry but I don't trust your research just based on yours words and cherry picked images.
Also research is incremental, so producing a proper code on which other can work on top of should be part of the CONTRIBUTION.
A very large number of research papers I've read in my life were so poorly written as to being near impossible to understand.
The time it took to infer the actual meaning of certain parts of certain papers was vastly larger than just reading the code would have been.
It has become very clear to me that being a strong CS researcher does not necessarily imply mastery of language or even the very basic ability to explain one's ideas. As a matter of fact, with very few exception (eg Feynmann), the two skills seem to be at odds with one another.
On the other hand, code can be read and understood much more readily, especially when it can be experimented with, if only to insert printfs in it to understand what it does.
The culture of scientific publication in CS must change. The "I'm pressured by my advisor and therefore have no time to publish my code" argument is complete bollocks.
The code should come first and your advisor should pressure you to publish that first (assuming it actually works).
If someone wants to read your poorly worded explanation of how it works, fine, publish a paper.
Show me the code first.
Being able to compile in some way and, possibly after some bug-fixing/compiler-pleasing yielding useful results that approach/relate to those in the paper, greatly increases my confidence in the researcher and his work. Pretty code is nice, but it's weighed separately, not affecting the appreciation for the original research work itself. Only the existence of code with an AGPL or less restricted license affects how I perceive the paper/original research.