I call them insincerely opensourced projects.
I call them insincerely opensourced projects.
That should be left to TensorFlow model zoo or GluonCV model zoo. I simply look at these research "open-source" as reproducible research.
This is why I think jupyter notebook and things alike are so important as the future publishing media. Reproducibility is very important.
It's just good news when you can find source code at all instead of just being told vague things about something the author did.
Absolutely correct, and not in a good way
Though to be fair, people doing research have other interests besides maintainable code (though half or more of the annoyances will probably come and bite them later)
This is extremely common, and I don't see why it should be allowed. I could easily write a paper claiming some kind of interesting result, fake a few "ground truth vs my algorithm" pairs, and talk a bit of mathy-sounding piffle about what it does - and it would be indistinguishable from many papers I have read.
In my opinion - provide runnable source, or what you have isn't a paper, it's a boast.
This sad situation is both bad for open source and science. Bad for science since inaccurate results that probably aren't even reproducible aren't the knowledge advancements we need. Funding agencies like NSF need to change their policies to address the root cause here. Universities will follow where the incentives lead them.
One big challenge the community faces is that if you want to get a paper published in machine learning now it's got to have a table in it, with all these different data sets across the top, and all these different methods along the side, and your method has to look like the best one. If it doesn’t look like that, it’s hard to get published. I don't think that's encouraging people to think about radically new ideas.
Now if you send in a paper that has a radically new idea, there's no chance in hell it will get accepted, because it's going to get some junior reviewer who doesn't understand it. Or it’s going to get a senior reviewer who's trying to review too many papers and doesn't understand it first time round and assumes it must be nonsense. Anything that makes the brain hurt is not going to get accepted. And I think that's really bad.
Geoff Hinton interview:
https://www.wired.com/story/googles-ai-guru-computers-think-...
I guess people look at statistical machine learning and deep learning, see all the formulae and hear all the calculus terminology and think - "oh, wow, that's a really rigorous field! Look at all the formalisms!".
But it's not. It's an extremely, almost exclusively, empirical field. The mathiness and the formulae are just unfortunate attempts to pass off the whole endeavour as something that it's not- some kind of careful science that uncovers deep truths about intelligence and cognition. In truth, it's all just about beating other peoples' systems in very specific benchmarks.
If it wasn't for this culture of pretensions to higher science, machine learning papers would most likely be written with much more clarity than they are now and mistakes like the one described in the above article would be rare.
Fully agree, and it's a necessary disease in young fields like ML (akin to grid search in fact).
But at some point, there will need to be some sort of theoretical foundation brought to bear or advancement will grind down to a halt
And the academic reward mechanism needs to start reflecting that fact.
You can get around it if you show BigName is wrong but if they've been vague enough that's hard.
I can understand it may be part of a meaningful personal journey for you, and I appreciate that. But if no one else can validate your research they're correct to discredit it and you.
So what is the optimal outcome here? Should we hold you to a standard of reproducibility even if it is as minimal as, "actually describe your algorithms correctly and don't misrepresent a piece of code and a paper?" Or should everyone just decide you can find your own research funding if it's not going to help anyone?
Now, that said, reproducibility is terrible in many fields. CS has an opportunity to act as a trailblazer here, but it should be noted that this would be holding themselves to a higher standard than their peers in other fields. As a result, there's going to be a learning process for everyone as they figure out how to make this all work. :)
And some pretty good science got done before computer scientists were gifted to the world.
I'm genuinely skeptical that modern software engineering practices are a good way of thinking about reproduction in science. Even in computer science. There's a lot that scientists can learn from software engineering (and in fact I've helped run workshops in the past on exactly this topic), but science is not engineering.
I'm happy to talk about this if you want. One of the most important aspects of this work was that people like Dijkstra started using notions that approached what real computers could read while remaining human-readable. This is some measure of classical "reproducibility". And work like McCarthy's was revolutionary in part because it was a definition of reproducibility as a result!
I can give examples of shockingly good papers that are struggling to see the light of day in their industry because they're written in ways that make them hard not only to understand, but to reproduce.
So don't presume to lecture me about this. Part of the reason the word2vec paper stands out is precisely because this is such a deviation from the norm to have a paper misrepresent its most fundamental component: the algorithm.
But similarly, if you say, "Here are our statistical models, they're in a notation for a private, custom MCMC system you can't use" that's probably failing the standard even if the work is good.
Word2vec is a great example of a piece of research that conveys a great idea with plausible results, where in fact, the numbers are less useful than the story behind them. Of course, google is a big corp that can just put out useful interesting stuff without worrying about how to play the science game.
This is the actually important definition of "reproducible".
"has an install.sh" is really nice, but it's far more important that an informed reader can recreate the artifact for themselves from the written description.
The "must be able to apt install or it's not real science" is a particularly dangerous path to go down. It's how you end up with e.g. shit loads of well-engineered LISP or FORTRAN code with no actual scientific insights or knowledge transfer. Which has actually happened in the past. A lot.
Isn’t the whole point of a paper, to obviate needing to talk to the authors? Like, so that science is a ratchet that doesn’t slip backward the moment the authors die?
Related to most ML code, it's not written in C but in Python frameworks that implement whole layers as building blocks. So it's not as hard to read/change. You can take something complex like the transformer and understand it in a few minutes. The most puzzling part with framework code is figuring out the shape of tensors and, sometimes, what each dimension means.
Regarding reproducibility - it's hard to achieve on account of parallelisation. You see, neural nets use floats, and floats are not real numbers. Half the float numbers lie between [-1, 1] and the rest outside. If you do 1e-10+1e10 you get 1e10. So if you add 1000 times 1e-10 to 1e+10 you still get 1e+10, but if you add them up together first, the result is different. So float summation is not commutative. Depending on race conditions the order in which summation occurs could change and the result change as well, even if you use the same random seeds. And if we give up parallelism then we can't run the experiments any more.
That's if float rounding issues are even a big problem in the first place. If your results are within .1% over a few runs it's reproducible enough for most purposes.
I am one of those people who actually did the extra weeks/months to properly test/review/document/release my code and data sets (you can apt-get install my "research artifacts").
In retrospect, it was a poor use of my time and a poor use of my sponsoring institution's time. "apt-get install" is NOT what we mean by "reproducible" in science.
High-quality or easy to install code is not necessary for a result to be reproducible. True reproduction would mean coding the algorithm from scratch by following the written description, and that's what it literally means in most other fields of science.
You can't download & install a Large Hadron Collider in an afternoon. Does that mean the LHC experiments are "not reproducible"? Of course not.
But that's not even the important point. The really important point is that, in most cases, high-quality code is not even sufficient for a result to be reproducible! See: the article we're discussing.
IMO, the "blindly rerun the code" definition of reproduction is actually a HUGE barrier to creating a true culture of reproducability in computer science. It results in super lazy reviewing where "public source code that's easy to install and puts the correct-sounding shit into STDOUT" becomes a stand-in for "paper actually describes a novel idea in enough detail that it can be truly reproduced".
> So what is the optimal outcome here?
An optimal allocation of scientists' time and effort.
As a scientist who has actually done that leg work, I don't think packaging code so that it runs with a single click is the best use of public money in science in 99.9999% of cases. That time is much better spent on writing and other dissemination explaining the ideas that make the code work (in some cases well-documented source code is the best description but in other cases prose is much more effective and illuminating). Or on coming up with new ideas that are even better than the old ones.
Which I guess is just another way of saying that scientists should spend their time on science, not engineering.
P.S. When shitting on "scientists" for not being good enough software engineers, please remember who's going to be doing the actual work you're demanding. It's mostly phd students who make $30K/yr. And they have to do this work in their free time because their 60 hr/wk day job is fully allocated to doing the actual science. I.e., treat scientists who maintain their code as you would treat FOSS contributors who are making 5x-10x+ less than you while working longer hours. Because maintaining high-quality code is something they are almost certainly doing in their free time.
Which failed here. However, you're going to have a difficult time convincing many people that the attributes you described would be bad properties, just that they might not justify their expense.
> You can't download & install a Large Hadron Collider in an afternoon. Does that mean the LHC experiments are "not reproducible"? Of course not.
By the same token though, the LHC repeats experiments and solicits feedback on how to improve their methods, which they go to great lengths to publish and simulate, because they're aware of this problem.
> IMO, the "blindly rerun the code" definition of reproduction is actually a HUGE barrier to creating a true culture of reproducability in computer science. It results in super lazy reviewing where "public source code that's easy to install and puts the correct-sounding shit into STDOUT" becomes a stand-in for "paper actually describes a novel idea in enough detail that it can be truly reproduced".
Ah, yes. Yes. "If this code is TOO reproducible then people might reproduce it, and handwave handwave the quality of papers would decline.
That's certainly NOT the case in pure CS papers, which have only improved since the days when folks felt that "Lenses, Bananas and Barbed Wire" was how folks should go about writing papers.
Now, physics might be different. But there is surely a middle ground between, "I've shipped you a LHC just plug in in lol" and "This paper doesn't even remotely describe how we achived the results."
If you believe that wasting the time of scientists is bad, then surely you're for clear papers with accurate descriptions of the methods so that those who go and reproduce your work are not sent on wild goose chases?
> As a scientist who has actually done that leg work, I don't think packaging code so that it runs with a single click is the best use of public money in science in 99.9999% of cases.
No, we got that part. But surely someone does and maybe you can design your work to leverage that rather than reproducing and discarding scaffolding. My big concern here is that a lot of scientists (like you claim to be) are underqualified and unpracticed at software, and thus are surely seeing at least some aspect of their work distorted by software and hardware issues.
> Which I guess is just another way of saying that scientists should spend their time on science, not engineering.
Scientists are not going to be able to escape engineering. No one else is going to build what they need besides them.
> please remember who's going to be doing the actual work you're demanding. It's mostly phd students who make $30K/yr. And they have to do this work in their free time because their 60 hr/wk day job is fully allocated to doing the actual science.
Yeah, I'm aware. I suspect their lot would be better if your attitude wasn't that their work is disposable and unimportant.
How wrong should the original work be before it ceases to be a case of 'clean your own pipettes'?
Could you PLEASE consider reading the article we're all discussing before you roll into the comments section of an article about it with your strong-but-loosely-held opinions? You're arguing against a point that almost no one is putting forward (that software should be one click).
> I have to fix bugs before the code can run. I'm very curious how did the author run that code with the bugs.
Pro tip: at paper submission time, md5sum your code and also git tag it in your private repo. When you release the code, if you have to / want to release with a clean history, make submission-time your initial commit and then make the current state of the repo your second commit. I've never encountered an institution that won't allow that level of history in the code release, even places that are pretty hard core about wiping pre-release history.
The software types love to sht on researchers for their poor quality code. Yet a few threads over, there’s always a discussion about how more tests don’t make it onto the kanban board (or whatever you guys are using these days) when crunch time approaches. It’s not just us.
Sorry but I don't trust your research just based on yours words and cherry picked images.
Also research is incremental, so producing a proper code on which other can work on top of should be part of the CONTRIBUTION.
A very large number of research papers I've read in my life were so poorly written as to being near impossible to understand.
The time it took to infer the actual meaning of certain parts of certain papers was vastly larger than just reading the code would have been.
It has become very clear to me that being a strong CS researcher does not necessarily imply mastery of language or even the very basic ability to explain one's ideas. As a matter of fact, with very few exception (eg Feynmann), the two skills seem to be at odds with one another.
On the other hand, code can be read and understood much more readily, especially when it can be experimented with, if only to insert printfs in it to understand what it does.
The culture of scientific publication in CS must change. The "I'm pressured by my advisor and therefore have no time to publish my code" argument is complete bollocks.
The code should come first and your advisor should pressure you to publish that first (assuming it actually works).
If someone wants to read your poorly worded explanation of how it works, fine, publish a paper.
Show me the code first.
Being able to compile in some way and, possibly after some bug-fixing/compiler-pleasing yielding useful results that approach/relate to those in the paper, greatly increases my confidence in the researcher and his work. Pretty code is nice, but it's weighed separately, not affecting the appreciation for the original research work itself. Only the existence of code with an AGPL or less restricted license affects how I perceive the paper/original research.
That problem is compounded with the issues you describe, which I'd categorise as probably the best case. I did not find even a single case of a properly opensourced repository that just compiled somewhere that wasn't the researcher's laptop.
Imagine that you are a graduate student or a newly minted PHD who happens to be pretty good at practical software engineering: You can leave academia for FAANG and a solid six figure income, or you can stay struggling, poorly paid, unable to get tenure in academia. In some groups software engineers are treated as glorified typists, too valuable enabling other people's work to be a principle researcher themselves. I've heard from a number of people that they avoid hiding their programming skills to avoid that trap.
This has many consequences beyond the obvious direct ones like papers with unusable implementations.
Academic work is often valuable even when it isn't practically useful but in many cases where academics are /trying/ to do something practical they fail because their community lacks the engineering experience-- I've seen a fair number of papers presenting optimizations as useful engineering when in reality their approach only makes sense because of inefficient non-programmer tools make a weird set of primitives fast (e.g. using a matrix multiply in matlab where in C you'd simply write a loop). As a practitioner this is frustrating because it sometimes requires re-implementing the approach to realize that it was only 'fast' compared to a pants-on-head-foolish approach.
It also can create essentially fake results.
Many times the lack of software engineering is compensated for by mocking components out in ways that would give faithful results if the researcher understood everything, but the whole point of expirementation is that the researcher doesn't! This mocking is seldom disclosed in papers. For example, I've encountered multiple papers that claimed to implement some enhancement to Bitcoin and tested it in a test bed but in reality they just added sleep()s with appropriate times, guessed out of specious reasoning like "our function X should be 10x slower than a signature", and hardcoded constants (computed in mathmatica or sage) for their messages. ... not realizing that sleep() isn't the best analog for cpu busywork in a multithreaded program!
Another kind of fake result I've encountered which is even less directly a result of the software engineering shortage is in signal processing literature. While working on audio/video compression I found it common for algorithms to be presented without various constants and after reimplementing and asking the authors for their constants I found that they'd been cherrypicked for the ten images used in the paper, and that the whole approach doesn't actually work. This is a kind of ineptitude (or outright dishonesty) that would be much less common in a world where reviewers received a working and usable implementation in source form-- but that can't be expected in a world where qualified software engineering is not readily available to researchers.
I don't have any proposed solutions but I think it's important to acknowledge that it is a common and serious limitation to the usefulness and accuracy of contemporary research.
In the past academia and government agencies developed the internet and its communication protocols; this was (imho) a much better internet than the corporation-dominated internet we have today.
I'd like to see a resurge of this type of internet. Let us use well-researched open protocols again; and let us use applications which are really built for people, not advertisers. Let corporations build the hardware, but let us keep our data far away from them.
And besides the computer science / software engineering branches, I'd like to see other fields (like industrial design, and for instance even sociology) to join the development of a better digital future. The kind of future where everybody profits, not just shareholders of big companies.
This happened so many times when I was working on NLP/CV algorithms - I've read and implemented many algorithms from papers just to find out that they only produce the amazing improvements on cherrypicked dataset. On pretty much all other practical data the algorithms performed worse and in many cases even crashed!
To compound the problem, the fact that you're competing against many of your peers for the few jobs there are, and jumping positions every 1-3 years after grad school, means that best practices are not disseminated by working with a stable set of coworkers. This is similar to your 4th paragraph point.
As a mentor of grad students on projects that these days are increasingly using machine learning, I am trying to figure out how to add this stuff to my teaching load, essentially, because the pain of my students using nested for loops when they could just use a single vectorized expression is real... but I myself have no formal CS training......
They didn't even bother to merge my pull requests.