How to implement an algorithm from a scientific paper
codecapsule.com
codecapsule.com
edit: Here's an example:
https://github.com/weecology/white-etal-2012-ecology
Run portions of the analysis pipeline:
Empirical analyses: python mete_sads.py ./data/ empir
Simulation analyses: python mete_sads.py ./data/ sims
Figures: python mete_sads.py ./data/ figsEspecially when pretty much all of these papers described results that meant the authors had already implemented their algorithms in executable form. But almost none of them made the code available in any form whatsoever.
Papers that I implemented or tried to implement in the past were often sorely lacking in details. Your point about unstated assumptions is absolutely true. Some papers are nothing more than glorified abstracts.
Free access to PUBLICLY funded research should be the default.
I agree. But some research is privately funded.
-A trivial bug in your program can go a long way to discrediting you if someone wants to. -If your methods get inlined into a popular library, the number of people who cite your work will drop to 0. Popular libraries (for Machine Learning at least) typically have a set of authors who will be cited, and if you implement some new method or improvement in that library, they will get the credit and you won't. Since being cited is the most common measure of your net worth, having your work be accessible only with your name attached is significantly better for your career than having your work be maximally accessible.
Releasing the code is another way that people can attack your conclusions, so maybe people are reluctant to do it?
"In the good old days physicists repeated each other's experiments, just to be sure. Today they stick to FORTRAN, so that they can share each other's programs, bugs included." -- EWD 498 ("How do we tell truths that might hurt?")
So yes, I agree, but there's going to need to be some education and cultural change before we get there, and I think fields that have embraced a computational subfield have a big head start.
Or worse yet, if you just run experiment.sh, see the results are the exact same, and declare that you recreated the results.
Too bad its adoption hasn't grown faster. Especially with all the recent focus on "big data", applied machine learning, and computational science papers. It's hard to actually measure the quality and contributions of most of them.
Not having an accepted method to cite and share datasets is also part of the problem.
1- http://www.computer.org/csdl/mags/cs/2009/01/mcs2009010005.p... [PDF]
If the research relies on a piece of "in-house" code, which isn't described, then that's a problem. But the OP is about research where the algorithm is the advance in the field. In a way, the algorithm is the result and not the method (even though the algorithm will probably be applied to a sample problem and the results thereof validated).
Now, if I develop a new method at the bench, you are expected to do the method at your own bench if you want to apply it to your research. This often starts as a pilot experiment just trying out the method, which can take weeks to months, before you integrate it fully into your methods. I don't have to provide a kit which consistently implements the method along with the paper. Certainly once the method gains some traction, an outside company may begin selling kits. Until then, it absolutely should be "the reader's job to ... reimplement the research" - that's an essential part of the process.
Similarly, an algorithmic method can be reproducible even if code isn't provided. Just like with bleeding edge wet-lab methods, if you want to use bleeding edge algorithms you will need to be able to code. You will need to be able to read a description of the algorithm and implement it accurately. The publicly funded result is the algorithm as an idea, and providing that the review process works, you're getting that idea. Later on, if the algorithm gains any traction, someone will implement it in a usable, robust package and more labs will be able to use it.
Finally, it takes significant time and experience to produce (and maintain) code that others can use and expect to work most of the time. That's a waste of time/money for a research lab, and an inefficient use of public funds. If you can't code yourself, leave code production to those who do it well and don't whine that you can't just plug-in whatever hot new method just came out without any effort or proper understanding.
I haven't seen people asking for maintained, (re)usable code. We just want the crappy code that was used to produce the results. There is even an appropriate license, the Community Research and Academic Programming License (CRAPL).
Trusting without the ability to verify goes against everything scientific.
If you think your code is too "crappy" for publication, why do you believe it is bug free enough to produce dependable answers?
Yes it does.
Very often, the data selected for publication is cherry-picked. Running the same crappy code on a more complete data set, (or alternatively, on a partial data set) would give a very quick indication of the robustness of the results - and unlike re-implementing, might be doable in a day rather than months of effort.
Furthermore, when you actually re-implement (if you do), it is extremely helpful to compare intermediate results, which is impossible unless you have the original everything.
> Re-implementing the algorithm they describe in the paper and getting the same result (or not) is far more interesting.
Yes, but very rarely done in fields that are not CS or EE (and not very common in these either). Usually, results are just taken as gospel.
Also, there is a ridiculous amount of negligence (and even fraud) in publications. just running the crappy code, seeing the results, and having a cursory look at the code and data would reveal a lot of that.
Hasn't this always been true about scientific papers? Descriptions can be verified by reproducing the experiment. Why is a paper any less trustworthy just because there's code involved?
It was always true to an extent.
Code is a force multiplier that makes it significantly harder to evaluate the paper with out it (and without reproducing an equivalent).
I'm not in academia myself, but I've heard from friends more than once that when they actually received code (and/or data) they requested from an author, the code turned out to be not precisely described in the paper, and the data is often massaged to fit in a way that's not precisely described either.
The question shouldn't be "why aren't you satisfied with what was good 20 years ago?", but rather "when sharing the bits that makes everything reproducible is a 'git push' away, why isn't it considered mandatory?"
It is a common error that science is about proving things; The scientific method is actually about trying to disprove things and failing to do so. If what you want to do is science, why don't you make it easiest to disprove your results?
People arguing against source code release often argue as if those of us in favor think that re-running the original simulation is the end-all, be-all of reproducibility. Clearly that is not the case. No one simulation can truly prove anything, and independent reverification will always have a place. But since we do have the source artifacts and original data, why not release them and show exactly what was done and how it was done? Again, the idea that experiments should not do so is merely an artifact of the fact that scientific papers could only be 10 very expensive pages or so in a journal; why carry unexamined assumptions based on that now outdated fact forward into the future?
Accidents of the past are nothing more than accidents of the past, not holy writ. And I'm not aware of a good argument against release of source code that doesn't boil down to well, that's just not how we do it when deeply examined.
If I can't reproduce the results because you're using methods that I can't reasonably find/replicate, then the paper should not be accepted. That's still true for code, and sometimes providing source code is the missing piece. That's definitely not what the OP is about.
This is missing the point. I can code, and I'm not promoting code sharing to more easily use other people's research. What I want is to be able to verify that the research I read is accurate, and there's simply not enough time available to reproduce everything I read from scratch. Reproducibility is important; I shouldn't have to just accept the authors' word that their results are exactly as described. Independent verification should be happening at the review stage at the very least, and authors should be required to make that possible if they want to publish.
A good example is VisTrails, a workflow application. If you use it to make an image, then you can click on that figure in a PDF, and the URL leads to an online record of what made the image, which downloads to your local machine and runs. You can pick up where the author left off (software, data, internet permitting). Running every program under such a workflow is cumbersome or impossible, though, but it's work in the right direction.
Yep, and we have no backup of science. If science is so important, why is there no backup of all these files? How important it must be that nobody is bothering- despite the moral implications- to make a complete backup? And why can't scientists access papers?
When deciding which paper to read, it can be a great hit where it was published. The link merely claims groundbreaking work was published in the best "journals". Especially in computer science, conferece papers are where recent, groundbreaking work is published and good conference are the ones that are hard to get into. However, I agree that groundbreaking work ofter gets a longer follow-up journal aritcle. But those usually appear years later and for those algorithms it is likely that there are existing implementations available by that time.
I have for some time been attempting to learn everything I can about automated theorem proving. A key paper is Robinson's 1965 paper "A Machine-Oriented Logic Based on the Resolution Principle". It uses quite different notations compared to, say, Kowalski's "Logic for Problem Solving". Robinson made an important contribution, but his work is like the assembly language of logic. Modern papers are much higher level. Kowalski's book is 1979, so its like C. My new domain map made these works much more comprehensible, especially as I switched between them.
The other good point is patents. That's why I'm reading papers from 1965 and books from 1979.
Is it only commercial applications that you have to worry about? I was under the impression that even a free implementation would be infringing the patent and make you liable for damages.
I have had University Council tell me I cannot open source code I have written implementing patented algorithms. It is unfortunately more common than most people think, and often not acknowledged in publications.
Data generation, input, manipulation, output and result checking are all very good.
Maybe things like Go can change that, or then some optionally typed language. There is no fundamental reason not to get massive improvements.
I don't know about CS, but in most scientific fields, this may be a bad sign. It can mean they're just trying to pump up their own reference counts, or it can mean they don't really know what other people are doing.
The only way to be sure that neither of these is the case is to know, a priori, that they're truly doing groundbreaking work.
1. Here is a paper about some interesting things we've been looking at, here is what we know, here are some ideas we are building on. These tend to be presented at conferences, as a "what do you guys think?" sort of introduction for the larger community. It is nice, because other researchers can then tell you if this is actually new, or just a repeat of an idea (that was hard to find because it used different terms and went nowhere), or something worth looking into, or the old "hey, here's some pitfalls I see from my expertise".
2. A couple more conference papers with results building on the ideas in the seminal paper in 1. These are just to keep awareness of your work to others in the field, get feedback, and to play the game right - you get much less credit if you don't have a history of showing you've been working on this for a while when someone comes along and "scoops" you.
3. Actually interesting/important work. This is the type of paper that the OP calls "groundbreaking". After all that work (see documentation over the years), here is something pretty awesome!
Another facet of this process:
When doing research, you have no idea what you are doing. I mean, you have expertise and goals and hypothesis/theory, but you don't know how it will pan out. You don't know if you suddenly find a spot where a left turn is required. So publishing about these new things is a good idea. Other people in the field can benefit from just that. Further, in my experience, each small result tends to spawn more questions/investigatory tracks, etc than it closes. So a lot of papers with honest "future work" sections are great places for grad students and others new to the field to dive in and get their feet wet. They can follow some of the tracks the original researchers just had no time for, and help fill in gaps.
Finally, the self reference is a good sign, because it establishes you aren't just some person coming out of left field with $BIG_IDEA (which looks a bit crack-pot-esque...)
"Groundbreaking papers...out of research teams in smaller universities that have been tackling the problem for about six to ten years. The later is easy to spot: they reference their own publications in the papers, showing that they have been on the problem for some time now, and that they base their new work on a proven record of publications."
Don't most bibliometrics exclude self-citations?
Some of the papers on the subject - the groundbreaking ones by folks like Goetz Graefe - were brilliant and very interesting reads, but at the same time were so involved I felt like I would need to dedicate years before even scratching the surface.
Walking away, I did learn to see the difference between good papers and bad, and learned a heck of a lot on DB internals (good through the end of the 90's at least). But I think I'll stick with books from now on :)
PS - Speaking from CS perspective.
Problem is, this kind of obfuscation is really difficult to spot unless you actually understand the paper and see all the steps necessary to implement the algorithm.
My research is in molecular dynamics. I've only been a grad student for one semester, but I've written a lot of code. The code I am currently working on takes a force-field description, combines it with a listing of atom coordinates, and completely automates the production of an input file to a simulation program. (This is less trivial than it sounds. All stretching, bending, torsional, and improper atom connections must be generated. Bond-orders must be determined, solely from the structure. And then this is combined with equations from the paper describing the force-field to give an energetic potential for each connection type).
I would think this code could have a lot of value to other researchers. Though I doubt anyone in non-CS departments have even heard of Github.
No. A Java class would certainly have worse performance than just float or double, since each instance would be individually heap allocated, have the usual per-Object overhead. Better use text preprocessing (or just "double") than that.
(otherwise, this is all fine advice)
> 6.4 – Avoid mathematical notations in your variable names Let’s say that some quantity in the algorithm is a matrix denoted A. Later, the algorithm requires the gradient of the matrix over the two dimensions, denoted dA = (dA/dx, dA/dy). Then the name of the the variables should not be “dA_dx” and “dA_dy”, but “gradient_x” and “gradient_y”. Similarly, if an equation system requires a convergence test, then the variables should not be “prev_dA_dx” and “dA_dx”, but “error_previous” and “error_current”. Always name things for what physical quantity they represent, not whatever letter notation the authors of the paper used (e.g. “gradient_x” and not “dA_dx”), and always express the more specific to the less specific from left to right (e.g. “gradient_x” and not “x_gradient”).
Especially when you're just starting out, creating your own naming scheme just creates more opportunities to do something wrong.
An equation and an algorithm might achieve the same result but the ways they get there are so different that using different notation styles makes perfect sense. For example, a simple finite summation is a compact block in mathematical notation but it's a multi-line for loop in C. Trying to force the constraints of the 'source' notation on the implementation makes no sense.
Remember, dA/dx means 'gradient'. You're not creating your own naming scheme, you're translating the concept of 'gradient' to the appropriate notation for the medium you're working in.
If the Linux compose key supported more of the common mathematical symbols, I would have so much trouble not using them in all of my JS code. It's already hard not to use names like â to denote unit vectors.
As a side note, I really want to make a JS library called "Eta" for creating progress bars (puns!), where the global namespace is under Η (the Greek letter), but I think that might piss people off, even if I did allow the visually identical H as an alias.
Google '.xcompose github' (without quotes) and see what you come up with. I use this one, myself:
https://github.com/kragen/xcompose
but there are a lot of others.
In case you happen to use vim, you can easily type Unicode characters by using digraphs (Ctrl-K and then two characters):
<c-k>a* → α
<c-k>b* → β
<c-k>g* → γ
<c-k>OK → ✓
<c-k>NB → ∇
<c-k>RT → √
<c-k>(- → ∈
<c-k>-> → →You say "F" for forces, "p" for probabilities, "x" for machine learning inputs, "y" for classes, etc. In some cases, you might want to use English terms for these things, but usually you'd want to use the standard letters. Especially if the source paper that you're following uses some flavor of standard notation. (So many standard notations to chose from!)
Of course, coming up with universal rules for naming things is impossible. The advice would probably have been better left out.
Figure out if the pseudocode notation in the paper is using 0- or 1-based array indexing.
If the paper doesn't match your implementation language, consider doing the initial implementation using an array adapter class. The adapters can be removed later if they reduce performance, but they will likely have save you a number of maddening errors in the meantime.