Missing data hinder replication of artificial intelligence studies
sciencemag.org
sciencemag.org
https://blog.projectpiglet.com/2018/01/causality-in-cryptoma...
I provided code and the data (since I knew no one would believe me, without evidence). Why do we believe researchers who have more incentive than I do to publish? Their jobs require publication, or at least are greatly improved by it, thus they have every incentive to be less than honest about their "experiments". Most of the "A.I. revolution" isn't about algorithms, it's about having more data and the right data. That's why Google and Facebook can out compete in research, as much as they do. They have all the data.
Personally, I've almost made it a life mission to try and replicate every study I can. After watching people in labs (private and public biology labs), just straight up make up data - I never believe anything.
I also wrote some stuff on backtesting (which in industry is viewed as the gold standard, but IMO is deeply prone to flaws): https://blog.projectpiglet.com/2018/01/perils-of-backtesting...
Having worked in industry and worked in academia... I don't believe 90% of the research I hear.
(No affiliation)
Even registered the domain: http://researchdonor.com/
I know a few people with terminal illness who would donate large sums of money for reproducible research.
If its not from a public/standard/replicable data source my bullshit meter hops up a couple steps.
"We improved on the current SotA by 15% but you can't see the data or the code - but trust us, its for real"
This is why about 90% of the papers Google publishes are useless.
They describe deep learning successes, with datasets that aren't public, on hardware that isn't public, with software that isn't public.
They describe technologies, be it spanner, borg, flumejava or others, relying on closed hard- and software.
If an average university professor can't replicate research 1:1 based solely on existing, open technology and the paper, the paper should not be accepted by any journal. In fact, it shouldn't even be accepted on conferences, or be able to get a doi code. That's not worthy of being called science.
And it’s problematic when a large amount of AI research happens at Google, Facebook and co., and none of it can be replicated. Science requires replication, if AI research can’t be replicated, it’s not science.
thats just fine. they don't really need to publish anything, but if they do want to publish it, it should be done in a way other people can reproduce their findings.
The world isn't as black and white as you'd like it to be.
You've got to vote with your feet (or in this case, attention).
The correct standard should be, if you claim you wrote a program that does X, the program should be available along with simple instructions that explain how to run it. The reviewers should then follow the instructions and verify that it actually works, and returns the results described in the paper! This basic process would sadly invalidate most computer science research.
E.g. the paper is about some new method to approach X. In general, the paper could be valid even without any program implementation whatsoever; but you'll likely might supplement the method description with some experimental evaluation - but the paper is not about the experimental evaluation or the "experimental apparatus" i.e. the code they used. I mean, the paper is likely making some claims about the method as such, not about any single particular implementation of the method including their own. As a part of the main claim "this method seems good and interesting" they're providing some evidence "we tried, and in certain conditions described here applying method X was 10% better than method Y" - but the actual code used to do that is just supplementary material; the code is a tool they used in research, not the result of the research they published. The code is the "telescope" you used to make an observation, but the paper is about the observation, not about the telescope.
It's just as with medicine - we generally accept papers that say "well, we tried this procedure on 100 patients and 73 of them got better" without requiring video evidence of those 73 patients actually getting better; in the same manner, we accept CS papers that say "well, we tried this procedure on 100 datapoints and 73 of them got correct results" and don't require the reviewers to reproduce the experiments and verify if the experimental results aren't falsified; just as reviewers don't try to reproduce experiments in pretty much any other science.
> well, we tried this procedure on 100 patients
"this procedure" is the most important part, and they describe it in detail in the paper, hopefully well enough that someone could attempt to replicate it, and some do attempt to replicate it. (Not that procedure descriptions in such papers are always sufficient for this.)
The difference between that and
> we tried this procedure on 100 datapoints
Is that it's nigh impossible to describe a ML procedure in enough detail to reproduce with just the description in the paper. Tiny changes in the parameters and construction can completely change the result; the only way to be able to reproduce it is if you had the source code. And also the source data which is just if not more important as the source code (see sibling thread).
The opportunity that academic CS has over every other science is that they could empower every reader with the capability to verify the results of every paper they read, and this is actually attainable. Reviewers of other sciences don't reproduce findings themselves for purely practical reasons that don't need to exist in CS.
I though ML folks would have the statistical background to know you cannot infer a true statement from a single occurrence?
Research papers should NOT release code. Ever.
Researchers don't care, and should not care about your setup, your IDE, your OS, your dependencies, your processor architecture etc. None of that matters, unless it's the central point of the paper. You literally ask researchers to write coding tutorials or something. It's simply comedic.
Most machine learning and CS research revolves around Methods, Architectures, Processes and other abstract approaches to solving a problem. Their algorithms are always reproducible with any implementation you like.
Sure, the training data should be supplied, because in AI/learning it's part of the algorithm, but even if sometimes they aren't, you can try and reproduce the relative changes with another similar database. If their point is that one training method works better than another, the result still should be reproducible with similar datasets.
I get where you're coming from, after all most people on this site are programmers, but research is not about publishing code that works, it's about publishing ideas that work.
The entire field has a reproducibility problem, which implies that as a academic community we're just churning out pubs and not actually doing science.
I'm helping form Nature's foray into ML (https://www.nature.com/natmachintell/), and we're strongly considering enforcing a requirement that FOSS source code be published in a public venue (like github), and requiring that at least some benchmarks are established on publicly (but perhaps not freely) available data sets. This won't be appropriate for every paper of course, but our editorial staff is certainly going to be focused on maintaining scientific quality.
Well, that's the dream. In practice their algorithms are often not reproducible at all. Requiring code would get rid of the papers that cannot be made to work, where the authors stretched the truth because without code they could not be caught.
Most reviewers seem to agree that besides the top and bottom 10% of papers the primary differentiator is luck, so might as well add code to that mix.
I think that it should be about shaming. If you shame those that don't provide enough information to replicate you might create an incentive to publish properly.
Replication problems create so much noise for further research that I would classify a researcher who doesn't publish replicable results as hostile to the research community.
Seriously though, publish your data in the supplementary materials. Quit being afraid of failure or someone catching your mistakes. This is science, not figure skating.
The other advantage of a repository is that it gets included in large data collections, and is in a standardized format.
Often it's a lot of work to upload to repositories, nearly as much as the analysis (if you have a good analysis pipeline), but is well worth it both for oneself and for the community.
https://www.crossref.org/blog/the-research-nexus---better-re...
(edit - missing data citations from scholarly publications cause/propagate the replication problem described in the article)