Groundhog: Addressing the Threat That R Poses to Reproducible Research
datacolada.org
datacolada.org
This is just a bad way for the author to promote their own library for dealing with this. The way their library seems to approach this (using dates instead of versions) seems horrible too - on any given date I can have a random selection of packages in my environment, some of them up-to-date, some of them not. So unless all researchers start using the author's library (and update to the latest versions of everything just before they publish), it's only making things worse and not really solving the problem it claims to solve.
[1]: https://cran.r-project.org/web/packages/versions/index.html
I've not tried renv yet but packrat was a pretty poor solution.
I agree packrat (which I created) was a poor solution for most users. renv is far better and more usable.
I think in julia this problem is solved quite nicely with the Project.toml (list of package that you directly dependent) and Manifest.toml file (the version numbers of the complete dependency tree which is automatically generated).
It seems that in groundhog you declare only direct dependencies. Is there a way to store the full dependency tree in R ?
"Update January 6th, 2021 A reader alerted me to a bug with the current groundhog (version 1.1.0) where you cannot set the groundhog library to be a folder containing spaces in the name."
So we are talking about software here that somehow made it to version 1.1 *without anyone ever using a directory with spaces in it with it". This can be interpreted in two ways: either very few people have spaces in their paths, or very few people have actually ever even tried (not even really used, I'm only talking about the most basic trial use) this package. I'm not a betting man, but if I were, I know where I'd put my money...
YMMV of course.
I have a friend who taught herself R for her research and it was basically one big procedural codebase.
Sarcasm aside, I've worked with codebases like that- thousand-line java methods and classes and the like. The problem is that there's nothing that really forces modularity on a codebase. There isn't even any consensus, objective way to modularise code. Otherwise, a machine could do it and we wouldn't have this kind of problem. But, a machine cannot, and so we do.
First of all, I don't think people report this type of stuff because they don't know how to report it, and secondly think it doesn't need to support this use case anyway since space is a latecomer to naming and path game.
> The original quote is from Bjarne Stroustrup, the creator of C++
i find this ironic, given the 'popularity' (either way) of C++
it's been years since I've seen anyone doing that - a main reason, is that a very widely used dev tool, make, does not handle spaces in paths:
http://savannah.gnu.org/bugs/?712
thus leading to inertia in the whole ecosystem - if make does not support spaces in paths, why bother
This is extremely common, especially on Linux. Basically anything that uses things like Bash or CMake will almost certainly not work in directories containing spaces.
Developers don't use paths containing spaces because it causes so many issues with badly written Bash scripts, and as a result they don't test their code with paths containing spaces.
Bash and CMake and similar hacked together languages have very error-prone quoting rules that make it very easy to accidentally make something work with paths without spaces but fail on paths with spaces.
It is also a PITA to use when typing in a shell, as you need two characters ( \ + space ) instead of one. So even though my scripts can handle them, I still avoid them if possible.
Today I wanted to send a screenshot by mail.
Should be simple, but with not Gnome. I make the screenshot, Gnome creates a file "Screenshot from ...", but does not tell you where. Then I search it in the file explorer, find it, copy the path. Then I paste the path in the mail program, file:///....Screenshot%20from%20. Then the mail program: "File not found"
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=193163
I hit this when trying to test libgmp (as an example of an important library you would lose).
This means in practice you can't really build most software which uses configure scripts and libraries in a directory with a space -- this may well be what they are hitting.
Which in our world would scream 'complete amateur, avoid, avoid, avoid', but perhaps it's different in the R world.
Unfortunately, it’s that world we live in for pretty much everything.
Reproducibility? What if all of the source were to depend on part of a CPU instruction set that we stop using? How long must things be reproducible? We don’t even make lab equipment exactly like we used to with the experiments our current sciences are based on.
However, I give a thumbs up to Groundhog for trying to do the right thing.
And even if you might disagree for the single-threaded case, most things running in parallel will eat that free lunch of bit-identical results due to timing differences.
It's definitely a specialized language. It's not the go-to for managing servers or anything with a lot of I/O, but it has those capabilities because they're useful for managing projects. And I'd be hard-pressed to justify using a language for statistical analysis if it doesn't focus on statistical analysis. It'd be like rolling my own cryptography.
You need to differentiate between "base R" (everything that comes with a new install) and community-contributed packages. Base R is amazingly reliable. It has detailed documentation[0].
User-package land is more of a Wild West, that's true. I would personally not use anything that's not on CRAN unless I can walk up to the maintainer's desk (in non-pandemic times).
And CRAN... well... let's just say that people used to point to CPAN as a strength of Perl, too... All that sort of archives, after the first few years which comprise mostly of contributors with deep knowledge and who can produce high quality libraries, turn into dumping grounds for trivial half-assed 'libraries' under the guise of 'community contributions'. Example: try to do trivial compound interest simulations in R. So basic that it's barealy worth calling 'finance'. There are (at least) three packages on CRAN that claim to do this, except that (depending on which variable in the equation you want to solve for) they all provide only part of the solution, in mostly incompatible ways. And this is because very few of the people putting code into CRAN know how to... well... write good code. This is not an indictment of those people; many of them are much more intelligent than a bunch of us combined. It's just that for them coding is a byproduct, and with good intentions they share what has been useful for them, it just leads to a situation of 'in the land of the blind one eye is king'.
This is completely, 100%, absolutely wrong.
Of course you can. There's packages, with excellent software engineering structure, that are designed to include documentation and tests.
R has so much good software engineering, that clever people with no software engineering background can easily make their own packages!
And come on, the R language is a masterpiece. It's not cobbled together like JavaScript or bash. It's got impeccable functional programming language pedigree, you can even look at the AST directly of a function directly inside code.
I'm not sure how you came to any of your conclusions, other than not bothering to understand the language to start. It's a beautiful language with a messy, user contributed set of stats code.
For me, the problem with R is that the language is inconsistent. Many packages arose to address many problems, but they all feel like a hack on top of the core language. Take the whole Tidyverse; it just does dataframes from R core but then from the ground up. Now, users can choose between the core language dataframes and the Tidyverse dataframes. Same holds for plotting. The core issue, I think, is that the core language misses some essential features which other languages do have nowadays. For example, a type system. In R, since types are missing, everything is a table (dataframe) which I find just weird.
> It's not cobbled together like JavaScript or bash.
But also not as good as my favorite: Julia. Comparing it to Bash is like saying that its better than COBOL. We all know Bash is quite old, but for certain situations it just works.
As far as type systems, there's really two different types of "types": individual types objects that can have generic functions attached to them, etc. This is not as well known, and there are actually several object systems for typing:
http://adv-r.had.co.nz/OO-essentials.html
But these sort of objects are not quite as commonly created by programmers, because the second type of "types" are much more useful: data frames, which is kind of a vectorization of structs. This is what would be used in data oriented design, which is apparently much more common in modern game design.
Ie maven will create a folder structure like "/home/user/.m2/repository/com/example/example.jar" which will never have spaces unless the username has spaces (Can linux usernames have spaces?).
Spaces in filenames are a reality though, especially on Windows (where the home directory itself used to have spaces in it, and also where many home directories on corporate networks are on network drives and start with \\), and any software that can't deal with those kinds of paths has just not been exposed to much (if any) real world use. That was the point I was trying to make - software that can't handle anything but the most bog-standard path names in its core configuration is 'hey guys look at what I hacked up yesterday evening' quality at best. (yes yes it is possible to imagine exceptions, like software that is decades old and ported across platforms; I'm talking about something new that is meant to solve a general problem).
Microsoft MRAN https://mran.microsoft.com/
> For the purpose of reproducibility, MRAN hosts daily snapshots of the CRAN R packages and R releases as far back as Sept. 17, 2014.
MRAN doesn't seem to be very well known or used in the R community, but I don't really know why?
Separately, Nix https://nixos.org/ also solves this problem for lots of different languages, but is difficult to get started with and still a bit rough around the edges. Probably not a good recommendation for a typical analyst or academic at this point.
Nixpkg/Nixos is obviously a useful technology for reproducibility, but note that the output of Nix scripts can depend on the time the system was built, the contents of URLs and the system architecture unless care is taken.
In general, we want the output to depend on the system architecture and the contents of URLs. Nix uses hashes to require that URL contents don't change over time, which protects from those contents changing arbitrarily.
The main problem with reproducibility in science is that most scientists are not actually interested in doing science. Of course software will not fix this problem.
mran is a great idea and if Rstudio (the defacto gate-keepers of the faith -- with Hadley the high priest) pushed to use mran, then the R community would follow suit (like they do for everything else).
This would do a lot to bring MS into the fold, which would actually be great for R.
https://rstudio.github.io/packrat/
and sell their own package management product
There's certainly benefits to being able to pull down research source code, and bug checking it. That's how programmers check code: tests and audits.
However I think reproducing research is more often then not done "from scratch", taking a new sample, treating it, checking results. "independent verification".
Re-using source code saves time, but I would argue not being able to shouldn't threaten reproducibility.
However, if you can't even reproduce an analysis with the authors' own data and code, that's a red flag before you even get to the starting line. Ensuring that level of reproducibility is, I think, an essential ingredient to enabling the stronger form of reproducibility.
Personally, I made the mistake during my graduate career of trying to reimplement an analysis using a certain rather complicated ML algorithm, from scratch, in a different language than the original authors had used. After struggling mightily to get it to work, I finally bothered to try to get their own code working. (I had been hesitant to do so because I wasn't proficient in the language they used, and it wasn't even clear they had released all the necessary code, aside from the core algorithm.) Once I did that, I discovered that I couldn't even get their own code working on their own data, and gave up. This was researched published in Science by a group from a top-tier research university. (I don't fully blame the authors, it may well have been my own incompetence that was the issue. But it just serves as yet another illustration of how pervasive and disregarded the reproducibility issue was for a long while.)
> You give me your code and enough information for me to produce and identical environment or (even better) your code is insenstive the environment, then your research is Repeatable.
> If you describe your study sufficiently well that I can re-implement your study from scratch, without looking at your code and still get the same answer, then it is Reproducible
> If I can arrive at the same conclusions as you, just from a description of its aims, then it is Replicable.
More often than not it‘s not clear from a paper what exactly the authors did to a achieve a specific result. Being able to exactly reproduce what previous authors did should improve reproducibility; also for new samples.
[0] https://www.emeraldcloudlab.com/ [1] https://nextjournal.com/ [2] https://www.youtube.com/watch?v=L1UgdoP2aeg
The coding standards are often abysmally, unexpectedly terrible. Often not even the help of the original authors is enough to be able to produce the same figures from a paper because things and settings and commands get forgotten. Some part of the analysis was done in one language, another part in Excel. Some of the code has now disappeared. Some of the libraries are no longer working. Some people left and their academic storage space was wiped and therefore the intermediate steps and results or notes are deleted. You wouldn't believe it.
Once a paper is published researchers are not really incentivized to document things or maintain the materials. They got the publication, they put it on their CV. On to the next project! No time to waste on work that's already completed. New work leads to new publications, messing around with the old code for the sake of a potential later person interested in it is a waste from the point of view of a researcher, career wise. Also most papers are never attempted to be reproduced ever.
It's slowly changing though but many people are grinding their teeth, because they can't torture the data as much if things are out in the open.
I disagree. It is not perfect, but it is certainly a process that enables scientific development, as it has for centuries. If we start to create more and more rules that researchers need to follow, it will become even harder to make scientific research and most institutions won't have resources to continue.
The culture needs to improve. Benchmarks shouldn't be everything but reviewers are inexperienced. Many reasons in many parts of the system. In other fields the issues are different. They are more about having to obtain statistically significant results or you fail your career.
It's a good thing that people are waking up to this. It's not about punishing the individual researchers, it's about our collective intellectual immune system. We can't digest this firehouse of papers if it's poisoned to such an extent. It's not about charity or burden. It's about being skeptical when we know we're dealing with unreliable data. Science is a massive endeavor with massive quality differences between works and researchers and groups. Blind trust is no longer enough if you care about keeping your beliefs curated.
Also, it appears that Groundhog is itself a CRAN package and the author recommends installing with install.packages(). So is the author committing to never making any backwards incompatible updates to their new package?
Also, because groundhog isn't made for the author to use, whether or not the interface changes is irrelevant. You'll never encounter library(groundhog) in a paper.
It reconstructs how the fully updated version of everything worked that day which isn't necessarily the same as the researcher's environment. It's a horrible idea to use dates instead of package versions for this. The author's library doesn't solve the problem it claims to solve.
Well, yes, probably. It's not all that hard, and groundhog seems to have a fairly simple API anyways.
And groundhog still uses CRAN packages, it just brings a method of pinning them to a specific version.
How else would you install the cran packages without using install.packages? Unless if you want them to recursively install it using groundhog but that seems unnecessary.
As long as you have the timestamp it should work, though I assume there will be some edge case.
What you're saying is like don't use pip because you don't install it using pip? Or don't use package-lock.json because you can't install npm through npm?
OP is not claiming that Groundhog itself is a threat to the R language ecosystem itself, whereas the author is claiming that the R language is itself a threat to Science itself...
Just a reminder: https://news.ycombinator.com/newsguidelines.html
This approach enables stricter validations against tampering with the package repositories as a hash of the package can be stored in the lockfile, however it is obviously a bit more complex to use than the groundhog approach.
Not to mention the given example for irreproducibility in base R looks at code that would be a bug in the script for 3.6. It's only useful to keep this reproducible if I'm debugging the script.
And, in this case, anyone who's proficient with R would recognize this problem from personal experience or the many warnings in tutorials. I usually wouldn't shoot down a given example as though it disproved the existence of any example, but I don't know if there is another example. Unless old code relied on undocumented or contrary-to-documented behavior.
This seems like a non-issue given renv. And renv gives a more reproducible, I think, solution as it pins to versions, not dates.
If you want to guarantee reproducible results you have to use a container/image with libraries added at build time. Anytime you are relying on floating versions or downloaded libraries you will have issues.
If you can't even get those numbers, then you can suspect any number of things. Maybe you're not using the right data, maybe there was a typo, maybe someone fraudulently manually tweaked the numbers, maybe you forgot to do a step in the processing chain etc etc. There's no way to know what's going on if you can't even be sure how the original numbers were created.
Its not a coincidence that the author gives an example from the tidyverse ecosystem. Authors and users of tidyverse value other things like consistency and new features over API stability and backward compatility. The base-R ecosystem is actually very stable and so the original package manager is very simple.
With R spreading out from the academic environment and with many new authors breaking their packages' APIs we observe new attempts to solve the issues with dependencies (such as renv or https://rsuite.io)
- use Microsoft MRAN which did the heavy lifting of hosting archives
- use date instead of version
- install package automatically in first time (which pacman::p_load has been doing for ages) and easier to use in script level.
It's not coincidence that most package manager solutions used version instead of date to control the environment:
- A paper published on 2017 may used a date in 2017.10.01, but there is a high possibility that some of the dependency packages might be of earlier date, unless the author update packages every day/week, which is not a good habit anyway because updating too frequently will break things more frequently.
- Then how can you reproduce the environment using a date? The underlying assumption that all packages will be latest till that date simply doesn't hold.
That's why packrat/renv etc will use a lock file to record all package versions, and why you will need a project to manage libraries, because you will need to maintain different library environments and cannot install to same location.
Yet the author take installing all packages to a single location as a feature since you don't need to install same package again, and try to avoid project and prefer script as much as possible when doing reproducible research?
I think the major sales point here is:
> A nice feature of groundhog is that it makes 'retrofitting' existing code quite easy. If you come across a script that no longer works, you can change its library() statements for groundhog.library() ones, using as the groundhog.day the date the code was probably written (say when it was posted on the internet), and it may work again.
I don't know how good ratpack is now a days. I've never met an R application that uses it, but at my old work, we would take a dated snapshot of CRAN at the beginning of every new project. If we needed to update a package we could then "update CRAN" for that project. When productionising a project it would be frozen to a date in CRAN.
This isn't true.
Given this, it almost seems more dangerous to imply through this package that a particular date's results are reproducible, since unless the user has the same version of R, they may see different results anyway.
[0]: https://stat.ethz.ch/pipermail/r-announce/2020/000653.html
You can install specific package versions recorded in environment.yml file.
There are probably many ways to do this but this is an approach I like.
https://docs.anaconda.com/anaconda/user-guide/tasks/using-r-...
My favorite description of the language comes from http://arrgh.tim-smith.us/:
> R is a shockingly dreadful language for an exceptionally useful data analysis environment.
I feel like this is just one more data point to support that statement.
Uh oh, someone just discovered the modern programming landscape!
Python, Node, R, Rust, and other langs/OSes with package managers are at the mercy of volunteers who keep important packages healthy. Once issues stop being fixed, y'all better have local copies. This used to be predominantly an OS issue, now it is a language issue, too.
> Python, Node, R, Rust,
Correct me if I'm wrong, but for binary programs, a lock file easily mitigates these issue. I know Node and Rust both support lock files.
My concern is more about packages going stale and don't peer-match with other packages that evolve, or major versions that change results: not so much for R pkgs but there have been cases of major versions breaking existing projects, or requiring significant effort to update. (One example that zinged me is the FFI interface for Node. The "official" package hasn't been touched in years, and the "replacement", FFI-NAPI, is still has lots of open issues. We were using in-house fixes for some time.)
Reproducability is a big problem all around. When I create releases I put the binaries as well as the source in version control, because changes in tools/libraries etc mean that I probably won't be able to create the exact same binary several years later from the same source.
There is always a tradeoff between flexibility and simplicity. Clearly software needs to be able to change, or you are never going to be able to improve it or fix bugs. And an assembly of constantly changing parts is clearly going to come with its own challenges.
My own software product, Easy Data Transform (which competes with R to some extent) trades off some flexibility for simplicity by having a single set of binaries for each platform. You can't add any components (without hacking). So the same version of software should always give the same result.
Is this person suggesting we never improve anything? :)
However, if I run my software on HPC cluster, that’s no longer an option. The HPC at my university doesn’t allow running Docker, only Singularity containers(which isn’t supported on Mac).