Making Git and Jupyter Notebooks play nice
timstaley.co.uk
timstaley.co.uk
Author here. A friend mentioned this was on the front page so I wanted to stop by and make explicit that this advice is OUT OF DATE as far as I'm concerned. It's way too much hassle (I work in a much larger team now than I did then!) and doesn't play well with rebase etc.
These days I either recommend the jupytext approach (not tried it but seems sensible) or personally I just use Sphinx-gallery.
Advantages:
* Plain python files play well with IDE refactoring, Black formatter, etc etc.
* You now have a readymade 'tutorial' page for your docs.
* Files are run with every docs build, so you can configure things to alert you when they're broken.
Disadvantages:
* You end up editing a throwaway notebook file. If you forget to copy-paste your edits back to the source, and rebuild, you have lost your edits. However, this forces me to keep the 'temporary, exploratory' nature at the front of my mind and not allow the notebook code to grow too large before performing some clean-up.
* database credentials.
* sensitive data
* a whole DataBricks webpage because the person didn’t understand how to export just the notebook.
* collections of notebooks named only what step in the process they are, and literally nothing about what they actually do.
* Whole base64 encoded images and zip files
* packages imported by manually manipulating system environment paths
* multi-processing/multithreading by shelling out and calling new python instances
* good old “don’t run these cells”
[0] https://code.visualstudio.com/docs/python/jupyter-support-py...
[1] https://code.visualstudio.com/docs/python/jupyter-support-py...
I believe a new design worth sacrificing backward compability.
In .rmd the notebook is just markdown where code segments inside triple backticks can be executed.
That sounds nice and git compatible (and it is), but you lose a lot of the convenience of Jupiter notebook, namely that output isn’t stored together with input.
The nice thing about Jupyter notebooks is you get story, code and results in a single package.
https://www.garrickadenbuie.com/blog/convert-r-markdown-rmd-...
Now you have:
- single file (Rmd, plain text) where you can do any edits, which, when compiled, produces:
- (1) you get story, code and results in a single package: one output file (usually, html), where all inputs and outputs are stored together (or, just the outputs, if you turn off echo'ing, e.g. for producing a report),
- (2) you get the code: another output file, which is a simple R script version of your Rmd file.
Code updates to the Rmd are viewable much easier on the output R script (#2), which is also much more convenient for debugging. For static text updates, I look at the Rmd file. For content updates (e.g. did my data change between the runs?), I look at the html (or Word or PDF etc) file (#1).
1. save code and output separately, so you can save the output (but not comparing the output versions), or always generate output from scratch when needed (this actually help to ensure reproducibility)
2. or you can use R notebook format, which save the result with document together, in some companion folders.
Plus you get the option of rendering an HTML which IMO is no different than having a jupyter file.
I set out to build ReviewNB[1], code review tool for Jupyter Notebooks. Turns out a lot of other people had this exact problem. One can see visual diffs & write review comments on notebook cell. Currently only works with GitHub though.
If you want to diff locally (before committing changes), you might like nbdime[2].
I was thinking about using RMarkdown files (vide https://towardsdatascience.com/version-control-with-jupyter-...) as for them diffs make sense.
Does any of you use this approach, or have insights on how to make it good for collaboration AND visible on GitHub?