The Jupyter+Git problem is now solved (2022)
nbdev.fast.ai
nbdev.fast.ai
- JupyterLab Git Extension[1] for local diffs (pre-commit diffs)
- nbdime[2] / nbdev[3] for resolving .ipynb git merge conflicts
- GitHub PR code reviews with ReviewNB[4]
- Alternatively, if you don't care about cell outputs then Jupytext[5] to sync .ipynb JSON to markdown
Disclaimer: I built ReviewNB. It's a completely bootstrapped business, 5 years in the making and now used by leading DS teams at Meta, AWS, NASA JPL, AirBnB, Lyft, Affirm, AMD, Microsoft & more[6] for Jupyter Notebook code reviews on GitHub / Bitbucket.
[1] https://github.com/jupyterlab/jupyterlab-git
[2] https://nbdime.readthedocs.io
Notice that using markdown is a possibility for jupytext, but not the only one. More interestingly, you can also store your notebooks as plain python files, whose comments are interpreted as the markdown cells of the notebook.
This is very useful, and not only for version control: if your notebooks are python files they can be executed easily in CI or by third parties just by launching the interpreter. No need even of the jupyterlab dependency.
With some care, you can craft a single python file "foo.py" that can be used at the same time as
1. an executable command-line program (that happens to be written in python)
2. an importable python module
3. a jupyter notebook (to open it you need the jupytext extension of jupyter)
4. the documentation with auto-generated figures, convertible to html or to pdf using "jupyter nbconvert --execute"
5. a regular .ipynb if for some reason you want to distribute the outputs in a re-executable format
For small simple projects, to showcase, describe and illustrate an independent algorithm, we have found this structure invaluable.
#Jupyter notebook and git
As much as Jupyter Notebooks have been a great tool for data science, the transition to deployment, and the general software engineering friendliness of Jupyter Notebooks could use some work. From time to time, I have explored how others have dealt with turning notebooks into an organized codebase and outputs. To date, I have not found a comfortable approach for me. The ideal approach for me would be to use something like 'node metadata' in the way of [Leo Editor](https://leo-editor.github.io/leo-editor/) to function as 'decorators' for a notebook cell for integration with git.
By this I mean using something like special markers in Python comments (since much of data science is done with Python) to map the content of a cell (or output) to a git repository. Better yet, define a special cell type for git metadata preceding a code cell. Then implement some basic git operations on the contents of a cell. Let's suppose we use @@git as a marker for metadata in comments for git. --- beginning of cell --- # @@git %upstream%=https://github.com/pyro-ppl/pyro # @@git %local%=~/repo/pyrodev # @@git %branch%=burnburnburn # @@git %file%=examples/cvae/util.py
# Here begins the contents of the util.py file ... --- end of cell ---
An extension would implement items in the menubar for various git operations: stage - stage the content as util.py file checkout - checkout from upstream, replace local copy, and refresh content of cell commit - commit stage file specified by %file% status - ...
Imagined workflow is that once a working idea scattered throughout a notebook has been sketched out, the user would mark the notebook cells that should be mapped to files in a git repository. Also this could be used in a mixed dev/data science environment where library code under development can be pulled right into a notebook.
Yes, there will be problems with committing code with comments that are specific to one user which is why a special cell type makes sense. Yes, there will be problems that I can't even imagine right now but ...
Please message me if you know of a cell-based git extension for Jupyter Notebooks.
Yes, even in a document model, merge conflicts can give you invalid documents. Programmers deal with this every day when they create invalid programs. Trying to hide the document into data complicates that in ways that are obvious in hindsight, and not that surprising with foresight.
And a host of other reasons
https://gist.github.com/softwaredoug/d527a18643f29832b0f41af...
I build all my projects, including a SQL lib, an EC2 interface, http and fastcgi frameworks, various web apps, and a mail merge system in notebooks, in nbdev and it’s made me much more productive.
Amongst folks that have used nbdev for 1+ years that I’ve spoken to, all report a 3+ multiple of improvement in productivity (based on non rigorous self assessment). This could however be biased because these people also report enjoying coding much more, so it’s possible some of that effect is the time just doesn’t seem as long.
Personally, after coding previously for over 20 years in various IDEs and editors (and for instance being prolific enough in vim I often gave talks about it) I wouldn’t ever want to go back to that old pre-notebooks time.
But it’s a huge investment and requires a lot of relearning - to get the most out of it you have to rethink just about everything. So it might be better for those earlier in their careers that haven’t got as many habits to change.
I like the ideas of literate programming, but I think there should be a way to do it independent of the "editor" being used.
I'm against the idea of doing all of your software development in notebooks.
There's a sane way to use notebooks. For me, nbdev is a step too far, as it really pushes notebooks as the primary dev artifact.
(All my opinions, people can do whatever they want)
While they make some great arguments about where notebooks can be really powerful for software development, I don't think the speaker makes entirely valid counterpoints to the original "I don't like notebooks" talk and most of the problems of notebooks still stand.
See this gist I posted above https://gist.github.com/softwaredoug/d527a18643f29832b0f41af...
My solution has been to switch over to Quarto notebooks (mentioned in the post with Jupytext), but I see the issue around saving cell outputs.
I'm curious why one would specifically want to save cell outputs as is in the Jupyter notebook, rather than archiving that in some other format. Sure, that might require putting a lot of information in one page (e.g., if that output is dependent on many other code cells and their outputs), but that just moves the linkage problem around - you'd have to have some way of indicating that the specific cell output was generated by a specific version of cell code (and the order in which they were run, sometimes multiple times).
I feel like Jupyter notebooks are the PDFs of data science. They are super useful for displaying results, but bake that data in a super inconvenient way for doing anything but rendering the data to look nice.
Though, I can see your point, I think. Why not include a build step that moves from your document to the generated output? My gut there is a large part of why the system got popular is that they worked hard on removing the friction that that would add.
As a comparison and to your point, I've seen people try to build "literate test suites" that were in a notebook, very happy with how the output looked. Only to find later that if they had used some of the more common test frameworks, those already create very nice reports. And moving the report format/creation out of the specification allowed a ton of flexibility.
What's wrong with rendering to HTML?
If someone starts playing with the inputs, they're going to lose the outputs you've created unless you have a saved rendered copy anyways
My blind guess is that it improves the readability of the notebook / promotes the literate programming mindset. But just a guess
This happens especially frequently if your team uses a lot of CIGARs (checked in generated artifacts)
In most cases writing a simple driver to automatically handle the conflict resolution is pretty straightforward (especially if the resolution is usually just to regenerate a generated file) and well worth the up front effort to eliminate ongoing conflict headaches.
https://git-scm.com/docs/gitattributes#_defining_a_custom_me...
The main hassle is that for security reasons all developers need to opt in by registering the merge driver , which you can put in a project bootstrap script if you have one. Would be great if GitHub (disclaimer: where I used to work) would integrate custom merge drivers in their internal conflict resolution flow.
We still support ipynb import/export, but using yaml for our internal representation of notebooks has made it hugely easier to do human-readable diffs and makes git operations way easier. (https://hex.tech/blog/github-sync/)
The problem is Jupyter notebooks+Git.
The solution is use Jupyter, but not Jupyter notebooks.
If I said "the problem is now solved, just use VS Code", most people would say "but that's not jupyter, i need the interactivity aspect", not realizing that VS Code is just as interactive (thank to jupyter).
Jupyter notebooks is an interactive notebook that implements that protocol.
Do you know what also implements the protocol? VS Code. And doesn't have any of the stupid problems that the Jupyter notebooks do.
In VS Code if you do as below:
#%%
print("my python print")
#%%
Then you can run it as a cell. Crucially however, note that those delimiters are ALSO valid python code. So you just version as with everything else.