Git and Jupyter Notebooks Guide
reviewnb.com
reviewnb.com
People: pip install jupytext. All your python files will become notebooks, and your notebooks will become python files.
The .ipynb stores inputs and outputs together in an unholy way. It is much cleaner to separate them. The inputs are python (or markdown) files that you can edit with a text editor and version control with git. The outputs are html, pdf, or whatever you want to nbconvert to and share.
The .ipynb file would only be useful if you want to share a stateful notebook, whose state cannot be easily reproduced by the people who you share it with. But that would be really bizarre and definitely in bad taste. Sharing the .ipynb is akin to sharing your .pyc files.
I love working with notebooks, but as a measure of hygiene I avoid .ipynb files altogether.
Another alternative if you want the outputs is to use nbconvert to convert the output to markdown, https://andrewpwheeler.com/2021/09/06/using-jupyter-notebook...
I’ve never used jupyter for taking notes in a lab setting, but with more and more instruments being computer/network connected, I imagine this would make total sense - put your data and notes with your analytics work.
Many jupyter users are not «software developers», they just use code to perform their work.
(How one does a diff of a data object look like? If there is a natural text format to save it in, it still is usually quite messy, and Git doesn't really like Gb sized csvs.)
My preferred workflow is to version the source files in Git and store the associated data objects in a separate archive directory with meaningful name and the hash of commit of generating code as metadata attribute.
Now if you had a version control "IDE" software that would render changes in figures and other blobs nicely, then it would make sense to build a workflow around it.
Jupyter notebooks store which Jupyter kernel they were run with to generate the outputs.
nbformat (.ipynb with inlined base64 outputs) isn't a sufficient package format: https://github.com/jupyter/enhancement-proposals/pull/103#is...
Papermill is one tool for running Jupyter notebooks as reports; with the date in the filename. https://papermill.readthedocs.io/en/latest/
I'm not sure whether you're unaware or just feigning ignorance, but notebooks are frequently used to share partial results, often in the context of "research", however you may interpret it. Imagine a grad student or data scientist preparing some code and plots to show during a weekly meeting.
In this context, the only thing that matters is quick progress and advancing understanding of a problem. The highly loaded words you employ while blasting the idea of uploading Jupyter notebooks are not relevant here. Wasting time on these things is seen as a bad thing. It's clear why someone using notebooks this way would want the interaction with Git and GitHub to be as seamless as possible: uploading something to GitHub is a very easy way to share it, even if this isn't the platonic ideal.
It will probably cause you some pain, but I've known people to commit binary objects and PDFs to git to accomplish the same ends. ;-)
As a concrete example, this one-liner of Python code is much more interesting to those who don't recognize it when it's presented with the associated output.
4*sum([(random.random()**2 + random.random()**2)**.5 < 1 for _ in range(10**7)])/10**7
This is also useful, e.g. when viewing the read-only export of a notebook.(The one-linear above is a monte-carlo simulation which approximates Pi. On one run, this result came to 3.1410416.)
Moreover, while research moves fast, reproducibility remains important. If your notebook is stateful, then when you share it I may not be able to recreate your result or you might have a bug due to something lingering in the notebook state. Having your outputs is convenient, but if I download the notebook, run it myself, and find that the code doesn’t run because there’s some variable that got defined earlier in your session but that code got deleted during iteration, that’s really not helpful. It’s the equivalent of handing someone your lab notebook but you kept erasing over early pages to make room for new content.
That’s one example of a bug. You could easily introduce more subtle bugs where the state leads to invalid results.
I'm a researcher and don't use notebooks for all the reasons you outlined and more. I have my own approach to dealing with reproducibility which is low tech and works for me and my collaborators.
My comment is meant to point out that there are many researchers who view all of the problems you describe as unimportant and not worth spending time on.
If it's specifically source code for anything that's intended to run, then avoiding including the outputs is a smart move. But then, if that's the case, there's a good chance you'd just be committing a .py file.
I like notebooks because they include output alongisde input. For example, Peter Norvig's Pytudes are all brilliant, quick notebooks that solve a particular puzzle[0]. The code itself might not be that interesting to run (unless you really want to confirm his strategy for wordle checks out) but reading through the notebooks makes for a great experience of simultaneously understanding his thought process, and seeing the solution.
I do a bunch of generative art stuff and have recently been experimenting with using notebooks as quick sketches[1]. I really like the workflow and end up with something like a journal that isn't necessarily intended to be ran repeatedly, but read over, where I can see the visual output created, as well as the method for it.
[0] Norvig's extremely cool pytudes, wordle example: https://github.com/norvig/pytudes/blob/main/ipynb/Wordle.ipy... [1] My not anywhere near as cool as Norvig's pytudes example: https://github.com/benrutter/jupyter-sketches
(Note that the jupytext paradigm does assume that the outputs can always be recomputed as a function of the inputs. I consider that a best practice, but some might disagree.)
I commit the Markdown-version, but I also use the py-version of notebooks for chained notebook imports. Allows me to split larger notebooks into multiple smaller ones. Both of these options are a blessing and Jupytext works super-robust.
Finally, when I want to archive (and share) notebooks _with_ outputs once in a while, I have a cell at the end to convert (nbconvert) to HTML, and I commit this html file. The Markdown-version remains as a clean basis for commit history. The HTML file is much better suited for sharing and archiving than the ipynb file.
I definitely do like to strip notebooks and make them run-idempotent to the best of my ability, but sometimes you just need stateful notebooks. And since .ipynb are technically json but in reality act more like a binary file format (with respect to diffing), DVC is the ideal tool to store them. Don't get me started on git annex or LFS, both of those took years off my life due to stress of using them and them bugging out.
Also I am hardly a fan of XML, but does anyone feel like notebook files would have been a near-ideal use-case of it? It's literally a collection of markup. The fact that json was chosen over xml I think is somewhat damning of xml as an application data storage format. I think xml is perfectly cromulent as a write-once-read-many presentation format or rendering target (html, svg, GeniCam api info), but it seems to flounder in virtually every other domain it's been shoehorned into, with the exception of office application formats.
Actually, downthread there is a link to a jupyer enhancement proposal for a .nb.md markdown based format. I think this is great. One theme I keep coming across in my computer science journey is that formats which have mandatory closing endcaps are kind of a PITA. It seems the stream-of-containers (with state machines as needed) is all-around better. JSON-LD is better than JSON, streaming video formats are better than ones that stick metadata at the end, zip is... an eldritch horror, etc.
jupyter nbconvert --clear-output --inplace my_notebook.ipynb
So you can use git as usual, like for code.I wish that non-emacs implementations of org were more commonplace, as it's a pretty sane markup language and supports embedded code and graphics, diffs nicely, and doesn't introduce the insanity of JSON.
- containing potentially sensitive data in your notebook
I realized that jupyter notebooks are a flawed idea when I've tried vs code. vs code uses jupyter-the-protocol (as opposed to jupyter-the-notebooks) in order to give you a notebook-like experience that doesn't involve the jupyter notebook file format. VS code's interactive files are valid python code.
To me that killed jupyter notebooks. Why use something that is strictly worse in every respect?
Nothing even comes close. There's a reason it's dominant in the data science field.
Your workflow works for you but the jupyter workflow works for millions of students, data scientists, and even developers. Heck I even know all the ways to avoid jupyter, and I still use it often, because it's so convenient.
> Nothing even comes close. There's a reason it's dominant in the data science field.
> Your workflow works for you but the jupyter workflow works for millions of students, data scientists, and even developers. Heck I even know all the ways to avoid jupyter, and I still use it often, because it's so convenient.
Copy pasting your comment here so when you eventually delete it people can still see the ignorance.
You have absolutely no clue what you're talking about. Worse, it seems like you didn't read what you're responding to.
Our preferred toolchain is based on make to build data science pipelines. Every step is scripted, and make ensures that upstream changes or script changes trigger downstream changes, ending with charting with gnuplot or similar. Our output charts all are not only timestamped but have a git commit id. And our source repositories contain a data manifest so we have commit IDs right into ETL stages into the DB.
End result is that in a couple of months, when the CxOb asks about some piece of work and pulls out a chart, we can trace the entire data pipeline used to create it, and reproduce it if required. That saves so much hassle!
That depends on how you use the notebooks.
With just a tiny bit of discipline, you can integrate notebook users into your sane workflow. For example, encourage people to restart the kernel and run all cells a few times per day (and definitely, before sharing anything). Meaningful output artifacts can be saved into files, that are later read by the notebook and displayed.
Then, when users are satisfied with their notebook, they save it as a python file thanks to jupytext, and commit it to git.
This workflow integrates well with your makefile setup: to reproduce the notebook and obtain its results you simply run it as a script. If you want a pdf or a static html that shows the notebook as-is, you can nbconvert it from your makefile.
For example, if your makefile has lines like these:
%.ipynb : %.py ; jupytext $< --to notebook
%.html : %.ipynb ; jupyter nbconvert --execute --to html $<
Then you run "make foo.html" and it will convert "foo.py" to "foo.ipynb", run all the cells, and produce a static visualization "foo.html". Since the intermediary notebook is not marked as a precious file, it is deleted automatically by make.Notice that you can simply run "python foo.py" as well, to produce the valuable output artifacts.
In the end, jupyter becomes just an editor of python files. A fancy editor, that allows interactive execution of pieces of code, which is great.
Why? I have no problem with reproducibility when I use a little bit of discipline.
Your workflow does indeed sound nice but also sounds like it involves way more tooling and institutional knowledge. Anywhere I can learn more about it or see the scripts you use?
What I need to ensure is that anyone picking up a piece of analysis 3 months later can reproduce exactly what was done. I've been burnt in the past by having to go back to the original analyst and be told "oh you run this bit of this notebook, then paste the results in over here, then run that". By insisting that everything is scripted and that there are no manual steps, we get a reproducible analytics pipeline.
The starting point for our methodology is the book "Guerilla Analytics" by Enda Ridge. It's worth reading.