Papermill: Parameterizing, executing, and analyzing Jupyter Notebooks
github.com
github.com
Not one of my proudest moments, but it got the job done.
If they can call the notebook like a function, the second person's job becomes much easier.
As a stop-gap solution, for cases like a single presentation / proof-of-concept that doesn't need to live on and be reused -- it would work. Anything that doesn't match this description will accumulate technical debt very quickly.
Although it does seem like packaging dependencies and handling parameters are separate problems, so I'm not sure if papermill is to be blamed for the fact that most notebooks are not ready to be handled like a black box, even after they're parameter-ready. Something like jupyenv is needed also.
It's not very common for Jupyter magic to be added ad hoc by users, but it typically creates a huge dependency on the environment, so no jupyenv is going to help (eg. all the workload-manager related magic to launch jobs in Slurm / OpenPBS).
Kernels... well, they can do all sorts of things... beyond your wildest dreams and imagination. And, unlike magic, they are readily available for the end-user to mess with. And, of course, there are a bunch of pre-packaged ones, supplied by all sorts of vendors who want, in this way, to promote their tech. Say, stuff like running Jupyter over Kubernetes with Ceph volumes exposed to the notebook. There's no easy way of making this into a "module" / "black box" that can be combined with some other Python code. It needs a ton of infra code to support this, if it's meant to be somewhat stand-alone.
It encapsulates the kernel, which encapsulates pretty much everything for the notebook, right? I haven't worked with Slurm or OpenPBS, but I think if you let nix build the images that your tasks are running in then I think you're covered for pretty much everything except things that only exist at runtime like database connections. Not a perfect black box, but close.
Not even close to everything. In real world the environment of a notebook consists of a bunch of things provided by whoever set up the lab, i.e. storage and tools.
Typical examples include setting up Lustre or Ceph in a way that it will be accessible from a notebook (but that would also involve authentication, potentially).
And, in terms of tools: a workload manager that's perhaps integrated with Jupyter to schedule notebook execution on available nodes, but also to simply run workloads. Also, just a bunch of stuff written by this or another research group. Just this week I had to install and configure some tool for arterial spin labeling, but that would be the case with any kind of research: there's plenty of stuff that researchers will rely on on, that is central to their research, but isn't directly related to Jupyter.
By and large, Jupyter is just a front-end to any particular system where research happens, it's not the system itself.
As you say, from a programmer's point of view the logical thing to do is to convert the notebook to a Python module. But that's an extra step that may not be necessary in some cases.
FWIW I used papermill in my Master's thesis to analyze a whole bunch of calibration data from IMUs. This gave me a nicely readable document with the test report, conclusions etc. for each device pretty easily.
I was aghast to learn that this person had never written non-notebook based code.
Code notebooks are great as notebooks, but should in no way replace libraries and well structured Python projects. Papermill to me is a huge anti-pattern and a sign that your team is using notebooks wrong.
It's not about preference, it's objectively a terrible idea to build complex workflows with notebooks.
The "scoff" was in my head, the action that came out of my mouth was to help them understand how to create reusable Python modules to help them organize their code.
The answer is to help these teams build an understanding of how to properly translate their notebook work into re-useable packages. There is really no need for data scientists to follow terrible practices, and I've worked on plenty of teams that have successfully been able to onboard DS as functioning software engineers. You just need a process and a culture that notebooks cannot be the last stage of a project.
Notebooks do that, and even leave a trace while doing it. Table outputs, plots, etc.
It is not like a python backend that listens to events and handle them as they come, sometimes even in parallel.
For data flow, the code has an inherent direction.
Perhaps the largest critique against notebooks is that they don't enforce a linear execution of cells. Every data scientist I know has been bitten by this at least once (not realizing they're in a stale cell that should have been updated).
Sure you could solve this by automating the entire notebook ensuring top-down execution order but then why in the world are you using a notebook like this? There is no case I can think of where this would be remotely better than just pulling out the code into shared libraries.
I've worked on a wide range of data science teams in my career and by far the most productive ones are the ones that have large shared libraries and have a process in place for getting code out of notebooks and into a proper production pipeline.
Normally I'm the person defending notebooks since there's a growing number of people who outright don't want to see them used ever. But they do have their place, as notebooks. I can't believe I'm getting down voted for suggesting one shouldn't build complex workflows using notebooks.
I think it is more the way you express your general attitude and how you look down upon your colleague just because they have found notebooks perfectly suited for their work.
If you do so, you gain:
* the rigor of the SDLC
* reusability by other developers
* more flexible deployment
But you lose the ability for a non-programmer to make significant changes. Every change needs to go through the programmer now.
That is fine if the code is worth it, but not every bit of code is.
In my experience, most of the time the problem is in the input and interpretation of the data. Not fixable by a unit test.
Jupyter is a tool to do some exploratory interactive programming. Most notebooks I've seen in my life (probably thousands at this point) are worthless as complete programs. They are more akin to shell sessions, which, for the most part, I wouldn't care storing for later.
Of course, Jupyter notebooks aren't the same as shell sessions, and there's value in being able to re-run a notebook, but they are so bad at being programs, that there's a probably a number N in low two-digits, where if you expect to have to run a notebook more than N times, you are better off writing an actual program instead.
This is how all intellectual work proceeds. Most of the stuff you write is crap. After many iterations you produce one that is good enough for others. Should we take away the typewriter from the novel writers too, along with Jupyter notebooks from scientists, because most typed pages are crap?
Can you possibly make Jupyter notebook act like a module in a program? -- with a lot of effort and determination, yes. Should you be doing this, especially since the alternative is very accessible and produces far superior results? -- Of course no.
Using your metaphor, I'm not arguing for taking the typewriter away from the not-so-good writers. I'm arguing that maybe they can use a computer with a word processor, so that they don't waste so much paper.
Some people like challenge in their lives... and I don't blame them. For sport, I would also rewrite some silly programs in languages I never intend to use, or do some code-golfing etc. Literate programming belongs in this general area of making extra effort to accomplish something that would've been trivial to do in a much simpler way.
That's why I'm stuck in Tolstoi's War and Peace. You have to know French to get past the first few pages.
But, more to the point of literal programming: it's not the only tool that wants programmers to write some sort of a plan or a sketch of the code before writing code. A much more popular technique is TDD, which, again, wants programmers to write something informally first, and then formalize it later in code. And, as with literal programming, my experience was that it's not helpful to the point of being a distraction.
There's a good reason to think that some sort of a sketch or a blueprint might be useful for the future program. It works like that in many other disciplines. Artists would make sketches before painting the picture, engineers make blueprints etc.
I think that the reason why literal programming doesn't work is because unlike a sketch or a blueprint, one has to carry it on forever (and propagate back the changes, once they are discovered) as long as the code is being worked on. It probably would've worked better if it was some sort of a plan that can be abandoned at any point, something to give the development the initial push, but not requiring any further maintenance.
I always wish they would take a hint from Emacs org mode and make notebooks more useful for development.
It let's data scientists work in the environment they work best in, and it makes it easier to productionize work. If you seperate them, then there's a translation process to move the code into whatever the production format is which means extra testing, and extra development.
For example, if your notebook runs into a bug, you can just run all the cells and then examine the locals after it breaks. This is extremely common when working with data (e.g. "data is missing on date X for column Y... why?").
I think most of the "real" use cases for notebooks is data analysis of various kinds, which is why a lot of people dismiss them. I wrote a blog post about this a while ago: https://rachitsingh.com/collaborating-jupyter/
For notebooks in an ML pipeline, I find that data issues are usually where things fail. Being able to run code "up to" a certain cell and create plots is invaluable. Creating reports by creating a data frame and displaying it as a cell is also super-handy.
You say, "dial some logic in", which is begging the wrong question (in my experience, at least). The logic in ML is usually very strait forward. It's about the data coming into your process and how your models are interacting with it.
I’m actually a software developer with 10 years experience and also happen to do data science. And found myself in situations where I parametrized a notebook to run in production. So it’s not that I can’t turn it to plain python. The main reasons are
1. I prototype in a notebook. Translating to python code requires extra work. In this case there’s no extra dev involved, it’s just me. Still it’s extra work.
2. You can isolate the code out of the notebook and in theory you’ve just turned your notebook into plain py. You could even log every cell output to your standard logging system. But you loose context of every log. Some cells might output graphs. The notebook just gives you a fast and complete picture that might be tedious to put together otherwise.
3. The saved notebook also acts as versioning. In DS work you could end up with lots of parameters or small variations of the same thing. In the end what has little variations I put in plain python code. What’s more experimental and subject to change I put in the notebook. In certain cases it’s easier than going through commit logs.
4. I’ve never done this but a notebook is just json so in theory you could further process the output with prestodb or similar.
I am often in front of folks who "aren't computer programmers" but need to use Python tools to be successful. One of my covert goals is to teach SWE best practices inside of notebooks. It requires a little more typing but eases the use of notebooks, refactoring, testing, moving to scripts, and using tooling like Papermill.
https://github.com/marimo-team/marimo
marimo notebooks are stored as pure Python (executable as scripts, versionable with git), and they largely eliminate the hidden state problem that affects Jupyter notebooks -- delete a variable and it's automatically removed from program memory, run a cell and all other cells that use its variables are marked as stale.
marimo notebooks are also readily parametrized with CLI arguments, so you can do: python notebook.py -- -foo 1 -bar 2 ...
Disclosure: I'm a marimo developer.
- You cannot extract live variables (needed for testing)
- Cannot use pdb for debugging
- Cannot profile memory usage
You can do all of that with ploomber-engine (https://github.com/ploomber/ploomber-engine).
Disclaimer: I'm the author of this package
>Ploomber (YC W22) co-founder.
Papermill is great, but yes: lots of room to hack on it and make it better.
https://github.com/tradingstrategy-ai/trade-executor/blob/ma...
Hell no. I want to rewrite all that as a proper script or Python module.
also, if it isn't maintained by the company that made it, then it is a good sign that they are no longer using it. it suggests that there is a better solution elsewhere.
https://news.ycombinator.com/item?id=38659315
A lot of people seemed to move to nbdev which tried to improve on many aspects of notebook dev culture, but nbdev's maintainers include GitHub employees who are can only comprehend doing anything via GitHub integrations and get huffy and how-dare-you-take-that-tone-with-us if you ever ask about how you would go about doing something without relying on GitHub.
So personally I've just been focusing on nbconvert and rolling my own things.
If the notebooks themselves contain assertions to check that expectations on the outputs are met, then you have an automated way to check that the notebooks behave the way you want on some test inputs. For long notebooks, this is more like integration/functional tests rather than unit tests, but I think this is already an improvement over manually run notebooks.
Note sure about strict types: you mean running mypy on a notebook? Maybe this can be helpful:
- https://pypi.org/project/nb-mypy/
About linters, you can install `jupyterlab-lsp` and `python-lsp-ruff` together for instance.