Jupyter Notebooks as E2E Tests
rakhim.exotext.com
rakhim.exotext.com
Whenever I've required a report which intermingles code, text, and outputs, a simple bash-script processing my (literate programming) codebase has always done a far more impressive job of generating a useful report. And, without having to sacrifice proper packaging / code-organisation semantics just for the sake of a report.
I find it a big shame that the iodide / pyodide notebooks didn't take off in the same way, at least. Those seemed tightly integrated to html in a way that was very elegant.
(they're not completely gone though, it was nice to see this example on the front page earlier: https://news.ycombinator.com/item?id=42425489 )
- move all of the code out of the notebook and into a nearby python module where it can be linted/mypy-d/version-controlled/code-reviewed easily
- tell the notebook to import that module and make a small number of function calls to get whatever data you need and make plots. at this point the notebook is really nothing more than a REPL with inlined plots/graphics
2) Reports are one of many reasons people use notebooks.
3) You can work with notebooks in proper IDEs like PyCharm or VS Code or Emacs.
Having said that, I do agree with you that notebooks enable poor programming habits. (I have seen quite a few notebooks with > 10k lines of code, which is insanity.) Among other things, diffing notebooks is typically a big pain too. It is just that for many folks, notebooks' pros outweigh their cons.
This workflow works because my primary job is no developing software. It's developing solutions to solve multimodal industrial and science reporting issues, and as needs change, having a complicated stepping as just a series of code snippets is wonderful.
Sure, text is cleaner but I treat much of my work as "living documents" so notebooks do that wonderfully.
But they're fantastic for what they were designed for -- which is quite literally "notebooks". AFAIK the idea was first popularized by Mathematica, and I still reach for that when I have some highly iterative, undefined math/data problem to sketch on. IMO the real issue is that Python is used both for this purpose and for software development, which leads to people using notebooks inappropriately.
A lot of transformations can be very time consuming to run, and the ability to cache(for lack of a better word) the computation without having to write to disk, or utilize the python repl (which is very obnoxious to use for anything that extends past a single line of code) really speeds the process up.
An example would be: pull a large json from a server. Then break out a new code block to do all of your different manipulations on that json. Lets you prototype around with the object you pulled from the server without having to worry about, among other things - Dealing with how slow reading it from server is - writing it to disk and then reading it from disk to deal with how slow that is - the latency of parsing the file on each run of your script - the list goes on.
These arn't things that you want to worry about when your current questions are "what does the structure of this look like" and "what are some of the basic statistical properties of this data". Notebooks are like the python repl with the benefits of having a proper multi-line text input.
I've never even heard of people using notebooks for report generation, and honestly I'd agree that sounds like a complete nightmare.
As a person who's often tasked with building reports for presentation, and more often, for passing on to analysts to expand on, I find that notebooks give a much more accessible workflow than other options. Excel is way too limiting, most of the downstream analysts aren't software devs and don't have, know, or want to know an IDE. Notebooks give a way to combine graphs, text and code (show your work) in a concise form.
Jupyter notebooks in the article are used for "learning-oriented tutorial" or "goal-oriented howtos" docs. You see explanation, code, results, and it is easy to try it. There are solutions with a single click that can get you the running editable copy.
Misleading docs may be worse than no docs, so using notebooks as executable documentation is a plus. Though it is not the best format for e2e tests in general. Tests may be too complex for that. Software (proper code) is better at handling complexity in the general case.
Notebooks are literally for notes, for exploring new ideas - not creating production artifacts from existing ideas. You get a stateful kernel that can incrementally build up state instead of re-executing. And you get visual artifacts and user interfaces inline. The value proposition is faster iteration and immediate feedback, not report writing.
If someone needs here is an some sample code to run notebooks programically, and tune the output and formatting:
https://github.com/tradingstrategy-ai/trade-executor/blob/ma...
https://youtrack.jetbrains.com/issue/PY-71195/Remote-Develop...
As I write more code, I increasingly find the most important thing about tests early on is that they are easy to write and maintain. The help that, I find one of the best 'quick test suites' is "run program, save output, run 'git diff' to see if anything changes".
This has several advantages. If you have lots of small programs it's trivial to parallelise. It's easy to see what outputs have changed. It's very easy to write weird one-off special tests that need to do something a bit unusual.
Yes, eventually you will probably want some nicer test framework, but even then I often keep this framework around, as there will still often be a few tests that don't fit nicely in whatever fancy testing library I'm trying to use (for example, checking a program's front end produces correct error messages when given invalid input).
This should make integration with pytest etc, much simpler.
[1] https://nbconvert.readthedocs.io/en/latest/execute_api.html#...
ipytest: https://github.com/chmp/ipytest
nbval: https://github.com/computationalmodelling/nbval
papermill: https://github.com/nteract/papermill
awesome-jupyter > Testing: https://github.com/markusschanta/awesome-jupyter
"Is there a cell tag convention to skip execution?" https://discourse.jupyter.org/t/is-there-a-cell-tag-conventi...
Methods for importing notebooks like regular Python modules: ipynb, import_ipnb, nbimporter, nbdev
nbimporter README on why not to import notebooks like modules: https://github.com/grst/nbimporter#update-2019-06-i-do-not-r...
nbdev parses jupyter notebook cell comments to determine whether to export a module to a .py file: https://nbdev1.fast.ai/export.html
There's a Jupyter kernel for running ansible playbooks as notebooks with output.
dask-labextension does more of a retry-on-failure workflow
The concept of running code examples inside documentation as a part of tests is well known, and extending it to end-to-end tests / user guides is a good idea.
Next step might be to add hidden code cells with asserts, to check that the code not only runs, but creates the expected output.
As per the article, they have to run the notebook locally, commit all the outputs, then the CI checks if the notebook can run, and then renders the _committed_ output to the docs. This means that the verified code and output can be out of sync, partly defeating the purpose of the CI.
On top of their reactive goodness, Marimo notebooks are .py files, which makes it very suitable for this kind of (ab)use.
In particular, there is very little that a "notebook" style environment can get you that you couldn't have gotten as output from any previous testing regime. Styling test results used to be a thing people spent a fair amount of time on so that test reports were pretty and informative. Reports on aggregate results could show trends in how they have been executing in the past.
Now, I grant that this article is subtly different. In particular, the notebooks are an artifact that they are testing anyway. So, having more reliance on that technology may itself be a good end goal. I still have a hard time shaking that notebooks are being treated as a hammer looking for nails by so many teams.
BTW, sometime ago I wrote an article about surprising things that you can build with Jupyter notebook https://mljar.com/blog/how-to-use-jupyter-notebook/ You will find in the list: blog, book, website, web app, dashboards, REST API, even full packages :)
I believe you can achieve that if you use jupytext library, right?
[0]: https://stackoverflow.com/a/56813896/212538 [1]: https://jupytext.readthedocs.io/en/latest/