After working on several large machine learning and data science projects over the years, I’ve sadly come to the conclusion that the notebook environment is a very wrong approach to take for shareable, extensible or reproduceable systems. On my team now, aside from the tiniest, most ephemeral prototypes that can be thrown away entirely, nobody ever uses Jupyter notebooks for anything.
The projects I’m referring to are very similar to those of this post too — an in-house image annotation tool we wrote with PyQT instead of notebook widgets, lots of interactive Flask web apps with simple React pages to demo and explore face detection results, diagnostic graphs, search engine output, and prototyping systems for layouts of image search results pages under different ML based ranking algorithms.
The main reason the notebook caused problems is that it leads to poor software craftsmanship, which should be a consideration even from the earliest stages of models and prototypes (because it will make you produce higher quality output faster, not for any spiritual commitment to coding standards).
When notebooks contain code you wouldn’t want merged into a given library or project’s main branch, it needs code review. And notebooks are awful units of work to review because they break models of automated testing, lack the same context for enforced style guidelines, and are generally written with a mix of priorities usually focused on the way the code author wants to think about the prototype, instead of separating these concerns into decoupled units.
Through review, you quickly realize that anything that needs code reuse has to be using good design principles, separating concerns, organizing things into decoupled functions or possibly classes.
When you factor these all out of the notebook and into a helper module (and write tests), then you see a bunch of the magic constants or initialized parameters that need to be factored out into a parameter file so that you can easily vary parameters, have it version controlled, and have a way to connect results artifacts (usually embedded plots or tables and saved data files) with parameter settings used to generate them.
Then you realize you need to make the overall software environment also reproducible, because a notebook itself is never “reproduceable” apart from the overall software environment it was run in, and other people picking up your notebook would need the same data and possibly a Docker container or other container or VM to match the same software, libraries, network settings, port number, etc.
And this process goes on until you realize you had to factor things into a tested module, automate the creation of the surrounding Docker container or other environment, extract out & document all your parameters, and so on. Until finally there is nothing left in the notebook but custom plotting code or data displays.
And the display units could be e.g. a version controlled latex file that picks up image files or renders table templates with data, or a separate web app that uses good design to implement reusable video display or something, all called from a boring non-notebook launch script, all of which could be driven by a simple custom Makefile — making the whole thing just as interactive, generally more so, as the notebook which suffered all the problems.
I fully agree this requires a team with good software craftsmanship skills. But anybody competent enough to write the original notebook can also learn these skills, and reorganize the work with just a few craftsmanship principles that almost immediately reveal the notebook as something that just gets in the way and slows you down.
And for teams needing to make reproducible data studies or reproducible model training environments for real business situations, this effectively makes a notebook inappropriate most of the time.
The main use cases where the notebook remains value-additive are situations where you can throw away the prototype at any time, it does not need any maintenance and it has no code or analysis within it that needs to be reused.
This does happen for some kinds of tutorials, some pedagogical uses, and some totally ad hoc work, like slinging some code to whip together a quick answer for someone on the business team. The notebook can still be good for these cases.
But these really represent a small number of use cases, certainly much smaller than the set of use cases the notebook is marketed towards.
My experience, after being a zealot for IPython notebooks in the early days, is that notebooks are just way oversold, and they encourage thinking in a low-craftsmanship manner, and are best avoided as a general rule of thumb.