A Visual Debugger for Jupyter
blog.jupyter.org
blog.jupyter.org
It looks to be a very welcome improvement for those of us who routinely use Jupyter notebooks for "REVL" (read-eval-visualize-loop) experimentation.
Has anyone here had a chance to test how well this works with objects from Python's scientific stack such as numpy arrays, pandas dataframes, matplotlib objects, PyTorch and TensorFlow CPU/GPU tensors, and so on?
I'm pretty sure that if your notebook is big enough to debug what's in it you'd be doing everyone a disservice keeping it contained like that.
This can convert a Jupyter notebook into a Python script but unless the notebook has been written with some care (with the intention of converting it into a script later), a lot of manual twiddling will be required.
What I generally do is use Jupyter notebooks when I'm exploring a problem or dataset. When I'm done with the exploration phase, I immediately mark the notebook as deprecated (trying to keep it in sync with code is a nightmare) and link the the relevant source code elsewhere in the project that it was deprecated for.
Databricks does a decent job productionizing their notebooks, and have seen people be quite productive on that platform.
Papermill and some light container tooling can take you a long way. It gets bonus points because data science folks are quicker to jump into a notebook and debug their own equations than they are a Python module.
This is what I would call a visual debugging experience: https://github.com/hediet/vscode-debug-visualizer/blob/maste...
(Disclaimer: I'm the author of that extension)
This new extension gives a graphical user interface that's similar to how debugging works in i.e. VS Code or PyCharm.
I guess that's visual about it.
I would be very happy if you could reach out on Twitter (@hediet_dev) ;)
What needs to be improved for Python use?
For TypeScript/JavaScript however, my extension injects some helpers that convert some common data structures (like number arrays) automatically to appropriate JSON. It requires a deep understanding of the target programming language to see how such helpers could be injected.
I don't have this unterstanding for Python and would be very happy if someone who really knows Python could help me out.
[1] https://github.com/hediet/vscode-debug-visualizer/blob/maste...
Thanks!
Thus, people in one camp see this as the standard name for the feature, and people in the other camp view it as misleading, and a continued diluting of the word in a way which trivializes their research.
After a decade of hearing "oh, visual programming, you mean like Visual C++?", I've learned to avoid the word entirely. It's a loaded term.
Every common word that is used as a popular brand name has this problem. I'm getting flashbacks to the 80's/90's and trying to explain my home computer to IBM PC people. "Do you have Windows?" "Well, it's an Apple. There are windows on the screen but it's not Microsoft Windows." "If you've got windows on the screen you can drag around with a mouse, that's Microsoft Windows."
As Jupyter has very powerful visualizations, I initially thought they somehow integrated their visualizations into the debugger, just to see that it's an ordinary debugger every modern IDE has. I know a debugger and it's UI is a crazy complicated beast, but I would expect any modern programming language to have such a debugger.
Just because there are text-based browser like lynx, modern browsers shouldn't start calling their products "visual browser", but I get your point.
Sorry for publicizing my extension here. It's free and open source, works with Python and might help a lot of people. According to github insights, traffic decreases whenever I don't publicize it somewhere.
I don't understand why VisualStudio doesn't have LightTable like expression playground yet, but until it does, Jupyter is what we have -- a fancy REPL and document publishing format being abused as an IDE.
That is only true in one very specific case: When you are executing all code cells in order, from top to bottom, exactly once.
If you start executing cells multiple times, or out of order, or even worse, execute only parts of cells (which is possibke in many Jupyter UIs), all bets are off. Anything can happen.
BTW, it looks pretty much like Spyder. AFAIR Spyder is a funded project by Jupyter. What happened between two projects?
Turns out that with JupyterLab, all you have to to is right-click and select "New Console for Notebook" and it opens a console pane below the notebook already attached to the notebook kernel. You can also instead do File > New > Console and select a kernel listed under "Use Kernel From Other Session".
The "New action runInConsole to allow line by line execution of cell content" "PR adds a notebook command `notebook:run-in-console`" but you have to add the associated keyboard shortcut to your config yourself; e.g. `Ctrl Shift Enter` or `Ctrl-G` that calls `notebook:run-in-console`. https://github.com/jupyterlab/jupyterlab/pull/4330
"In Jupyter Lab, execute editor code in Python console" describes how to add the associated keyboard shortcut to your config: https://stackoverflow.com/questions/38648286/in-jupyter-lab-...
Everyone should aim to minimize the amount of work they do in Jupyter Lab / Notebook.
It shocks me a bit to find myself saying that, as it is such a beautiful piece of work. Furthermore the people who wrote it are better software engineers than I'll ever be: the frontend, the zeromq-mediated communication with the kernel, the fact that the architecture has generalized so successfully to other language kernels, its huge popularity and reach. Nevertheless, I believe I'm serious. It really comes down to just two related issues, but they're extremely important: debugging and version control.
If you're a software engineer, and not a data scientist, here's how you probably debug already, or if not then how you should debug:
- You identify (a) commit(s) on which the behavior is correct, and (a) commit(s) where it is not correct.
- You experiment with fixes. Perhaps you stash them, perhaps you create experimental commits.
The critical point is that you use your version control system (probably Git) to navigate between alternative versions of the code. With a single command, you can switch the version of your code base, and the subsequent process you invoke to test your code is a fresh process, unpolluted by any state from the version of your code that you were on 30 seconds ago.
In contrast, Jupyter notebook does not encourage this style of work at all. In practice, what you will do when trying to debug some code in Jupyter is comment out lines, temporarily delete code, add experimental lines, add experimental new cells, etc. All creating a working tree, and a collection of in-memory python objects, that is a baffling mixture of changes related to the original feature development, and changes related to experimental debugging. Debugging will wear you out, as the state of your notebook gradually approaches complete incomprehensibility.
If you're a software engineer, you'll already know the benefits of being able to make precise adjustments to the state of your code with git commands. You want to learn statistics and data analysis skills from data scientists, but in doing so you should not regress to a worse style of development by starting to write much of your code in Jupyter notebooks.
And if you're a data scientist, you will want to acquire the debugging skills of software engineers. If you are not using git, you want to start learning it now.
Crudely, we can imagine a 2-dimensional diagram with one axis for engineering skills and another for data science skills. Everyone wants to be in the top-right quadrant. In that quadrant, version control is used, and the version control system is used for debugging. Debugging is rather important in developing all software, whether scientific/numerical or not.
So both groups should be minimizing the amount of code written in the Jupyter notebook UI: instead, write code in a standard Python package, in a virtualenv, installed in editable mode with `pip install -e`. If you need to use a notebook for graphical display, or HTML display of Pandas dataframes, or display of an audio playing widget, or any of the other amazing things it does so well then fine: use importlib.reload in your notebook to load and reload the bulk of your code from your Python package. The notebook should just feature calls to plotting routines etc that you have implemented in standard code files using your text editor/IDE. You could even aim for your notebook to contain so few lines of code that in some projects you might not even bother committing it.
Many of our PhD students learned programming using matlab and it's a mess, they never use version control, everything is in a single script, nobody properly debugs, etc.. I believe this is because matlab encourages that type of programming.
I decided very early on to use python instead of matlab, significantly before ipython notebooks became a thing. Because most of the python resources came from computer scientists, using software engineering methods was really "forced" onto me, and I really enjoy it now.
If I look at the people who have been converted to python using jupyter I see a very similar phenomenon that you talk about. People are creating a huge mess, in ways it's even worse than the monolithic matlab scripts, because the notebooks can't even be run as a whole unit, because working and not working cells are mixed, there are several cells that define the same function, but only one is the correct one (good luck remembering which one it is after a couple of months) ...
I thing jupyter is a great tool for teaching and exchanging and presenting analyses. But it is a terrible tool for programming and especially learning to program, because it really encourages bad practices.
Typically, I have a git repo with the final code products, some of the more complex code gets written in notebooks, then transferred to git and thoroughly tested. I've been dreaming of this debugging experience in jupyter because that's still not a task that's suitable for notebooks, but I am hoping that it will come for vanilla python kernels before I can hope to adopt it.
I agree with you that it's not a tool for writing software. It's probably best thought of as a really good REPL. And there are tons of uses cases for just that (at least in my discipline): analyzing an experiment, pulling data from a DB and plotting it, sharing boilerplate code, sharing analyses.
Sometimes I use it to test something out in isolation—something that I want to see functioning outside of the larger context of a production system—or to run a local version of an application, but that's not my primary use case.
You correctly observe that it's not a good fit for the tasks you have at hand, but I hope the above illustrates that there are lots of tasks that it's a great tool for.
I very often use jupyter to "pop open a shell intro a production service and start interactively debugging stuff live" (over an ssh tunnel, no public open ports and other security considerations, mind you).
It's amazing the feel you get the first time when you open a notebook that acts like a live REPL to smth. like a Django app and you start investigating and trying out stuff by just stitching snippets of code together in a notebook that imports your app and uses its db!
Now I code all API services regardless of tech they use so I can easily "pop open a jupyter REPL into a running system importing app code and running it agains its db".
Interactivity and REPL-driven-development-and-debugging is awesome if you have good discipline to contain the chaos and keep your notebooks aggressively short-lived (any useful code will be refactored and copied into its place in the regular codebase, most notebooks get deleted before merging a branch into dev/master).
So not to be argumentative, but to be clear about this discussion, I'm going to say that your comment is 70% irrelevant, since using an interactive REPL is routine in python development.
However, it is 30% relevant, because retrieving and archiving the code you ran is going to be much more convenient in a notebook than by using %history or whatever in a shell-based ipython.
notebooks are about non-linear execution
sure, there's enough rope to hang the whole neighborhood in that, you can totally f things up with no chance of recovery by doing that, or end up with data you have no idea how you got at and no way to re-trace 100% deterministically your steps
but if you have some discipline, stick to read-only-wrt-db, and you just delete whole notebook when things stop making any sense, it's... magical to have all that power and freedom at you finger tips, without having to keep much stuff in your working memory since you can dump it in a var or cell anytime, and mix your text notes through the code too!
It's not for everyone, but I love this beautiful chaos :)
All of which does seem to paint a picture of a programming environment which is handy for ad-hoc interventions and graphical/audio/video/HTML output but highly unsuitable for organized development of a code base (even a small one), highly unsuitable for systematic debugging, and highly unsuitable for beginners learning to program beyond their first baby steps (as @cycomanic points out elsewhere in this discussion).
The danger when that is done in an IDE setting is that because you build up commitment in a solution as you go, a reluctance sets in to rework it. So the final output is nothing like what you would write if you did it from scratch - it's littered with historical quirks of how you arrived at that implementation.
So I actually think that breaking out of the IDE and doing an exploratory phase in something like Jupyter is a really useful way to get your ideas into reasonable shape before you write your "real" code.
There were a couple of things you wrote that make me think that. Firstly: "without requiring the overhead of things like version control". There is a miniscule overhead to using git in the way I describe: `git init`, `git add`, `git commit`, `git reset` are the only commands you need and they take a second to invoke. Secondly "you build up commitment in a solution as you go": as i said above, I believe that as one gets more comfortable with git, it no longer feels a constraint -- quite the opposite, you feel liberated to experiment because you always know you can get back to any state you wish.
No amount of git or anything else an IDE can do achieves those things.
import matplotlib
matplotlib.use("Qt5Agg")
But yes, even I might start up a notebook to iteratively refine a plot! And the HTML table output for dataframes is perfect also. import mymodule
mymodule.myfunction()
That allows you to do this to reload your module and pick up the updates to myfunction that you've made: from importlib import reload
reload(mymodule)
However, how do you ensure that python can find your code so that `import mymodule` even works? Don't mess about with PYTHONPATH and sys.path. What you really want to do is house your work in its own python package. So, the milestones you want to get to are (not implying you don't already do these things!):- Always use a virtualenv when working with python
- Create proper package structure for your python project. This means your directory structure will look like this
myproject/myproject/__init__.py
myproject/myproject/mymodule.py
myproject/setup.py
- Google for how to create a minimal setup.py. Just put what you need in there, it's not much.- Now, with your virtualenv activated, so that `which pip` resolves to `myvirtualenv/bin/pip`, do this:
cd myproject
pip install -e .
- That pip command will execute your setup.py and "install" your library into the virtualenv. But it will install it in such a way that you can edit the code and the edits will be picked up by the installed version (it uses symlinks).- Now install jupyter in that same virtualenv and start your notebook. You should now be able to do `from myproject import mymodule` and `reload(mymodule)`. And your project is now a real python library so you can create subdirectories, etc e.g. `from myproject/plots import create_boxplot`.
https://github.com/jupyterlab/jupyterlab-git
https://github.com/elyra-ai/elyra#notebook-versioning-based-...
Add poor version control practices and the dawning realization that you don't know whether your code ever worked or just seemed to work because some variable was in scope that shouldn't have been, and it quickly becomes chaos.