JupyterLab 4.0
blog.jupyter.org
blog.jupyter.org
- I get an analysis that I like, but there isn't a good way to share it with others, so I end up just taking screenshots.
- There isn't a good way to take the same analysis and plug new data into it, other than to copy-paste the entire notebook.
- The process to "promote" fragments of a notebook into being reusable functions seemed very high-friction: basically you're rewriting it as a normal Python package and then adding that to Jupyter's environment.
- There aren't good boundaries between Jupyter's own Python environment, and that of your notebooks— if you have a dependency which conflicts with one of Jupyter's dependencies, then good luck.
Some or all of these may be wrong or out of date— Jupyter definitely passes the "oooh nifty" smell test, and I'd love to figure out how to make it usable for these longer-term workflows rather than just banging out one-off toys.
Don't get in that situation to begin with. Pop an `%load_ext` `%autoreload 2` at the top and just write functions in an imported .py file from the get go.
I would take a notebook that demonstrates what's happening to my data any day over a docstring that may or may not be correct. Particularly if I have to render it in yet a third system to juggle and manually keep in sync.
So we've gone from notebooks to... notebook + py + docstring... I'm sure we can think of another useless layer of indirection to bolt on.
I'm not sure you know what a docstring is if this is your response.
I really feel like you do not understand what a docstring is. This does not make any sense.
> the way you could if you broke the function up into cells
If your logic is broken up across cells then you cannot use it anywhere but in that notebook...
> If your logic is broken up across cells then you cannot use it anywhere but in that notebook...
You don't seem to understand how notebooks can be used and processed themselves. They are just data and there are libraries for loading and transforming them. You might check out pydev and papermill to get an idea about this (and maybe also literate programming in general).
Using notebooks as functions is a bad idea. I think papermill is trash.
Notebooks are good for illustrating what you did and what you got. They are not good for illustrating how you did it.
I’m not sure you should be able to just pop in and do “data science”
What would be your preferable way of sharing your analysis with others?
- You can turn jupyter notebooks into pdfs directly in the jupyter UI.
- You can upload them to Gitlab/Github and share the link to the rendered result.
- You can upload them to Colab/Binder/Kaggle and let people play with the code themselves.
- You can turn jupyter notebooks into beautiful websites/documents with: https://quarto.org/docs/tools/jupyter-lab.html
- You can add jupyter notebooks to you docs using nbsphinx: https://nbsphinx.readthedocs.io/en/0.9.1/
- You can turn jupyter notebooks into interactive web apps with voila: https://github.com/voila-dashboards/voila
- You can turn jupyter notebooks into presentations with rise: https://github.com/voila-dashboards/voila
But it sounds from both your comment and the many sibling replies that there are a number of tools now directly addressing this space, so I should definitely re-evaluate what is available.
I don't know if I understood it correctly but maybe you could:
- Upload your notebook to Github, then create a url with Binder (part of the jupyter ecosystem) directly to an editing/fiddling playground: https://mybinder.org/
- If by user-local you mean on their own machine, they can clone your repo and run their own jupyterlab to fiddle
- If everything should stay on your own computer/server, you could share a link to your own jupyterlab and collaborate with others in real-time: https://jupyterlab-realtime-collaboration.readthedocs.io/en/... (doing this securely might be a bit of a hassle)
- Plotly can be used directly as pandas backend: https://plotly.com/python/pandas-backend/
- The plotly.express module makes it easy to create interactive plots in html/js format from pandas dataframes: https://plotly.com/python/plotly-express/#gallery
It's cumbersome, and I'm not totally sure it's the correct way, but I remember getting around this by creating a virtualenv for my projects and then using that virtualenv's python as Jupyter's "kernel".
I bet a lot of people end up installing jupyter into each virtual environment instead.
This sort of works, as (I think) most people are only working on one or two notebooks at a time, and aren't using notebooks that relate to more than one virtual environment.
[0] https://stackoverflow.com/questions/39604271/conda-environme....
Try micromamba, it will shave years off dependency resolution
I’m to the point now where if anything other than venv/pip is required I won’t use it. Unfortunately there are many things that insist on conda.
Conda isn’t perfect but takes on a lot of problems that pip doesn’t deal with at all. Regular conda is really slow these days but you can use mamba instead or just configure conda to use the libmamba solver and it’s much nicer.
The folks at prefix.dev seem to be building some pretty cool drop in replacements for conda too.
Most shells cache executable paths, so the path for jupyter will be the global path, not the one for your virtual environment. This is unfortunately not at all obvious and leads to very hard to track down bugs that seem to disappear and reappear if you aren't familiar with the issue.
I have a recipe here which always works: https://github.com/nlothian/m1_huggingface_diffusers_demo#se...
If you don't have requirements.txt then do this: `pip3 install jupyter` for that line, then `deactivate` and `source ./venv/bin/activate`.
I primarily use Jupyter for prototyping: trying ideas, plotting results and sharing notebooks for others to improve on (or punch holes in). Once a piece clearly shows promise I move it into a module. While this usually means a significant rewrite I actually see it as a benefit: I can be messy in original prototyping plus a rewrite after experimentation often leads to a better code with a small time investment.
Beyond the minimal self-discipline of actually moving code into modules I still have two three-character friction points: vim and git. Pointers on addressing those appreciated!
I would love, love, love a tight vim and Jupyter integration to be able to switch, easily and frequently, between editing a set of cells in vim and in Jupyter with solid sync between them. I am perfectly OK with vim ignoring the output. And I would love to have a git mode that only checks in the changes that cause actual differences in the python code; not timestamps or the output.
EDIT: actually, you can do it directly in neovim - https://quarto.org/docs/tools/neovim.html
1. Only checking in semantic differences (not output, timestamps, etc.):
Use the `jupytext`-extension [0] to seamlessly pair you notebook with an lightweight markup version of the input only (which can be used to generate full notebook).
2. Being able to switch between text editor and notebook interface
As other have mentioned, there are integrations for multiple editors.
Another approach is to move the central code out to a python module which you edit in a text editor, and then use the `%autoreload` magic [1] to reimport that module whenever you execute a cell in the notebook.
[0]: https://jupytext.readthedocs.io/en/latest/install.html [1]: https://ipython.readthedocs.io/en/stable/config/extensions/a...
Enclose scripts in functions within the notebook, which minimizes the clutter of hidden state. I also have a habit of not walking away from a notebook without doing a "restart kernel and run all cells" to make sure the notebook works. I'm not dealing with giant data sets, so this doesn't cost me much.
Frequently used functions go into .py files, using auto-reload to keep things synchronized while I'm working on them.
Mature .py files that I might want to re-use in different projects get turned into pip-installable packages. The notebooks become informal tests of the packages.
I've never used venv, and never encountered dependency version problems. Some of the dependency horror stories may be obsolete due to the maturation of the big packages such as numpy and matplotlib.
https://github.com/glacambre/firenvim
Haven't tested it in combination with Jupyter but I imagine it should work
Once I figure things out of course, either I'm done with the notebook and can copy a few plots out for inclusion in a powerpoint, or I'm done with the notebook and extract a pile of functions into a utility script or package.
Either way, it has served its purpose for rapid prototyping and one-off analysis.
jupyter notebook as IDE, but with .py files instead of .ipynb
The notebook .py files are just regular python files with comments that can be edited at hand wih any text editor. Thus you can easily collaborate with your local graybeards that will dislike editing text on their web browsers.
I'm not sure why you think it's viable to collaborate with others who don't use jupyter? I mean on any software project, if we don't agree how to compile or run a project, then we usually can't collaborate IMO. Changes become nonsensical (breaking the one or other mode that is not tested by the author of the change.)
I could not find a way how to check-in .ipynb file and let reviewer review code only, and ignore the metadata part
if jupytext will allow me to work in jupyter, check in file to git and create MR and let other people easily see, review and comment my code that would be great
Jupyter supports printing out as a PDF, an all-in-one-webpage and all sorts of other formats. I've used this and the support for Markdown to create a number of documents that I've presented up the leadership chain, or to other engineers across the org "Here's the data, here's the narrative, and here's the code that shows you how I got to it so you can reproduce". Can even share the Jupyter notebook itself.
Is the issue that you do not want to save the data and report into a folder and distribute that? That is, you want an entirely self-contained notebook? Or is there something else going on here?
I'm sure that's possible but it seems kind of wrong to put your binary "data" in with your analysis and presentation-making code.
The testing I do is annual equipment performance evaluation. I'd like to be able to process each test and then feed the results into longitudinal monitoring. One thing I am considering is adding a library or extension to papermill that automatically creates a workspace hdf5 or dill or whatever that I can store individual variables into. After studying the ipynb JSON it just seems odd that you can't just store blobs as attachments. But what I understand is it has to do with the kernels and notebooks running as separate processes and passing things around as notifications. So basically the kernels don't have any access to the cells or any sorts of introspection.
With papermill you have parameters, there's just not any "return values" in the processed notebooks. If it existed you could treat "reports" more easily as cached function evaluations.
https://stackoverflow.com/questions/34342155/how-to-pickle-o...
(I just got it by googling “pickle an object in jupyter,” so sorry if this is something obvious that you’ve already seen and doesn’t quite solve your problem).
Only way I can think to make the html file truly portable while fully functional would be to embed a python interpreter and all required libraries as wasm.
Could be possible with pyodide? I haven't used it.
Papermill does a great job of recording the input parameters and execution history. But there isn't anything equivalent to the "return foo, bar" part of a function which makes it difficult to build up modules. You don't want to have to digitize a plot to carry on to the next step is what I'm saying.
Sharing: there hasn't been a good solution, so you have to hack something together. For us (small team) symlinking a "shared_notebooks" network path works fine (but looking forward to the RTC mentioned in the post). It's very important to restart your notebook before sharing.
New data: I've never really run into this in a form that "restart + run all cells" didn't really work. Are you trying to keep older versions?
Reusability: I highly recommend installing a personal Python package in development mode to your kernel (i.e. pip install -e .). Then just move functions from the notebook to the package and reload.
Boundaries: As others have said, I would just make a "user" kernel right away and never touch the Jupyter environment (as TLJH does by default).
I am kind of hopeful that Armin Ronacher's rye fixes Python environment hell and we end up with better solutions for many of these things, but there are definitely some issues with version management in Jupyter.
Typically a data engineer will be the one who helps bridge that gap, but that has a problem of does data engineer's output == scientist output, which can be time consuming to handle.
To shrink the gap from dev->prod - we have 2 notebooks, one for model development and one model deployment in production. We use papermill[0] to execute directly notebooks in production.
Shared functions between the dev/prod that are built by the scientist are put into a separate notebook and then imported via `run`. If I'm honest, our scientist don't do this and simply copy/paste the functions if we don't yell at them to fix it.
This basically allows us to stay within the jupyter environment entirely so that dependancies are isolated.
So, it's far from perfect, but it's allowed us to shrink the dev->prod life cycle time. Love to hear what others have done towards the same end.
nb_conda_kernels is pretty reliable but not actively maintained. Gator from the mamba folks is new and still a bit rough around the edges but looks like it will be pretty slick eventually.
To plug in new data, I experiment with papermill.
There are solutions that can put jupyter and the kernel in different environments. It's even not too hard to setup, but it astounds me it's not a common way to set it up. It's natural to have kernel + the analysis's deps in its own environment.
> - There aren't good boundaries between Jupyter's own Python environment, and that of your notebooks— if you have a dependency which conflicts with one of Jupyter's dependencies, then good luck.
The best Jupyter UX for me now is VSCode. Just put an .ipynb file in your workspace and you get the notebook interface inside VSCode. Put `%load_ext autoreload` and `%autoreload 2` in the first cell, and use the same python environment you're using in your workspace for the Jupyter kernel. Then you can import libraries from your project, use them, and it's very easy to promote code from the notebook into a library. You can just cut a function from the notebook, paste it into a library, add an import, and rerun the subsequent cells to verify it still works as expected.
Cant you just set the parameters in the first cell?
>- I get an analysis that I like, but there isn't a good way to share it with others, so I end up just taking screenshots.
Export to md, pdf or html?
I believe that you can use https://github.com/tweag/jupyenv for this.
- I get an analysis that I like, but there isn't a good way to share it with others, so I end up just taking screenshots.
You can publish any Hex notebook with literally just a few clicks, and anyone you share it with can access it, or edit it, or fork it, without installing anything— or you can even make it public. You can easily turn a notebook into an "app" or interactive report if you want, hiding/showing certain cells or choosing cells to show only code/only output. You can just share the raw notebook though too.
- There isn't a good way to take the same analysis and plug new data into it, other than to copy-paste the entire notebook.
Super easy to duplicate a Hex project and hit a different table or data source, or you can use input parameters (like ipywidgets) to make one notebook parameterized and work on a bunch of different data sources.
- The process to "promote" fragments of a notebook into being reusable functions seemed very high-friction: basically you're rewriting it as a normal Python package and then adding that to Jupyter's environment.
You can promote any part of a project to a "Component" (docs: https://learn.hex.tech/docs/develop-logic/components) that you can import into other projects. They can be data sources, function definitions, anything. If you make upstream changes to the component, you can sync them down into projects that import it.
- There aren't good boundaries between Jupyter's own Python environment, and that of your notebooks— if you have a dependency which conflicts with one of Jupyter's dependencies, then good luck.
Hex has a ton of default packages in its already installed standard library, and all the dependencies are ironed out— if you have packages you want to use that aren't there, you can pip install them, pull them in from a private github repo, or ask us to add them to the base image. You can also run Hex projects using a custom-provided docker image if you have super custom needs.
You should *definitely* check it out if you have these pain points. Here's an example of a pretty complicated public Hex project: https://app.hex.tech/hex-public/app/9b882bc1-ead3-4f0b-87d1-...
And here's a simpler one I just made the other day on a cool Silk Road dataset https://app.hex.tech/hex-public/app/cdc1b8fe-144b-4a74-a5ef-.... There's a bunch more examples at https://hex.tech/use-cases. Happy to answer any questions!
The experience is a lot better than JupyterLab (which I am forced to use from time to time on SageMaker). The VS Code UI is cleaner plus I get a full language server which means I can rename variables and refactor fearlessly.
I also get full access to VS Code plugins.
And on a separate but related note, does it change the way you think about how you spend your time coding? (Assuming the costs do ramp up with usage such that time literally does equal money?)
There’s no IT and I can provision instances of any type (subject to limits) at any time.
Anyhow there a wealth of free extensions to customize it and the setup is really straightforward. I have version management git in a private GitHub project for version management. You can add extensions for rendering graphs in good quality and importing and exporting stuff is easy.
I have not been able to figure out why some people prefer to use Jupiter notebook as it is.
The other nice thing about VSCode is that you can extend it with VSCode Neovim (https://marketplace.visualstudio.com/items?itemName=asvetlia...), which runs a headless version of Neovim and allows you to do all the wonderful things that that entails, including stuff like VSCode's native multiple cursor implementation (and Lua config files!). All in all it's a great workflow, it's pretty light, and if you're paying for (or self-hosting) a beefy server it can turn any laptop into a powerhouse.
VS code takes care of spinning up the remote Jupyter server. All I have to do is create a new .ipynb file and everything happens automatically. Execution and disk are remote, only the UI is local. This is the magic.
It’s exactly like SSH except you have a rich client IDE in VS Code. The only data that moves over the network are your keystrokes and pastes and what is needed to display output in VS Code. You have to try it to see.
That is nice the company pays for all that cloud compute but for an individual it would just seem more practical to build a beast of a machine.
Most of my analytical work now is done in .py files, broken up into blocks with `#%%`. Real notebooks feel really clunky since adopting the approach.
EDIT: Googled and answered my own question. Here are docs describing the feature: https://code.visualstudio.com/docs/python/jupyter-support-py
And a video demoing what you are describing: https://www.youtube.com/watch?v=lwN4-W1WR84
Personally I've since gone full literate programming mode to the point that I care far more about the narrative and documentation (of methods and results) that I will build and modify tools rather than go back to the Matlab way. I have been looking at Quarto but haven't had the time to see if I can transition my existing (and target/ideal) workflows.
I know it gets a lot of hate but ipynb have a lot of advantages as a format for building small custom tools for modification/transformation. Most of the complaints ultimately seem to boil down to not having tools that do what you want. Only want to diff the code cells? That's easy in a python utility that loads the notebook and looks at it intelligently. You can also use pre-commit to modify the notebook and strip out things that don't belong in git.
(Also nbdev... exists... and is a good example of how tools can help. Unfortunately it's too tied to GitHub functionality and the developer is a GitHub zealot who is oddly brittle and takes offense and demands justification if anyone mentions not wanting to rely on GitHub)
- Same interface for analysis, scripting, and building more complex multi-file pipelines. I can also use the #%% notation to break up and debug scripts, which is probably teaching me all sorts of bad habits but it's something I find helpful.
- Similarly, as another commenter in this thread notes, .ipynbs just don't play as nicely with the other dev tools (e.g., Git, Black) and generally feel like second-class citizens in VSCode.
- I much prefer having the VSCode interactive window on the right, as opposed to having my output dumped out below my code block. I now find using the classic notebook style makes the document much longer and harder to navigate, particularly as I work with text a lot and I'm often outputting large chunks of text for inspection.
This noted, I think this is all possible because I'm rarely producing my final products in notebook format. Neither my boss nor the stakeholders I typically present to can (or have any inclination to) read code, so I don't really need a format others can execute or inspect. I just take the charts and figures and dump to presentations and other normie-friendly documents.
Anyway, thanks for spreading the workflow, and I will definitely try it out in the coming weeks.
I like to have the relevant code and output side-by-side, and dislike scrolling past outputs to get at code. Again, pure preference.
My screen copes fine with two tabs and the sidebar hidden most of the time, but more real estate would be nice.
What I'd love would be to pull tabs out into separate windows, like in a browser, and have the Jupyter output and variable inspector on a second screen. If anyone knows a way to do this (not new window) I'd love to hear. Last time I looked seriously this wasn't possible.
Naturally, if you need these things then .ipynb makes sense.
Big positives are how it integrates with the rest of the IDE so go to definition, debug cell, and data explorer just work.
Some negatives are a possibly onerous setup if not already using VSCode as your IDE (to get some of the IDE-like stuff to work), and how there isn’t exact parity on hot keys so muscle memory fails you occasionally.
It splits the editor into a UI that is run locally, and a server that does the heavy lifting on the remote machine. Conceptually it’s very similar to Jupiter, where you have a user facing front end with the UI run on JavaScript and rendered by your browser, and a python kernel backend, and the two communicate over pipes that can be run over the internet.
What it effectively means for VSCode is that you get a more seamless experience than I experienced with Pycharm remote development.
In any case, jupyterlab + jupyter-lsp gets most of the benefits for me.
JupyterLab feels like a clunky web based IDE. I check it every year or so and go back to Notebooks.
I used to and still run a Littlest Jupyter Hub: https://tljh.jupyter.org/en/latest/ for my org.
I keep thinking whether migrating to full blown JupyterLab is worth the pain.
With the improvements that Visual Studio Code has made in ipynb support there is even less reason these days.
The biggest thing keeping me on VS Code of course is full blown Copilot support. Whenever I have to fall back to Colab I feel 2-3x less productive.
My workflow is:
* Notebook for exploration/fiddling around 90% of the time is spent here - keeping state open is so convenient
* Extract/export code to regular .py for production
Even a simple improvement like remembering the sizes of various subpanels in the debugger sidebar will make me feel like I am not pulling teeth when I use it.
And don't get me started on inspecting the value of variables. If you are looking for more than an object within an object, you might as well go back to print statements.
At some point I felt like I was a bad dev for not using a debugger, but at this point I think I'm more versatile since I'm less dependent on finicky tooling to figure out what some code is doing... Every language has it's own debugger to learn, but logging (and good strategies for logging) works about the same everywhere.
something like
>>> d = { "a":1, "b":[1,2,3] }
>>> import json
>>> print( json.dumps(d, indent=2) )
{
"a": 1,
"b": [
1,
2,
3,
]
} >>> d = { "a":1, "b":[1,2,3] }
>>> d["d"] = d
>>> import pprint
>>> pprint.pprint(d)
{'a': 1, 'b': [1, 2, 3], 'd': <Recursion on dict with id=4516761856>}If it's for student, you could maybe get educational license, that is completely free for all their products, but I imagine it could be quite a hassle
The VS Code implementation of notebooks has too much vertical space for me.
I like Jupyter, but after trying out Livebook with Elixir I wish there was something similar in Python.
Smart cells and Toggling parts of code on/off is extremely useful features in a Notebook app.
The ipynb files are output artifacts. Why would you want to store them into git? It would be like storing compiled program binaries.
I actually started in notebooks and then learned to love the REPL as a simplified "scratchpad notebook." I'd say in many ways notebooks are an improvement that cater heavily to REPL-lovers, but that for some quick tasks, the extra complexity isn't always worth it.
Jupyter Lab is an IDE where you can open notebooks, files, terminals, etc all in one interface. Additionally, you can easily adjust the layout of open files (want notebooks side by side? Click one notebook's tab and drag it to one side of the screen. Want a terminal on the bottom of the screen? Open a terminal and drag its tab to the bottom of the screen. Want 3+ notebooks side by side? Click and drag. Etc).
I basically only use the Jupyter Lab interface when I'm working with notebooks (sorry about using the term notebooks so much, it's a synonym for a .ipynb file as well as the name of a server mode Jupyter offers).
I still dont quite get that is diff btw jupyter nb vs jupyterlab.
My understanding is notebook for single user locally, while jupyterlab is multi-user running on server, something like that
Jupyter Lab lets you have multiple files/directories/terminal/csv files/json files/html pages/etc open at once in the same browser window.
For now, I need jupytext (which is a great extension), so then I stay with v3.6 until jupytext is resolved.
For this announcement - or in the changelog - it would be great to know why it is a major version bump - what is the major compatibility change?
P.S. Juptext is a pretty cool extension https://pypi.org/project/jupytext/
jupytext is great, you can also review your notebooks as `.ipynb` files using GitNotebooks [https://gitnotebooks.com]. (I'm the solo dev)
I think if they made breaking changes to the notebook format, they'd mention this right away.
Additionally they often point you at their terrible discourse forum for asking questions. More often that not I don't see a good answer there either, when I merely search for one. I think their gitter channel has worked best for me so far, when they did not point me to that forum.
Typescript also helps a bit when compiling.
Sometimes I visit an old bookmark, that looked like part of official docs by the URL, but find it 404ing. Ultimately I agree, that good and accessible docs is not the project's strong side.
I guess I will have to figure some things out again soon when updating extensions to version 4.
I've been trying to work out exactly how it works myself for the last hour, and there is no clear indicator on how to activate this mode. No documentation either, barring: https://github.com/jupyterlab/jupyter_collaboration
What am I supposed to be doing here?
But, yeah, you can develop python software in Jupyterlab if that makes sense for you.