Show HN: Jupyter kernel using Poetry for reproducible Python package management
github.com
github.com
Author/OP here. Wrote this small but mighty package to make it easy to create, run, and share Jupyter notebooks with reproducible environments. Was definitely inspired by Julia's Project.toml/Manifest.toml package management.
Ended up building it to support my startup Pathbird which lets instructors build engaging, interactive courses using computational lessons. A big value proposition there is that every student gets exactly the same environment (no dependency version issues!) so we want to make sure that the environment running on Pathbird exactly matches the environment on the instructor's computer.
Example[1]
[1] https://github.com/linz/stac/blob/6a82e92432945777fbd49631b4...
NixOS on servers, Nixpkgs-darwin instead of homebrew (mostly), and Nixpkgs on Fedora for a laptop. Solid results so far.
I worked on a similar project for my last company. However the goal was portable notebooks that could be executed anywhere, and rapidly deployed to various environments with minimal installation. I ended up using a combination of LinkedIn's shiv library with Netflix's papermill library.
The point was to turn a notebook into an executable runtime with all the dependencies embedded. I don't want to get into the specifics of why we were doing this but it had to do with my previous employers product which is targeted at no code and low code folks, and integrated Jupyter notebooks.
I think the use of poetry is very elegant here. And the fact that you can reuse the same kernel easily is a huge plus.
I have been doing this with a hand edited kernel.json for a year or two and it works perfectly.. hadn't been looking forward to demoing it to coworkers due to hacky setup.
You've solved the only problem then perfectly!
Way more than half of the effort on this project when to looking in to how to distribute the kernel.json file (I ended up copying from ipython). The actual code that runs is little more than a Popen.
Translation: I've never used poetry, don't know anything about it, and I'm going to talk out of my rear end and assume that it doesn't use conventional, existing venv/pip under the hood.
There isn't a single python project out there that can't be worked on with poetry. It creates a venv. Poetry is doing a lot of things to manage automatically activating and deactivating it, etc, but it's just a venv.
https://github.com/python-poetry/poetry/issues/4231
Is one issue. If you even transitively depend on pytorch poetry will break and you will need hacks to make it work or just give up.
also, a lot of python packages don't seem to follow the idea of semvar very well. Especially when you compare it to other communities like Go.
> The Python environmental protection agency wants to seal it in a cement chamber [...]
Also, it really doesn't work super well if you need to work on multiple development projects simultaneously, and when "just change the way you structure your entire project" is poetry's answer to this, the solution is "okay we won't use poetry".
Have you ever used poetry before? Are you implying that it can't be used to install Python dependencies that have c/fortran like pandas?
If that's what you're saying that is not correct.
Poetry does seem much more tied to python (for better or worse).
Keep in mind at least in mamba's case, not being tied to the python interpreter is very much a feature, not a bug: https://medium.com/@QuantStack/open-software-packaging-for-s...
BTW, sorry for being off topic, but Lex Fridman has two great recent long form interviews with Anaconda founders Peter Wang and Travis Oliphant. Good discussions!
I assume there is some political difference or acrimony which means the actual sources about core conda are entirely silent on the issue - I’ve privately replied to people expressing concern over this on the conda mailing list several times and people are always relieved to know there is a solution.
Seriously, it’s the only way we can tolerate using conda as an ecosystem any more, and only ever feel this pain of 15-20 minute feedback cycles preparing packages with conda-build (for which boa is a promising new replacement we haven’t moved to yet)
In conda, you can declare this properly, and you can have one true eigen in your environment.
No, they are implying that many scientific python packages and dependencies of python packages have dependencies on non-python libraries and packages written in c/c++/fortran. If you want to manage and reproducibly pin these they also need to be tracked, because otherwise you just end up with "whatever the underlying OS has installed", if you are lucky, or "compile failed" when installing the package if you aren't.