The Jupyter+Git problem is now solved
fast.ai
fast.ai
I don’t know whether it’s the Data Science culture or Jupyter but there is a big lack of discipline in writing maintainable code in DS and non-existent git support is part of that.
I always strongly discouraged developing models using notebooks, instead advocating for using .py files and then using notebooks for sanity checking data.
I don’t have any clever ideas for how we can move past Jupyter but the sooner we do the better.
In line with a nephew comment of mine, I feel that bringing the immediate interactivity or iteration cycle of notebooks to the development experience would help a lot, and not be too bad a thing for common development either.
I've heard of the related nbdev project, which seems like an interesting and compelling idea. But it'd be nice to see the reverse: something that makes ordinary python development more immediate than using a debugger/vanilla REPL.
Once you have a consistent set of questions, and methods to answer them, then yes, copy off the relevant chunks into their own scripts and source these when using similar data to bring you up to speed, and modify them to your tastes.
The issue with starting off with an external script initially is the distracting temptation to refine your code so it can be better used with future data, despite not yet having seen or not knowing what that future data is like. The initial "play and explore" phase of an analysis is very important imo, and notebooks really facilitate that.
The Jupyter allow you to load big chunk of data or some large model only once, and then use it for experiments in other cells. It is hard to replace this feature with plain `*.py` file. For me, this is the killer feature.
- I work with several people who are purely data scientists, and I lean on "culture" rather than "Jupyter". In some circles, probably influenced by academia, programming is considered to be low status work. You are not going to solve the problem by switching to .py files, even though for most tasks literally anything is better than Jupyter.
That they don't use git, or that git wasn't originally even a concern, is a consequence of that low-status perception. If you pitched something to the developer community and told them "oh, by the way, you can't use git", they'd synthesize tomatoes out of thin air to throw at you.
- I'd carve an exception for stuff that satisfies ALL of the following: (a) is self-contained in a single notebook, (b) has no dependencies on anything non-standard, (c) is demonstrative in nature, or a personal exercise rather than production software. For example, I wrote my solutions to the Advent of Code problems in a notebook and I liked the experience, especially how you could mix math and code.
I think notebooks are a great way of presenting findings and showing your workings at the same time. If that was their main use then I wouldn’t have any issues
If you want to make the argument about "two kinds of people", I think it's more about A-type/B-type data scientists in the industry. I'm really mostly a developer and not a data scientist, but when I assist in DS tasks I wear a distinct B-type hat, and that informs my perspective. A-type people have different priorities and that's fine; my gripe is when you try to import A-type practices in a B-type scenario.
That way, if you put your classifiers in joblib pipelines, once you're done with fitting steps you can just export your trained classifier with:
joblib.dump(pipe, "trained_classifier.dump")
And resume your work with: joblib.load("trained_classifier.dump")
Considering this works for any Python object, a lot of heavy lifting can be exported for later (swift) use this way.One of the main differences for me is that the .py file is run in its entirety (outside of if/else blocks for loading data). That usually corresponds to multiple cells of a Jupyter notebook that one would need to Ctrl+Enter through, where missing one would cause a problem.
The second is just how you can decouple the code and the terminal - it's a personal pet peeve that Jupyter notebooks jump around when running through cells - I don't want to be scrolling all around just to reset some variables to their original values, and it's really nice to run a whole .py script and see an output side by side, where the script is much longer than my screen. I can keep it open at the important part in VSCode, and change some intermediate process, and let all of the ad hoc plotting code remain at the bottom.
Finally, the biggest difference for me is how figures behave - the way I have it set up is that they open in their own window and remain interactive (can zoom/pan). I know you can do it in Jupyter as well, but the workflow really emphasizes inline plotting with non-interactive plots, especially when it comes to sharing them. But with the .py script and IPython command line, I can open up 5 figures, tile them however I'd like, and then refer to them by name/number in my script, so I can clear and overwrite them however I'd like, and they don't close or move around. This makes comparing things very easy, like how changing a parameter changes the rest of my analysis.
Lastly - the way it is set up is more like Matlab... whoops, but I think their workflow is much more ergonomic than a notebook. However, for sharing with other people, I usually just copy and paste the various parts of my scripts into a notebook, as that is the de facto standard.
one of the key things that a good notebook system must allow you to do is to mix something like markup format + LaTeX + source code. writing math-heavy documentation and explanations is simply impractical and limited (readability suffers) if done in comments. jupyter however is severely limited as it is unreadable in its raw format and therefore does not play well with a version control system such as git
instead there is a solution that allows one to do everything jupyter does good with the additional benefit that it plays with version control really well - ie org-mode [1]. the only difference is that instead of using a browser to interact with it, you use emacs. the added benefit to this is that you can also use full-featured key bindings (emacs / vim) and even integrate a language server for auto-completion [2]
EDIT: moreover the list of supported languages in orgmode far exceeds that of jupyter [3] (or did the last time i made this comparison)
[2] https://emacs-lsp.github.io/lsp-mode/manual-language-docs/ls...
[3] https://orgmode.org/worg/org-contrib/babel/languages/index.h...
i'm sorry to burst your strong held convictions but you can choose any of the following
a) use any key-bindings you like including emacs, vim, cua, or combination of
b) use org-mode without any knowledge of more advanced emacs commands (except basic knowledge of using an editor)
c) drink some milk (gotta have strong bones) and learn how to use the emacs system including emacs lisp and have one of the most advanced computing environments in existence at your service
sorry, but that is just the reality
The latter needs to accept that most users, particularly scientists, reject out-of-hand anything requiring configuration or compilation no matter how trivial.
But it's all moot since org-mode is largely promoted by non-scientists (computer science is not a science), and should a wysiwig-inclined scientist ever get past the emacs obstacle, he'll balk at the awkward BEGIN_SRC incantations.
My academic training is in physics and mathematics. I was introduced to programming in my computational physics class. We used emacs as our editor
Emacs is such a fundamentally different paradigm from all other IT tools/editors that it just doesn't make sense to recommend a specialized tool with a steep learning curve and non-transferable skills when it's not ubiquitous and more standard alternatives exist without it which do OK. It doesn't matter that emacs was historically the first and everyone else decided to go in different directions, that's just the reality of today.
im curious which transferable skills you think Jupyter has
> Emacs is such a fundamentally different paradigm from all other IT tools/editors
its not. you use a mouse and click where you want your pointer to go. you use a keyboard to type
> steep learning curve
this is very much like saying that linux has a steep learning curve and you refuse to touch ubuntu because you are scared to blow up your computer
Most people do say that. What programmers cannot grasp is progress is about giving people what they want, not what is rationally best. Most scientists want to knock out a paper or presentation, and make it home on time for dinner.
are you their union rep ?
I think you're just suffering from selection bias, given that this is on HN.
Add me to the list of people who began using Emacs and Org mode during academia in a non-CS program.
Furthermore, go look at the Emacs conference - you'll find a significant number of speakers are not CS folks.
That being said, nobody's gonna learn Emacs for org-mode, unfortunately.
If org-mode wasn't backed by Emacs, it would merely be a markdown substitute hence much less useful. There are many org-mode clones for modern editors like neovim or VSCode, except all they offer is front-end features (highlighting, folding, node manipulation etc). There is simply no reason to use those over a decent markdown editor. So I think you have this backwards; Emacs isn't holding back org-mode, rather much of advanced org-mode features are made possible and distinguished by the fact that it builds on Emacs.
also the ecosystem is huge and chances are that the configuration you are after is just a package-install away
Configurability is a strength of a system, but it is no an answer to a difficult learning curve. A user must first understand the system in order to configure it appropriately.
Even at the level of key bindings, the user needs to understand the relative frequency and importance of an operation to choose an appropriate key combination. Universal reconfiguration may even make the system less learnable, if documentation and tutorials can't assume a reasonable default configuration.
In my opinion, configuration is great as one of the final steps of a user's journey, taking the system from something that works to something that sings. It's just the wrong level to sell benefits to beginners.
i have a feeling that people who write these things have never really tried emacs beyond opening it and getting annoyed that ctrl-c/v/x don't work (at first) the way they are used to
emacs is not key-binding-based, it is command based. if you change a key binding its not like you can wreck anything as you can always call the command prompt by M-x and search for the command that you wanted some key binding to perform. key-bindings are just shortcuts to commands so i think its best to listen to your fingers and form muscle memory and then assign them
what are your most basic commands? copy, paste, select, start/end of line/function/class/paragraph/etc, move by word/sentence/etc, save, exit? these are not that many to set to whatever key combinations you want. i wish my browser had at least this level of extensibility
We get it - what makes you think we don't. We are merely pointing out a superior solution.
Like back in 2004 I would tell people how many of their problems would be resolved if they switched to Linux. Fast forward two decades later, the statement is still true, and most people still don't use Linux. But it wasn't a problematic thing to point it out to them - be it in 2004 or now.
(It's a lot easier to use Emacs than switch to Linux.)
If that were easy, then that other editor would be an emacs. There's a reason emacs gets all these modes, and its because writing extensions is easy. If I want to extend something like VScode or Pycharm, or whatever, it is a massive undertaking.
I guess you meant literate programming.
Literate programming is different from interleaving console inputs and outputs and random paragraphs in the same document.
Even if we expand the original idea to comprise that, literate programming is much more than that.
How many python packages are written in literate programming style?
How many programs written as notebooks would be actually better if they were structured differently?
The issues with notebooks - in general - are unrelated to literate programming. The notebook format is convenient to have some kind of “interactive” programming though, rather than “literate”.
interactive programming is usually handled by the repl, for which you do not need a notebook
Note how babel is presented, by the way (last point in particular):
Babel augments Org code blocks by providing:
interactive and programmatic execution of code blocks;
code blocks as functions that accept parameters, refer to other code blocks, and can be called remotely; and
export to files for literate programming.
https://orgmode.org/worg/org-contrib/babel/intro.htmlhttps://en.m.wikipedia.org/wiki/Read%E2%80%93eval%E2%80%93pr...
When you get into MLOps as well, having .py templates actually makes the Data Scientist’s job easier as they can plug and play their models into a system that tracks inputs, outputs and changes for them
Personally, I start every notebook with
%load_ext autoreload
%autoreload 2
then develop production quality code in .py files.1. Experiments in notebooks. Notebooks are saved under git but mostly as a backup, I don't care how nicely they play together. I don't get why you would discourage notebooks for running experiments, doing it with .py files sounds kind of miserable.
2. Services and library code in .py files, under version control, just like any other software we write.
Having your services and library code as .py files you can import in is great.
The issue comes with how to move from experimentation to deployment. If you already have services/library code as .py files you make your life a lot easier. The issue comes when everything is spread across multiple, poorly documented notebooks. If you're working with an MLOps team it makes their life a nightmare to take those notebooks and conform them into something usable.
Jupyter is great when it is used in the right way.
People use a great tool in a poor way and then broadly condemn the tool.
And any tool that is sufficiently flexible to be broadly useful can be used in very poor ways.
Jupyter is great, it gets me over the barrier potential for starting a task every time. I build and prove out an algorithm/task piece by piece. Once I'm happy, I move the meat of it to a function in a .py file, and move the code I used to test the algorithm to a unit test function. Delete the duplicated bits and replace with imports, and then what remains is a tutorial/demonstrator notebook using the function I wrote and maybe some nice plots to go along with that, that I wouldn't put in a unit test (nor that show up in docstrings). This can be converted to sphinx docs if the code gets big enough.
What a great tool for incrementally building software! In my world, I build brick by brick, not all at once. Jupyter is a key to that process.
Notebooks can be clean if you follow some rules:
1. Code flow always goes down: holding Option+Enter should execute all fields without any errors. Don't do `x += 1` if `x` is defined underneath.
2. All blocks are idempotent: running any block 5 times should produce the same result as running it 1 time. Don't do `x += 1` unless `x` is defined in that block.
3. Keep block-local variables short and block-global variables long. Don't do `x += 1` unless you are not using `x` anywhere else.
Also, the Table of Contents extension [1] is a life-saver for making long analyses workable.
[1] https://jupyter-contrib-nbextensions.readthedocs.io/en/lates...
Personally, I like to label block-global variables in capital case (like PEP8 constants), so as to make them easy to spot. Being formatted like constants also causes me to think twice about altering it after instantiation.
It's not perfect for in-person working because a single person should always keep their work under version control, and they should be able to view meaningful diffs to understand the history.
It would be nice if there were something that would make out-of-order problems light up, the way that code editors can highlight errors while you're editing. A limitation of "browser as editor" is that it misses out on some of the powerful things that code editors do today.
Another thing is to put things in functions, so temporary variables are disposed of. That's a halfway step to putting things in .py files. A benefit if .py files is not always that jupyter is bad, but that variable scoping is good hygiene.
As a table of contents, I usually write some markdown up top with links to markdown HTML anchors elsewhere in the page, which themselves also have a link back to the TOC. Works pretty well. Will have to check out that extension.
4. Use a function in each block, that is defined and the called with appropriate arguments, and return the values of the block. This prevents proliferation of global state - and the functions are really easy to move out to .py modules when things have solidified a bit.
While most programmers have reached this conclusion, they're generally not day-in day-out jupyter users. They need to understand *everything* is transient for scientists who optimize for proof-of-concept and publish-and-forget-it paper writing.
Which itself is a huge problem.
Happily this mindset is changing, at least in some scientific all fields. For example, in particle physics proposals a document ("data management plan") much be written describing how that unconscionable attitude will not be taken with the experiment's data and software. That said, this transient mindset and derision of real software skills is still fairly prevalent in this field.
More "nature of the beast" in my opinion. Science measures itself by how many alluring women it can date; engineering, by how long it can keep the wife happy.
Even for short-lived experiments reproducibility is important. So much of today's experiments ultimately rely on complex stacks of software to get their results and on data which humans can afford to acquire only once. Preserving both is necessary for future re-validation or reuse.
This problem should not be excused.
I also use it a lot for experimenting, parameter tuning etc. It's not too bad to have it explicitly distinct from production level code. Run/tune/experiment in notebook, once you're happy with the model -> code it up in .py file(s). Also great for quick presentations :)
However, the fast.ai team is actually doing a pretty solid job running everything off notebooks. So if I wanted to go that direction (and skip the .py files) it's that project I'd look at for how to do it.
Do you mean you _don't_ want to give the students shell access?
By default you can run shell commands from within a Jupyter notebook be prefixing them with `!`.
and if it's just the professor's lab, just a jupyter lab instance with one password works too.
It was easy to use as a presentation, with figures and plots embedded. With controls enabled, you could demonstrate what varying certain parameters would do and pitch proposed cleaning profiles.
I was rather easily able to send a directory and it's notebooks/data sources to colleagues in the water sciences team so they could validate my results on their own (they were luckily also familiar with Python and Jupyter), and caught some minor bugs in the pipeline.
This was all much more collaborative and concise, and I feel Jupyter played a huge part in it.
Once it was done, it and a "final draft" pdf were added to the Docs in the repo and the pipeline was written out into a full application of it's own.
Add in # %% and it makes the code executable in isolation. Like a cell in jupyter. This means i can write a py file. Execute bits indepentantly and when im ready to check it in just remove the block comments
Notebooks never caught on with me.
And then serve the results as a static page under my.logs.intranet/my-job/2022-02-14/my_recurring_task_notebook.ipynb.html
You can get so much more context regarding what went well or wrong compared to browsing through log lines in some more or less user friendly tool.
From the beginning of the article: "With nbdev2, the Jupyter+git problem has been totally solved. It provides a set of hooks which provide clean git diffs, solve most git conflicts automatically, and ensure that any remaining conflicts can be resolved entirely within the standard Jupyter notebook environment."
It makes notebooks testable, composable, versionable, and more.
"we" have to learn and teach the next sets of people new to computer science
Partially the lack of discipline comes from the implicit data dependancies between cells. Variables are all globally scoped and unless you ensure the notebook can be ran top to bottom its easy to introduce subtle bugs. I believe Julia's https://github.com/fonsp/Pluto.jl solves this issue quite well.
Another part comes from cells that should really be functions. In my opinion this is because functions are 2nd class citizens compared to cells, and could be improved with UI (function cells? node based programming?).
Programming is more than just manipulating text, so why shouldn't tools move in a direction of just being fancy text editors?
Once this approach is found, a good clean-up/refactor is strongly recommended, to then start a proper software development that will create a live product from the found approach. I call this the switch between research mode and development mode, and it has strong parallels to the way R&D is done in many industries. I believe a lack of understanding of this dual nature of ML is what causes many of the problems in MLOps: plans that don't take into account the research time and risk, mixed teams where engineers don't understand the initial nature of DS work, attempts to put notebooks containing research code in production, etc. Even planning for the refactor doesn't solve it all - what will happen when the next generation of a model has to be created? Will the refactor Ed code be forced on the DS and ruin their research productivity? Will they start from scratch again and not only lose all the refactor/dev cost but also make this a recurring cost? I have been looking for answers for this for years now, and found none so far.
Source: I've been working with data for 27 years, as a data engineer, data architect and data scientist. When I do DE, my code is considered high quality by my peers, but when I'm doing DS research, I know I write bad code - and I won't change that. It's more productive to work this way and do the big refactor (possibly leaving the notebook env behind along the way) than the alternative.
It looks like this tool, nbdev2, solves a real-world problem for Jupyter users, including me, with zero effort required to use it every day. It relies on clever hooks to get git to treat cells as first-class citizens (as opposed to lines of text, the default). Nice! Based on that alone, I would expect nbdev2 to be widely adopted over time. In fact, if it works as well as advertised, it should be incorporated into Jupyter. I, for one, will be giving it a try!
If you use Jupyter to solve problems in your domain of expertise, feel free to ignore all the smart-sounding software engineers who will poo-pooh this tool only because they don't like notebooks and don't want anyone to use them. No matter what you do, there will always be people who look down on easy-to-use tools that enable scientists and practitioners from other disciplines to write, run, and explore ad-hoc code on-the-fly.
EDIT: nbdev2's authors are on this page, answering questions. Thank you again!
There seems to be a lot of discussion in here around the pitfalls of jupyter, and notebooks, and the poor coding practices of data scientists. If you haven’t read the article or used the software I’d like to highlight that all of these (legitimate) complaints are exactly what nbdev2 was created to address, and in my opinion very successfully solves.
The way it works is that everything runs off a master notebook, and then with one command: libraries are built, git diffs are magically fixed, tests are run, documentation is automatically created. It doesn’t fundamentally change your workflow in any way, it just abstracts and automates away all of these pain points.
There’s a reason that everyone uses jupyter notebooks. They are fun to use, they are great for exploring and developing ideas. And (minus the aforementioned git collaboration issues) they are great for sharing with others, which is a huge part of the wider data science ecosystem. We don’t need to recommend avoiding notebooks, and allege they are just for beginners. We need to use tooling which addresses some of these final issues with writing mature software. And I'd like to thank the authors of nbdev for this.
The people who look down their noses at notebooks can continue to do so – but what they will find is that nbdev quite effortlessly leap-frogs over these sneered complaints, and allows you to write better software more productively.
IMHO the Jupyter+Git problem stems from the ipynb format. Jupytext does it "right" in the sense that you can work in .ipynb, and diff in .md. But as long as the base format is diff-unfriendly, all tools are methods of indirection of format complexity to tool complexity.
That's not to take away from the tool -- it looks great. It also takes the D out of BYOD, which is a win. But I think "solving" it means that anybody who receives an ipynb is able to just look at it out of the box, like plain text, so we're still a ways off.
We realised after working with Jupyter+Git for a while that the pain-points were actually with Jupyter editors (and/or their conventions) rather than the format, because they do things like store user-metadata in the file which pollutes diffs and leads to merge conflicts.
In fact, if Jupyter editors could handle merge conflicted files, we wouldn't need a custom merge driver either.
I also don't understand what you mean by discipline. Yes you need to make sure that everyone has the jupytext extension installed, but that just becomes part of the needed dev environment. After that the whole experience becomes completely seemless.
I've started rolling my own little plugin utilities and, so far, I have a (very) rudimentary notebook-like interface in plain text.
Combine a proper attempt at such a thing with a good interface to a background ipython kernel system (for which ipython could do with some minor enhancements AFAICT), and you'd basically have the best of all worlds (all plain text editor features including version control and personalisation and customisation, and the iterative advantages of notebook code-cell-based runtimes).
Hopefully, with such a combination functioning well, there'd be an emergent feature that allows one to more easily get interactive with a code-base for the purposes of understanding, debugging or developing it.
Personally, my biggest gripe with Jupyter at the moment is that a few years ago they decided to try to create a quasi-IDE (where they'll probably be beaten by VSCode) rather than improve the general utility of the kernel (or kernel protocol/interface?) and/or the essential notebook UI.
It's a personal gripe, and there's clearly value in the web-first interface they've made with JupyterLab (despite the not insubstantial growing pains that project has faced), but, watching ObservableHQ and Pluto (for Julia) focus just on the core notebook interface, while VSCode have focused on the IDE side and easily incorporated or recreated the now rather old/simple Jupyter Notebook interface, both with success, seem like some vindication on my gripe.
The old notebook was painful to extend in JS compared to writing lab extensions in TS. In Jupyter Lab 3 they have taken questionable steps, but so far I have been able to work around issues.
The ObservableHQ interface, for instance, I'd classify as just a notebook interface. IE, individually manipulatable code-cells with a shared runtime.
And yea, JupyterLab is better than the classic IMO. But, until recently I'd say, the notebook part of the interface hasn't gotten much love at all, while there've been steps, due to popular demand it seems, to provide alternative UIs that strip away much of what they've added on top of the notebook (ie, simple mode, and now Jupyter Lite).
I haven't really got experience writing extensions in the old Jupyter notebook, and hardly any with JupyterLab, but my experience with JupyterLab was frustrating because it felt like they really killed the ability to implement small and hacky plugins like you could with with the old. This always struck me as a shame. A necessary one perhaps given what I presume is the increased power of their new framework. But it always felt like there was a mismatch between the complexity of the plugin framework (which is a full web-dev experience) and the base features of the "product", where customising my test-editor is now much easier AFACT.
What issues and questionable steps were you thinking of?
2 things come to mind right now:
Starting with JupyterLab 3 (maybe 3.1), JupyterLab removes query arguments from the URL. Query arguments were the only way I know, to give arguments from the outside of JupyterLab to JupyterLab. Any extension, that relies on arguments given from the outside would break, just because JupyterLab removes query arguments, which were there since the beginning and did not do any harm, at aleast any I could tell. But suddenly this was taken away, without proper alternative. Now you have to hook into their "router" to quickly grab those arguments, before they are gone. This seems silly to me. Why randomly delete query arguments? They are there for a reason and since JupyterLab does not add any of its own, I cannot understand this decision. Simply seems to make it less powerful a tool.
The constant nagging about posting in their community JS-only forum. ("You should post this in the forum.", "Have you seen this post in the forum? links to forum") Why can this community not handle issues in issues, which can be easily found using a search engine. Why hide everything behind a JS-only forum, which one has to create another account for or associate ones Github account with? Whenever anyone gives me a link to the forum, where supposedly the answer to my question is, I keep thinking: "Ahh great, why did you have to hide it in there? If you had documented this in an issue, I would have found it via search engine and the thing would not have wasted my time and neither would I have had to waste yours." -- something along those lines. When I find an issue and its solution, I still post it as Github issue, so that other people can easily find it, without signing up to their forum.
> I haven't really got experience writing extensions in the old Jupyter notebook [...]
I have done that a few years ago, when JupyterLab was still alpha versions. It worked, but the typical JS mistakes plagued me. JupyterLab is of course using TypeScript, which helps a lot with avoiding silly mistakes. However, I do think there is something to what you say about no longer encouraging the quick hack. Some functionality took years to appear in JupyterLab, but was already available for Jupyter Notebook, before JupyterLab took off.
Jupytext will convert Notebooks (.ipynb) files to Markdown (md) and Python (py) 'on the fly' (while working in Notebooks).
- Markdown files can be added to git
- Python and .ipynb files are added to .gitignore
- Python files allow 'chained' import of notebooks (*.py verions), which allows to split larger notebooks into multiple smaller ones
This is my folder structure: .
├── notebooks
| ├── notebook1.ipynb # automatically generated from md
| └── notebook2.ipynb # automatically generated from md
├── md
| ├── notebook1.md # versioned in git
| └── notebook2.md # versioned in git
├── py
| ├──modules
| | ├──__init__.py # empty
| | └──tools.py # use for cross-project base tools
| ├──__init__.py # empty
| ├── notebook1.py # automatically generated from md
| └── notebook2.py # automatically generated from md
├──jupytext.toml
├──.git
└── README.md
See an example here [2]
Jupytext is mentioned as a 'potential' alternative. Re the "save" cell output: I usually produce html-files at the end of my notebooks (see the example), and add those either to git or auto-upload to an external webserver. The html is standalone and includes outputs, table of contents, and images (example [3]). I would advice against versioning all outputs (images) in git.
Very happy with this approach for a long time now. Jupytext increased my productivity by a hundred percent.
[1]: https://github.com/mwouts/jupytext
[2]: https://gitlab.vgiscience.de/ad/yfcc_gridagg
[3]: https://ad.vgiscience.org/tagmaps-mapnik-jupyter/01_mapnik-t...
Specifically, it doesn't handle the situation where you need cell outputs in version control -- since in that case, you still need the notebook, which results in all the usual problems occuring. With nbdev2, you don't need to think about anything or do anything special, and stuff like GitHub notebook rendering, nbviewer, ReviewNB, etc all just work. You just run a single command (`nb_install_hooks`) and that's it.
Also, no-one has to install anything extra to view your notebooks, since they're stored in the regular notebook format.
If you do a diff between two revisions where some figure changed you essentially will be swamped by the diff in the figure making it difficult to find what actually changed. Now tools like nbreview get around that, but now you're forcing everyone to use the same dev tools, and can't look at diffs any other way really.
> but now you’re fixing everyone to use the same dev tools
No they’re not. You can continue using whatever approach you’re using. Attempting to shut down alternatives like this though could be seen as forcing everyone to accept whatever that status quo and lowest common denominator solution, even if their dev tools could support something better.
Jupyter's ipynb format is only slightly more amenable to git than say an MSWord doc. Nbdime and friends will never get you to a point where git+jupyter will be worth the ugly.
That sounds like a nightmare. Why would you want to develop a library in a jupyter notebook?
> The solution presented here is the result of years of work by many people.
It's a bit depressing that it came to this. It's hard not to think that it was a mistake from the beginning and that the format should have been based on using special comment markers in valid code, together with an accompanying JSON metadata file. Or something like that. One way or another, we have a very strong tradition of storing code in plain text files, not embedded in strings in JSON or otherwise embedded in any opaque format. Maybe there'll come a day when it's appropriate to abandon that to get some advantages, but I don't think that day was the original creation of Jupyter. I know it was created by thoughtful and expert software engineers, but I feel that it was a mistake and it's actually made a lot of data science / academia-oriented people less qualified to participate in industry software engineering, because of the poor practices forced upon them by the inability to use git with Jupyter, and notions like developing library code in notebook cells.
Is it because Jupyter users in particular don't typically understand that there is a formatted text file behind the notebook, or how merge conflicts work?
That's what it feels like to be forced to drop into a normal text editor rather than using the normal notebook ui to fix the conflicts.
Suddenly binary formats could become mergeable.
Semantic diffing needs something like pijul, but a system taking advantage of this doesn't yet exist. Pijul avoids some merge conflicts by design, won't do the wrong thing, and handles conflicts correctly: we still need tools with a fuller awareness of what strings mean to have rich semantic diffs.
Has been my go to for this. It seems like nbdev2 is fastais own cooked solution with a bunch of other tools.
JupyterHub
Jupytext - converts ipynb to py
Nbstripout - strips all output from ipynb
Nbmerge - resolves merge conflicts
Vim-jupytext - vim plugin to auto convert ipynb to py
Papermill - parameterize notebooks
Git
Pandas, Altair - data analysis / Visualization
Phabricator - code reviews of notebooks
Vimdiff + vim-jupytext - diffs in terminal
This solved all my jupyter problems.
The nice thing about markdown-like notebooks is that they play well with git. The nice thing about jupyter style notebooks is that they contain all the content needed to actually _read_ the notebook.
https://github.com/jupyterlab/jupyterlab-git
Edit: Looking at the source, it does appear to use nbdime under the hood.
Definitely recommend checking it out if you haven't already!
Do I have that right? Because that sounds /insane/.
Also, unneeded metadata is removed from the notebook when saving, so there's less changes to merge.
Both these two things are done using standard hooks built into each of git and Jupyter. That is: git is written in such a way that it can fully support non line-oriented formats. We just took advantage of that capability.
<<<
side1
|||
ancestor
===
side2
>>>Idk, like importing some data and doing some analysis / forecasting?
Most notebooks appear really bad quality. Worse internally.
Better off looking at some excel https://github.com/martinshkreli/models
It stains the article from the very first paragraph.
I think it's quite unfortunate that they did not consider that the format would integrate well with version control systems when first designing ipython notebooks.
If you give me a counter example, good for you, but my statement holds true 99%.
But you are saying here: "If you leave git diffs in your files, whether Jupyter notebooks or otherwise, and run/compile them... They will break."
Have you changed your mind in this thread? Or what's your objection?
But if you don't know that git modifies files when conflicts, then you're an interesting and rather unexpected audience, I assume.
Meaning that for the typical git user, meaning, knowing about git diffs, the behavior is expected hence not broken. The files end up in an expected broken state, but git does not break them per se.
If you still disagree, let's just settle that we disagree and be done with it.