Nbdev: A literate programming environment that democratizes best practices
github.blog
github.blog
My question is about the coding style - @jph00 I’ve read your fast.ai style guide and worked with APLs like q/KDB (written by Arthur Whitney who you cite).
My experience is that brevity is great, until you need to collaborate or have individuals working on small parts. That was my experience as well trying to write an extension to the fast.ai code (where I had to read large amounts of source to understand how to implement a small change).
Given that a key motivator for literate programming is collaboration/communication, how do you think about this?
Case in point, here is a random notebook from the fastai repository, a python file would be simpler to read and shorter: https://github.com/fastai/fastai/blob/master/nbs/09b_vision....
Another example (and bonus points if you can spot the rogue semicolon).
https://github.com/fastai/fastcore/blob/master/nbs/02_founda...
https://github.com/fastai/fastcore/blob/master/fastcore/foun...
The packaging is also really well thought out. I don't have to stress out about connecting setup.py with whatever publishing system we have. The settings.ini makes things sane and I can bump the version whenever I want.
A get a lot of skeptical looks when I say the source code is in notebooks, but that's just syntactic sugar for the raw source code. You still get to edit the raw code files and with one command sync everything with the notebooks. From my point of you it is close to a pareto improvement over traditional python library development.
R illustrates this very well - as an R user you can have your pick between RStudio, Jupyter and Rmarkdown and the overwhelming majority of users pick RStudio and notebooks are reserved only for a niche set of use cases. It also speaks volumes that almost no one writes R in Jupyter even though it's supported very well - R users just have better options available to them.
I wrote a detailed blog post about the differences between Jupyter and R Notebooks years ago: https://minimaxir.com/2017/06/r-notebooks/
There's a big difference between a line-oriented plain text REPL, and a Mathematica/Jupyter-style notebook REPL, especially when you want to mix and match your charts, image outputs, rich table outputs, interactive JS outputs, and so forth. Also, for experimentation, where you want to go back and change things to see what happens (e.g. very common in data science) I find it much easier and more understandable in a notebook.
I have a video where I show the difference between these styles of working in some detail: https://www.youtube.com/watch?v=9Q6sLbz37gk
Though, I think I get what you mean as I reflect on my own dev experience. As a mainly C# dev which has a very very limited repl experience (I would and should say no repl experience but someone will yell at me about csi.exe or dotnet-script) I have seen people using notebooks for want of a good repl, but I’m curious why anyone writing python would.
(try "sudo apt-get install ipython && ipython" on your Ubuntu/Debian system to try it out)
I would not call Idle unpopular. But it had a limit for growing for sure.
But Notebooks are mostly coming from the science-corner of python. People there used notebook-like tools and workflows for decades and some brought that over to ipython-project. I remember 15(?) Years ago when the project started focusing more and more on the cluster-aspect of their shell, they brought up many differenct tools for this. One of them was notebook-like and what become later Jupyter. It quickly became popular in certain groups for those reasons.
Comparing R and Python also could be done better. Python is a general purpose language. R is for statistics and science. Even then, Python is extremely popular for these very domains with Jupyter notebooks.
This is why I shrug looking at "Jupyter killer" notebook alternatives tackling machine learning: in my experience delivering machine learning products to large, paying, clients, the bottleneck never was slicker stylesheets or cool animations. It was the nuts and bolts of things.
"make (something) accessible to everyone"
So the usage here is entirely consistent with standard English usage. It is also consistent with the French etymology (démocratiser), which has as a dictionary definition "Rendre démocratique, populaire" (i.e. to make popular).
dēmo- (people) -kratía (rule) has only indirect relation to popularity.
https://cdn.digg.com/images/e160ad4bb9c845f894155145539af3df...
https://medium.com/the-fourth-wave/the-great-democratization...
I don't know. How good did it feel to eat Freedom Fries after 9/11?
In this case, I agree it's not equitable.
The word 'democratize' here doesn't seem to add any meaning that 'literate programming' doesn't already cover.
Pretty good for an ide.
> we decided to assist fastai in their development of a new, literate programming environment for Python, called nbdev.
but this is followed by:
> nbdev builds on top of Jupyter notebooks to fill these gaps and provides the following features
is it a new environment, or is it an extended Jupyter Notebook? It looks like Jupyter Notebook to me. Why not Jupyter Lab?
> JupyterLab: Jupyter’s Next-Generation Notebook Interface https://jupyter.org
The earlier usage mentioned in Wikipedia is entirely uncited there however, and seems to have only been used in one academic project AFAICT.
Here's one from 1988 https://dl.acm.org/doi/abs/10.1145/51607.51614
Are these unrelated? Is Nbdev not only a "new programming environment", but also a new concept that needs a new name?
Kent Beck and Martin Fowler
http://index-of.es/Java/Planning%20Extreme%20Programming.pdf
Its a commonly understood term AFAIK
I agree - it's what made me raise an eyebrow... However, based on their comment above, the lead author of Nbdev believes to have coined the term.
My opinion is that due diligence and attribution are important. If I believed I'd coined a new term, I'd check first. Mistakes are easy to make, but when highlighted, perhaps corrections are more appropriate than negotiating with the person highlighting them:
From the lead author (jph00): ... If someone else can think of a better term that's never been used before, then I'll happily use that instead.
This is my second time on hacker news, and I don't think I will be back. Why not offer to help the project, or show support? Why try to find something to fight about? It's just demoralizing to see the lack of kindness from people.
It’s worth calling them out and discouraging this behavior because it leads to missing entire fields of previous work for both the author and people who build on the work.
And if the arguement is that this is a gateway action that leads to other bad things, that's not an arguement I put any stock in. It's used a lot in many places and many discussions, but just focus calling out the bad. It isn't his responsibility to ensure that some person building off his work at some point in the future does the appropriate level of research appropriate for their project. Assuming that an action is bad because you think someday down the road it might lead someone else to do something and that that something will be a bad thing is just ridiculous.
Whether they choose to update the materials and reference existing work is up to them.
But he is, he just claimed he coined “exploratory programming” FFS. It takes a shocking lack of hubris to announce that you are on the cutting edge of a field where you get to coin terms without doing the trivial amount of searching to verify it first.
- Nbdev started right around the 1.0 release of jupyterlab, and it might not have been on their radar
- Nbdev came out of fast.ai, and I wouldn't be surprised if they were using a ton of jupyter-specific features already which weren't supported by jupyterlab
(I'm the lead author of nbdev).
Neither? It doesn't change the features of jupyter notebooks, and its not an improved/expanded UI like jupyter lab (you could use nbdev with jupyter lab). Its utilities and automation to make package/library development a better experience if jupyter is where you write your code.
From https://github.com/fastai/nbdev:
"nbdev is a library that allows you to develop a python library in Jupyter Notebooks, putting all your code, tests and documentation in one place."
When a notebook gets large, it can be difficult to keep track of dependencies between cells. For workflows in which you have to run cells n_1, n_2, ..., n_k before running cell n.
I try to organize my cells so that if I run them from first to last, all dependencies are covered (e.g. "Restart kernel and run all cells).
Unfortunately, this doesn't help when I discover a bug in cell n_2 and don't want to run ALL cells n_2 + 1, ..., n-1, n because some of them carry out expensive operations.
When working in my editor, the way I resolve this is to make a light CLI wrapper around my program (if __name__ == "__main__": import argparse; ...) and my CLI commands encode all this dependency information.
Is it possible to get this kind of experience in a Jupyter notebook without building a custom plugin (I think a frontend plugin would suffice)?
Running all cells has its disadvantages, but what I found out was that often, there are bugs elsewhere than in cell n_2. I treat that as unit tests: run them all because even though I think only this is breaking, fixing it could have broken some other part.
Many other "Jupyter killers", as sensationalist blog post titles call them, claim they have done away with Jupyter's dirty "hidden state", but reading further about how the "Jupyter killer" avoids re-doing heavy computation by "caching" compute-intensive results and you want to tell them "get your mind right, which is it?".
The way we do it is schedule the notebook[0] to make sure everything works.
At first glance, this seems like it would address all of my pain points? Will be interesting to try it out.
- [0]: https://www.fast.ai/
- [1]: https://iko.ai
Here is a small project of mine highlighting some of the capabilities of nbdev: https://github.com/tmabraham/UPIT
Maybe we should say "ubiquitise" instead?
Literate programming involves having a meta language that is extended into the target source code through nested macros.
The killer feature was the ability to see everywhere a piece of code was used on dead paper by looking at the auto-generated index, with the chunks being logical rather than language driven. In a literate program you wouldn't care that something was a class or a function, you would just have it be described by what it does, not how it does it.
This is marginally better documentation for python notebooks.
It's more literate programming than not and chasing the promise of the ideal literate programming environment is what created the novel notebook environment in the first place.
2). Words have meanings and literate programming is defined extremely well by Knuth in his 1983 paper. This is not what was described there any more than the WWW is Xanadu.
3). That is not what No True Scotsman means.
Setting aside whether you are or are not right, can we just appreciate the irony in that statement for a moment?
The GitHub flavor is likely just because that is what the author was familiar with and what they were using.
If there are enough people interested we could get together to make PRs to add other remote version control systems and other static site hosts. I know an integration into the Atlassian world would really help me at work as that's my employer's chosen code repo and doc manager.