Pluto.jl – a reactive, lightweight, simple notebook
github.com
github.com
* The fact that I can actually use the source files later because they're just Julia files is incredibly useful. I often copy-paste from them into actual REPL-code, and sometimes I just polish the notebook until its source becomes usable as a command-line tool.
* I like the reactive notebook concept. It does really help with bugs
* Pluto is still rough around the edges. Too few keyboard shortcuts. Buttons and text are tiny, afloat in an ocean of useless whitespace. pushing to LOAD_PATH doesn't work properly. Pluto is a very young project and just now gaining attention in the Julia community, so I'm confident these usability issues will improve.
Can you say more? An example, maybe?
(edit: I suppose it's not really global state I'm talking about as much as hidden state left over from overwritten or deleted cells)
For example, I typically create tonnes of code cells when I visualize and try to get a sense of some data. The large majority of cells (probably >80%) are then deleted, and whenever I find a trivial bug in some code, I fix the cell where the bug occured.
But now - which cells had I rerun after fixing the bug? And did any of the run cells depend on some variable in a deleted cell? If so, the notebook will no longer be reproducible when I re-run it? It's impossible to keep track of. So when I use Jupyter, I frequently press the "restart kernel and run all" option. But of course, that is slow. So I need to serialize a lot of data, which is troublesome. Pluto completely circumvents that problem.
a = 1
println(a)
a = 2
Does it show 1 or 2?Edit: tested it, it throws an error "Multiple definitions for a: Combine all definitions into a single reactive cell using a `begin ... end` block."
Not sure I like that way of working.
Every other cell that depends upon `a` will then automatically update.
A downside to using multiple cells is vertical spacing/visual noise. This is something that the package authors are currently thinking about addressing.
The primary complaint there is that notebooks have a disconnect between the state of the program and the display of the cells. By using a completely reactive mode the state is no longer hidden. It's more akin to a spreadsheet than a notebook. A number of other complaints are completely circumvented by using a file format that's simply a pure Julia file with clever comments.
Video: https://www.youtube.com/watch?v=7jiPeIFXb6U
Slides: https://docs.google.com/presentation/d/1n2RlMdmv1p25Xy5thJUh...
Previous discussion: https://news.ycombinator.com/item?id=17856700
I don't understand the point of the "reactive" cell order, instead of conventionally doing top to bottom. It seems like it goes against the idea of "the program state being completely described by the code you see".
And notebooks are mostly for exploration, and you don't really have a fixed order. You start getting the data, then you run the model, then you go back and change the data in the import... The order of your internal logic (or how you want to explain it to others) isn't necessarily linear, and of course you can just go back and write like a normal program but you lose the chain of changes that led you to the result (in this case the compromise is that your chain is restricted, like you said you can't define the same thing twice, but in exchange the notebook will provide with you the program order for free).
Also, I guess that you haven't used notebook environments like Jupyter before, so a bit of historical context might help. In Jupyter, cells aren't necessarily executed top-to-bottom, they are executed when the user asks it too. The result is then stored in the global state (well, assuming there is a global variable that the data is assigned to). This means that cells that depend on other cells also depend on the order in which those cells were executed. Worse still, if you write your notebook in a sloppy manner, you can end up with a state that you cannot reproduce from the still-remaining code (for example, you can have variables A and B, B is generated from the result of A, then you remove A. Because Jupyter is not reactive this does not update B, so your notebook keeps working just fine... until you decide to edit B). So previously, notebook-like environments made it really easy to introduce bugs like this.
With reactive cells you don't have to think of state. It's kind of like pure functional programming: it removes global side-effects. And note how on a technical level, making cells execute top-to-bottom is really just very a simple way to enforce that cells must executed in order of dependency! So either option insists that this global-state-that-does-not-respect-dependencies is a big problem that should be avoided, they just present different solutions for the problem.
This leads to a much nicer code organization, particularly when you're writing something resembling an interactive document (e.g. an "explorable explanation"). Much of the documents I do on ObservableHQ has roughly the following structure:
Title
Prose
Visualization
Interactive UI elements (possibly mixed
with further visualizations)
------------------------------
All the actual source code
powering the implementation,
organized in readability order.In my opinion, these kinds of apps are the future of data science / data analyst work. Forget no-code, just enable these professionals to work in a single programming language that they're familar with and give them visualization superpowers. The Python ecosystem has https://www.streamlit.io/ and https://gradio.app/ now. R has https://shiny.rstudio.com/. I think we'll see more.
Julia is a remarkable programming language, and pure Julia projects like this show that it is also good for general purpose development.
I want to try this combined with the Flux DL library.
Does it work well with large-datasets given its reactive nature?
2. git-friendly.
It also does a trick of creating new modules to manipulate scope to make deleted variables/import/cells invisible (and therefore free to be garbage collected).
With Pluto and other reactive notebooks, you have a guarantee that the code you see on the screen will produce the same results. So if you go back and edit cells out of order, save the notebook, then open it and re-run later, it will always be in the same state you left it in.
Even though notebooks are a common cause of headache in my world (I work on ML deployment), I think they're an incredibly valuable tool, and the familiar, visual interface of the browser plays a big part.
It clicked for me when I took a statistical genetics class taught by a team member of Hail.is (open source genomic analysis library). Coming from a dev background, I found working in a browser to be a clunky, awful experience—until I saw the way my classmates, most of whom were scientists by focus, used it. My instinct is to think of code in terms of the architecture of a program, but for them, code blocks were like buttons on a calculator. The speed at which they could iterate, and their ability to jump around, really drove home the value of the browser interface.
Would I want to write an API in one? Absolutely not. But for tinkering with genomic data? They're ideal in many ways.
Jeremy Howard of fast.ai talks about this a lot: https://twitter.com/jeremyphoward/status/1072555920029376512
Two examples from a previous work experience (remote sensing) :
(1) A colleague where creating SSH tunnel to create and explore data with the Jupyter process was on the calculation server. He was able to launch heavy calculation, fast-feedback loop for satellite images and shapefiles, manipulate the results and write the explanation next to each cell. As you would use a real notebook in fact. I had the same workflow and when I realized that the script will be used more than once, I moved the code to python script with command-line arguments support (just plug-in `argparse` to the script) and moved the text to comment the script.
(2) Teaching, we held seminar about different API and tools and used jupyter notebooks to teach everyone. The fast feedback loop was essential for anything with figures, plots, images, etc.
Pluto.jl while not yet perfect for me address a lot of broken things that made using Jupyter notebook driving me crazy (I had to broke my cells in a way that I can rerun everything when needed to update the global space, it was aweful).
One of our colleagues had around thirty students who had to prepare their final year pojects in machine learning. We deployed our internal platform and gave them access so they wouldn't lose an academic year, as they were in a zone that was hit really hard with COVID-19.
They are mostly on Windows, are not comfortable with the CLI, have never used Git, they don't do Docker, have poor connectivity [4kB/s], don't have access to powerful machines, found it difficult to handle dependencies, and needed to work on +200GB datasets. They also were split into groups of two or three, and needed to be able to share the work with each other, and with their supervisor (our colleague).
So, one reason to use a browser based solution is to de-couple the user's computer from the dependencies, internals, or infra the work happens on, and simplify collaboration. This, or the main tool you rely on does not play nice with other tools, even if you're proficient.
We started to push for remote work in late 2018, and we started really going after it in 2019 because commute was draining our colleagues' energy. It really bothered us to see them arrive at work completely washed out, or see them worry about transportation at the end of the day, so we made remote work a priority. But they mainly trained models with notebooks, and there was a need to be able to do actual work as a team, so we built the tooling around our workflow and we've had to add in missing features to accomodate our colleagues who needed the notebook.
Why are results displayed above the code cell instead of below it (as in Jupyter/IPython)? Was a bit confusing at first.
I'm noticing it crashes a lot. Get messages like:
Worker 2 terminated.
Distributed.ProcessExitedException(2)
Really like the reactive aspect, though.I also strongly agreed with this at first, but after working with it some more I've found it compelling. The key mental model is to think of the code as something akin to a "figure caption."
There is also a great deal of overlap between IDEs and Notebook platforms.
[1] https://github.com/JuliaLang/IJulia.jl
Edit: "pure Julia" => "Julia centric" based on jakobnissen's comment
The main advantages of Pluto is that
* The sources files are executable Julia files with minimal metadata, so it plays nice with Git. Also, the code of the source files is ordered to reflect the execution order of the cells, to keep the source code and the notebook in sync.
* It attempts to remove all global state. If you change a cell, and dependent cells will change as well (similar to Excel). This makes bugs less likely.
From what I understand, results of cells are cached though, and they don't update unless something upstream changes, so there is a form of memoization happening. Which of course also has both implications for performance as well as memory usage.
Personally, I will never code in a web browser and don't understand how people can do so with large code bases.