I built ReviewNb[1] to solve one of those problems (diff). Note that, there is nbdime[2] which works well for local diff/merge. The idea for ReviewNb is to have much tighter integration with GitHub etc.
I built ReviewNb[1] to solve one of those problems (diff). Note that, there is nbdime[2] which works well for local diff/merge. The idea for ReviewNb is to have much tighter integration with GitHub etc.
I wonder whether there is a solution along the lines of auto-committing each cell before it’s executed and the results just after the cell is executed. Otherwise a user has to do too much manual organizing, which is a problem the notebook should ideally solve. When a user is happy with the experiments and the provenance of their results, they should be able to use an interactive rebase to create a cleaner version to share/archive.
As a project moves from exploration toward production, the entire thing is wrapped into a Makefile that can flow from raw data to publication in a single call to make.
https://gist.github.com/ontouchstart/854a3c280b81f530d3ae9cb...
The notebook generated by nbconvert (see the instruction in the Makefile) is too big to display in GitHub gist Web UI but works fine in nbviewer.
https://nbviewer.jupyter.org/gist/ontouchstart/854a3c280b81f...
How do you manage encapsulating each step, and passing data between them?
https://github.com/4kbt/ReplicableAnalysis
and
https://github.com/4kbt/PlateWash
as examples. The former is smaller/less complicated. The latter was my thesis work -- more complicated and (unfortunately) abuses recursive calls to Make.
I think coupling to github makes sense if you are a building a dev-support service, but for a end user it makes little sense to wed the vcs to a specific website.
No inline rendering of markdown.
Opening an .Rmd file is a lottery to see if rendered graphs and tables still exists.
Tables render completely differently in editor, HTML and pdf
Your last point also has an upside - it's using different engines (Rmarkdown vs. Sweave). I can write whatever HTML or LaTeX code I want, depending on what's appropriate. I wouldn't want to have to make web documents with LaTeX, nor would I want to make PDFs with HTML.
That's incorrect, take a look here -> https://blog.rstudio.com/2016/10/05/r-notebooks
In json the code has to be escaped into strings, and json is really finicky about syntax (e.g. no trailing commas). So it doesn't work well.
I never got the chance to redo it, however the solution I was leaning to for my post "I won the lottery, I can work on fun stuff" attempt was to store the meta-code in a version of the host language(s), with some simple syntax that could live comfortably in the comments of various different languages to do things like encode the cell divisions and so on.
Basically something like: #notebook[lang=python]
#cell[lang=python] def add(x,y): return x + y #endcell
//notebook[lang=scala]
//cell[lang=python] def add(x: Int, y: Int) = x + y //endcell
This I think would be beneficial for a couple of reasons.
1. Better diffing / merging.
2. One click toggle between show source and view as notebook mode, which would really allow this to work in an IDE like vscode pretty seamlessly. The cells become something akin to //#regions in the IDE. But at the end of the day you are still editing a source code file, so you can edit the whole file easily.
3. The keyboard shortcuts for executing and jumping between cells would generally work in raw code mode, so you could just edit there continuously and manually writing out //cell //endcell. Also the executing results could appear in block comments inline in the editor, off to the side, or in a popup above, the code you are editing.
4. The IDEs could uprender the comment-syntax into cells as they gained better support for the paradigm (similar to how they do for code folding / syntax higlighting already).
5. Eventually, perhaps a cross language, metasyntax could be established to make things a bit more concrete than magic comments (get ready for some serious bikeshed painting though!)
The closest I have seen anything come in this regard is Quokka however it's not quite all the way there.
Your example could be rewritten as:
# -*- notebook-lang: python -*-
or // -*- notebook-lang: scala -*-
Still, the usual way of using Emacs for "interactive notebooks" is via org-mode, which is a better Markdown with support for (among other things) executing code blocks straight in the org document you're writing. This way, Emacs support all your points 1 to 5, and is generally more powerful than Jupyter or other similar things, but it also means you can kiss any kind of collaboration goodbye.For some weird reason, the more powerful a tool, the less likely it is other people will be using it.
--
[0] - https://www.gnu.org/software/emacs/manual/html_node/emacs/Sp...
For those who don't, less powerful tools take their place and proliferate.
As you say, emacs checks all the boxes but the majority is not prepared to learn it and prefers to program throught their browser.
mynotebook.py
### (cell boundary)
"""Top-level unused strings (docstring-esque) rendered as markdown"""
def add(x, y):
return x + y
# jupyter-output-hash: 0123abc (which would link to some external key-value storage for the project)
Anything in something other than the primary language could be in something like `execute_scala(""" scala code """)` - which would execute properly given proper globals.As long as the output-hash storage is treated as append-only and is highly available (output cells could even be encrypted for security if this was a public cloud service, or you could even use a local or shared filesystem), then this file would not only parse and run as a perfectly valid Python file, but it would also hold references to outputs in a source-control friendly way. IDEs could show the cell outputs inline. If you rerun your notebook and get different outputs for some reason, `git diff` tells you exactly where things changed without being too messy. Basically, put outputs in off-chain storage, and just be a literate code file.
I feel like this would address most people's needs, no?
There were tons of
## JOHN: DONT RUN PAST HERE, EVERYTHING BROKEN
comments.
https://medium.com/netflix-techblog/notebook-innovation-591e...
This approach seems promising, particularly as it facilitates cross-disciplinary collaboration.
I wonder how much of the "3 engineers for 1 data scientist" ratio I hear all the time is due to Data Engineering being assigned the role of cleanup to code that should be better in the first place.
I've seen cases where the wrong choice here ends up requiring three SDEs for half a year, where if they gave up a tiny benefit of the best model, they could have done it with 1 SDE in 1 month.
Version control is transparent and integrated and it's possible to work with workbooks collaboratively.
So you can import 2 JSON files and diff them in Powershell.
Edit: to clarify, jq -S does deep keys sorting.
$ echo '{"z":{"b": "second", "a": "first"}, "x": 4, "y": 7}' | jq -S
{
"x": 4,
"y": 7,
"z": {
"a": "first",
"b": "second"
}
}