Why Jupyter is data scientists’ computational notebook of choice
nature.com
nature.com
a) Just like with a paper, you can present scientific or mathematical ideas with accompanying visualizations or simulations. From the REPL side, as a bonus, you get interactivity, and the reader can pause and experiment with the examples you're giving to improve their understanding or test their hypotheses. If I change this variable, how will the system react? You can just try it!
b) Just like with a REPL, you can type in and execute commands step by step, viewing the output of the previous command instead of running the whole thing at once. From the document side, as a bonus, you get nicer presentation (charts, interactivity, nice and wide sortable tables, etc) than you would in a shell, which comes in handy when doing things like data exploration or mathematical simulation.
It's decidedly NOT there for you to type all your code in like an editor and make a huge mess. It's apples and oranges w.r.t and a poor substitute for something like PyCharm or VS Code or vim. It is there for you to a) try things out yourself, and whatever you discover hopefully eventually make it into proper python modules b) make interesting ideas presentable and explorable for others. That's all!
When I see stuff like "out of order execution is confusing", I don't disagree, but it does make me wonder how long and convoluted the notebooks these people work with are - probably a ripe candidate to refactor stuff out into python modules as functions. When I see stuff around notebooks for "reproducibility", I'm a bit confused in that notebooks often don't specify any guidance on installation and dependencies, let alone things like arguments and options that a regular old script would. In that regard I think it's barely an improvement over .py files lying around. When I hear "how do I import a notebook like a python module", I'm very very scared.
Granted, I've seen huge notebooks that are a mess, so I understand the frustration, but it's not like we all haven't seen the single file of code with 5000 lines and 10 nested layers of conditionals at some point in our lives.
The irony, to me, is that I actually typically argue for the mixing of presentation and content. But to me, notebooks look like an attempt by people to make a WYSIWYG out of JUnit/TestNG/whatever style reports. Only, without the repeatability.
There is also the entire bend where these are taking off in a way that doesn't make sense. Do they do the things you are saying? Well, yeah. But no better than plenty of tools before them. Mathematica and Matlab both had "notebook" like features for a long long time. Complete with optimized libraries. And this is ignoring the interactivity of the old LISP machines. (You can see from my history I have a soft spot for emacs org-mode.)
Jupyter is a lot of things. Bad isn't necessarily one of them, but exceptional isn't, either. Heavily marketed is.
The exact same thing happened with the arrival of the www.
They probably didn't take off to the same extent as Jupyter because they're not free. IIRC MATLAB was quite expensive, particularly if you wanted to do anything specialised.
Notebooks have been in Mathematica for ages and are really powerful and difficult to describe to those who haven't used them. To give an example, I was building a tool and embedded images as variables in a way reminiscent of being an engineer on the USS Enterprise. You can point to a file in Python as a variable, but you can't just copy-paste an image in as a variable last I checked (don't think Jupyter is there yet).
It makes perfect sense. Just not to a lot of HN readers.
The average HN reader is approaching this from a perspective of "I am a professional programmer who might occasionally dabble in scientific computing, and therefore I hate this thing because it's not a professional programmer's tool designed by and for professional programmers according to the best practices of professional programmers".
The people who are actually using notebooks, meanwhile, are not professional programmers. They're scientists who increasingly have to do programming as part of their science. And notebooks are a godsend for them. We don't need to drag them all the way into our world; we need to pay attention to what they actually want, need, and find useful, and accept that it's going to differ from what we want, need, and find useful.
2/3 of scientific research cannot be reproduced by other scientists. But tell us more about why scientists should ignore best practices from other fields.
Okay, that was extreme, and if you think I was talking about programming, it's because you have a guilty conscience. ;-) It actually applies to all interesting fields -- programming, engineering, management, classical music composition, etc. Those fields don't even know what their best practices are, and acknowledge that things take too long and can't be managed. No manager would say: "Our programmers have best practices, so the work will be done next week." Why should scientists have such faith?
Meanwhile, do you trust Maxwell's Equations, Darwinian evolution, quantum mechanics, etc.? How did we establish the physical constants to mostly better than 8 digits of precision? Science has somehow figured out how to make progress despite the messy business of research.
For me, it's not that I "have" to do programming, but that physical science has been computation driven since before the 1940s. Programming is how I think and work. With apologies to Richelieu, "programming is too important to be left to the programmers."
"Best practices" are a chimera. The issue at hand isn't about what is "best", but whether or not a software engineer's "good enough" practices are more likely to achieve science's goals than a graduate student's "good enough" practices.
It's also disingenuous to claim that classical music composition doesn't have "best practices" when the field of music theory exists as an explicit manifestation of "best practices" in music. Having gone to a school with a conservatory, I also believe that I know several individuals who would would disagree with your mindset regarding how the creative process can't be managed. Indeed, if creativity, as it relates to musical composition, couldn't be managed most orchestras would be brimming with anger at the number of commissions that weren't finished on time for the concert, and most Hollywood studios and Broadway shows would screech to a halt.
As for being a a good REPL, I feel that an actual REPL (+ editor integration) works better than notebooks: You can combine a literate document with a REPL but still get the benefits of a proper editor/IDE and a proper execution environment, rather than a half-hearted mix of both that’s hosted inside a HTML contenteditable (= Jupyter), and you also get “charts, interactivity, nice and wide sortable tables, etc” if you want). RMarkdown inside RStudio or Nvim-R does this well. — I just don’t want to give up the advantages of a proper editor for the very slight increase in integration that Jupyter gives me.
I myself, as someone who likes to create really solid and maintainable tools, have fallen into the notebook trap and written things like "change the month in cell 22 then execute cells 1 through 3 and 20 through 27 to update the report".
The notebook format was great for prototyping what was really a small app. You don't really have those problems when you're just generating a document.
To really get something great out of excel you have to learn excel. I think that difference is almost as important as the excel stigma.
My gut is that the amount of data it can work with more than compensates for the disadvantage.
R and Python are supported. No Julia, unfortunately. VS Code and Atom support similar workflows with Julia. However, the Julia Language server in VS Code is extremely unstable and I regularly lose LaTeX completions. The REPL in Atom is mind boggling laggy and slow to the point that it is much less frustrating to copy and paste code into a REPL running in your favorite terminal emulator.
However, it doesn't seem to complete variable names. Eg, if I defined "foobar", it won't let me tab-complete that later when I start typing "foo".
I just looked through Julia's settings tab in Atom, and saw the option "Fallback Renderer" with the note "Enable this if you're experiencing slowdowns in the built-in terminals." It was disabled by default, so I've just enabled it.
Subjectively, I think it feels fine now. Longer use will tell, but I suspect I was just running into a known issue some setups run into, and they already provided the workaround.
EDIT: Comparing running some code in Atom's terminal and a REPL running in the GNOME Terminal, the regular REPL still feels notably snappier -- even though I'm `using OhMyREPL`, which makes the REPL a bit less responsive.
I'd say Atom feels acceptable (and definitely not "mind boggling laggy" right now), and shift/ctrl + enter more convenient than switching tabs. So I will stick with it (for Julia). More time shall tell.
With the reticulate package in R Markdown you can run python chunks, by putting, e.g.
```{python}
for i in range(1:10):
print("{}:{}".format(i, i*i))
# etc```
And in emacs org-mode you can:
#+begin_src python
for i in range(1:10):
print("{}:{}".format(i, i*i))
#+end_srcLanguage support in org-mode is pretty comprehensive, afaik.
I do not know the details of the implementations behind these, but my own source code is plain and simple unserialized text, and that means a lot to me.
By that I mean that that to share the r-markdown doc it appears that you need to rerun the whole thing. It does some tricks to do concurrent visualization, but to actually share the doc you have to rerun all the R/python from scratch.
In jupyter OTOH, if I have a long running ML pipeline as part of my doc, I can render without rerunning the pipeline.
Emacs org-mode is proof that a simple text format with markup rules is all you really need to support multiple languages in a single file. You lose some of the simplicity of parsing the file, but you gain a ton more.
Org-mode is great, but you still have to install emacs.
On point, a browser can not remember ifa notebook. Just parse the json. It can also parse text/plain. So, could show the org document without styling. The org document is actually readable. Json... Not so much.
To see the notebook, you have to have a Jupyter setup somewhere.
Edit: For example, see https://raw.githubusercontent.com/taeric/taeric.github.io/ma... which is the source for http://taeric.github.io/ChangeForDollar.html Not styled, and that is a short document, so probably woudn't be that tough to read in a json document, but I'm glad I don't have to.
Multi-language parsing is much harder to solve than simply enforcing some escaping mechanism in the inner protocol level and having tools do the "heavy" lifting (basically a solved problem).
That said, you can define away a large part of the problem.
Edit: For trivial examples of "org-mode" in an org-mode document, you need only look at the documentation of org-mode. That said, I expect there to be limitations, because they make sense. Similar to how you can pretty print json inside a jupyter notebook, but don't expect to have a notebook interpreted in the notebook. (If that makes sense.)
It's also trivial to export notebooks to .py files.
That said, my goodness do notebooks wreak havoc on git. I hope this in particular gets fixed as popularity grows.
But this is useless if you cannot edit those py files and obtain notebooks from them.
Actually, fixing the fileformat-mess should be very simple. Just change the file-load/save-functions. Use a folder-structure with every cell being a seperate file. Or switch to XML. Or make a generic interface and allow to save in whatever the people want. Saving notebooks in Mongodb or some SQL-Database seems like a good goal for dedicated services.
https://rviews.rstudio.com/2017/03/15/why-i-love-r-notebooks...
This is the reason why it is great.
> 1) Plain text representation
> In the case of an analyst, the domain of "software engineering" lies close to their own domain. Projects in both areas require code which (ideally) exhibits clarity and reproducibility. Obfuscated software is bad [...] and idempotency is good.
> The problem, then, is when the analyst takes a core tool from their domain and applies it to a slightly different domain like software engineering. Things go south fast: your notebook has not-quite-imperative code that is untested and unmonitored. It is, in other words, bad software.
As for the point about "refactoring stuff out into python modules as functions," the problem is that the new crop of data scientists aren't learning how to do this. The role of "machine learning engineer" is emerging to address this shortcoming in SWE skill throughout the data science community. It honestly cannot happen quickly enough.
[1] https://buttondown.email/oneshotlearning/archive/c06a0ded-74...
At the core of this, as some others may have already alluded to already, is that many academic scientists have not been socialized to make a distinction between development and production environments. Jupyter notebooks are clearly beneficial for sandboxing and trying out analyses creatively (with many wrong turns) before running "production" analyses, which ideally should be the ones that are reproducible. For many scientific papers, the analysis stops at "I was messing around in SPSS and MATLAB at 3 AM and got this result" without much consideration for reformulating what the researcher did and rewriting code/scripts so that they can be re-run consistently.
Geologist here - definitely true in my field. Nonetheless, while I don't develop in notebooks at all, I do use them for "reproducibility" in a sense -- by putting a bit of dependency info in a github repo along with a .ipynb file, I can do things like this: https://mybinder.org/v2/gh/brenhinkeller/Chron.jl/master?fil...
Which ends up being useful when a lot of folks in my field don't do any computational work at all, so being able to just click on a link and have something work in browser is a big help.
The image after: "For example (KJ04-70)"
(I also re-ran the preceding cells).
<img src="DenverUPbExampleData/KJ04-70_distribution.pdf" align="left"/>
works in Safari but not Chrome or Firefox. Switched to SVGs for now.
So much for reproducibility :/
In fact, there might be something about what attracts people to be scientists rather than engineers, that makes us bristle at doing what engineers consider to be "good" engineering.
Maybe your awful notebook gets the same answer you got the day before on the blackboard. Or the same answer your collaborator got independently, perhaps with different tools. Those might be great checks that you understand what you're doing. Spending time on them might be more valuable for finding errors than spending time on making one approach run without human intervention.
Not to say that there aren't some scientists who would benefit from better engineering. But it's too strong to say that fixing everything that looks wrong to engineer's eyes is automatically a good idea.
For my work, reproducing a result may involve collecting more data, because a notebook might be a piece of a bigger puzzle that includes hardware and physical data. This is where scripting is a two edged sword. On the one hand, it's easy to get sloppy in all of the ways that horrify real programmers. On the other hand, scripting an experiment so it runs with little manual intervention means that you can run it several times.
What I find objectionable is the inability of scientists to explicitly delegate tasks to domain specialists in their everyday work when it makes sense. I think that it's unrealistic of you to believe that engineers always work with "a fully dimensioned and toleranced drawing" before starting work on a project and that your would work "grind to a halt". Indeed, there's a reason for the qualifier rapid in the term "rapid prototyping". If you can give an engineer general specifications for what you want and then leave him/her alone, he/she should be able to produce something that mostly fits your needs while avoiding all of the pitfalls that wouldn't have occurred to you. It would also be incorrect to assume that engineering does not involve creativity and is purely bound by rigid processes- if your requirements were strange enough, something fresh would inevitably be built.
This sort of delegation of course, is actually more efficient, since you can work on other tasks in parallel with the engineer (such as writing your next grant proposal or article or gasp teaching). Most scientists also already do this implicitly by choosing to purchase instrumentation from manufacturers like Olympus, Phillips, or Siemens rather than building it themselves.
Part of the reason for why I have such strong opinions about this matter, is that I've actually witnessed scientists waste more time messing around in fields where they were clearly out of their depth. As an example, there was a thread on a listserv in my (former) field that lasted for literally months that was solely devoted to the appearance of a website. Everyone wanted to turn the website design into an academic debate, when the website's creation (which had little to do with the substance of the scholarship itself) could have been turned over to a seasoned web developer and finished in less than a week or two.
With that said, Jupyter has greatly improved my ability to find my own mistakes, and to reproduce my own results later on.
No, the majority of complaints are that notebooks are great, but Jupyter is a bad notebook. I mean maybe it’s impressive to someone who’s never seen a notebook before but to someone used to Mathematica, MathCAD, RMarkdown, org-mode, whatever, it just seems clunky as hell. I wonder how many “data scientists” claiming it as their top choice have ever tried anything else?
I built ReviewNb[1] to solve one of those problems (diff). Note that, there is nbdime[2] which works well for local diff/merge. The idea for ReviewNb is to have much tighter integration with GitHub etc.
So you can import 2 JSON files and diff them in Powershell.
Edit: to clarify, jq -S does deep keys sorting.
$ echo '{"z":{"b": "second", "a": "first"}, "x": 4, "y": 7}' | jq -S
{
"x": 4,
"y": 7,
"z": {
"a": "first",
"b": "second"
}
}I think coupling to github makes sense if you are a building a dev-support service, but for a end user it makes little sense to wed the vcs to a specific website.
https://medium.com/netflix-techblog/notebook-innovation-591e...
This approach seems promising, particularly as it facilitates cross-disciplinary collaboration.
I wonder how much of the "3 engineers for 1 data scientist" ratio I hear all the time is due to Data Engineering being assigned the role of cleanup to code that should be better in the first place.
I've seen cases where the wrong choice here ends up requiring three SDEs for half a year, where if they gave up a tiny benefit of the best model, they could have done it with 1 SDE in 1 month.
I wonder whether there is a solution along the lines of auto-committing each cell before it’s executed and the results just after the cell is executed. Otherwise a user has to do too much manual organizing, which is a problem the notebook should ideally solve. When a user is happy with the experiments and the provenance of their results, they should be able to use an interactive rebase to create a cleaner version to share/archive.
As a project moves from exploration toward production, the entire thing is wrapped into a Makefile that can flow from raw data to publication in a single call to make.
https://gist.github.com/ontouchstart/854a3c280b81f530d3ae9cb...
The notebook generated by nbconvert (see the instruction in the Makefile) is too big to display in GitHub gist Web UI but works fine in nbviewer.
https://nbviewer.jupyter.org/gist/ontouchstart/854a3c280b81f...
How do you manage encapsulating each step, and passing data between them?
https://github.com/4kbt/ReplicableAnalysis
and
https://github.com/4kbt/PlateWash
as examples. The former is smaller/less complicated. The latter was my thesis work -- more complicated and (unfortunately) abuses recursive calls to Make.
No inline rendering of markdown.
Opening an .Rmd file is a lottery to see if rendered graphs and tables still exists.
Tables render completely differently in editor, HTML and pdf
Your last point also has an upside - it's using different engines (Rmarkdown vs. Sweave). I can write whatever HTML or LaTeX code I want, depending on what's appropriate. I wouldn't want to have to make web documents with LaTeX, nor would I want to make PDFs with HTML.
That's incorrect, take a look here -> https://blog.rstudio.com/2016/10/05/r-notebooks
Version control is transparent and integrated and it's possible to work with workbooks collaboratively.
In json the code has to be escaped into strings, and json is really finicky about syntax (e.g. no trailing commas). So it doesn't work well.
I never got the chance to redo it, however the solution I was leaning to for my post "I won the lottery, I can work on fun stuff" attempt was to store the meta-code in a version of the host language(s), with some simple syntax that could live comfortably in the comments of various different languages to do things like encode the cell divisions and so on.
Basically something like: #notebook[lang=python]
#cell[lang=python] def add(x,y): return x + y #endcell
//notebook[lang=scala]
//cell[lang=python] def add(x: Int, y: Int) = x + y //endcell
This I think would be beneficial for a couple of reasons.
1. Better diffing / merging.
2. One click toggle between show source and view as notebook mode, which would really allow this to work in an IDE like vscode pretty seamlessly. The cells become something akin to //#regions in the IDE. But at the end of the day you are still editing a source code file, so you can edit the whole file easily.
3. The keyboard shortcuts for executing and jumping between cells would generally work in raw code mode, so you could just edit there continuously and manually writing out //cell //endcell. Also the executing results could appear in block comments inline in the editor, off to the side, or in a popup above, the code you are editing.
4. The IDEs could uprender the comment-syntax into cells as they gained better support for the paradigm (similar to how they do for code folding / syntax higlighting already).
5. Eventually, perhaps a cross language, metasyntax could be established to make things a bit more concrete than magic comments (get ready for some serious bikeshed painting though!)
The closest I have seen anything come in this regard is Quokka however it's not quite all the way there.
Your example could be rewritten as:
# -*- notebook-lang: python -*-
or // -*- notebook-lang: scala -*-
Still, the usual way of using Emacs for "interactive notebooks" is via org-mode, which is a better Markdown with support for (among other things) executing code blocks straight in the org document you're writing. This way, Emacs support all your points 1 to 5, and is generally more powerful than Jupyter or other similar things, but it also means you can kiss any kind of collaboration goodbye.For some weird reason, the more powerful a tool, the less likely it is other people will be using it.
--
[0] - https://www.gnu.org/software/emacs/manual/html_node/emacs/Sp...
For those who don't, less powerful tools take their place and proliferate.
As you say, emacs checks all the boxes but the majority is not prepared to learn it and prefers to program throught their browser.
mynotebook.py
### (cell boundary)
"""Top-level unused strings (docstring-esque) rendered as markdown"""
def add(x, y):
return x + y
# jupyter-output-hash: 0123abc (which would link to some external key-value storage for the project)
Anything in something other than the primary language could be in something like `execute_scala(""" scala code """)` - which would execute properly given proper globals.As long as the output-hash storage is treated as append-only and is highly available (output cells could even be encrypted for security if this was a public cloud service, or you could even use a local or shared filesystem), then this file would not only parse and run as a perfectly valid Python file, but it would also hold references to outputs in a source-control friendly way. IDEs could show the cell outputs inline. If you rerun your notebook and get different outputs for some reason, `git diff` tells you exactly where things changed without being too messy. Basically, put outputs in off-chain storage, and just be a literate code file.
I feel like this would address most people's needs, no?
There were tons of
## JOHN: DONT RUN PAST HERE, EVERYTHING BROKEN
comments.
From my perspective, it's a dumpster fire - in 2018 there should be something so much better than this. RStudio is a thousand times better but only does R. I used to like Beaker Notebook but it gave up due to Jupyter's popularity and converted itself into a bunch of Jupyter extensions which now have all of Jupyter's limitations.
Yet despite all this I can see that there's this enormous community that loves this and keeps developing and contributing to it.
In 1998 I was using a tool called MathCAD that provided a notebook interface running as a plugin to MS Word. In 2018, Jupyter is still not as good as that. Some things are just not meant to be webpages.
This is how I feel about most of the single page apps I've worked on.
And all JupyterCon 2018 talks if anyone is interested: https://www.youtube.com/playlist?list=PL055Epbe6d5b572IRmYAH...
At least Atom has good integration with hydrogen
I use R for 90% of my work, but most of it has been happening in Jupyter notebooks (which I'm not a huge fan of, despite practically living in them for the past 4 years of my life).
Thanks for sharing!
Scimax version: https://github.com/jkitchin/scimax/blob/master/scimax-ipytho...
Video of Scimax version: https://www.youtube.com/watch?v=dMira3QsUdg
Previous HN discussion highlighting key features: https://news.ycombinator.com/item?id=17839926
Some relevant blog posts:
- https://vxlabs.com/2017/11/24/getting-ob-ipython-to-show-doc...
- https://vxlabs.com/2017/11/30/run-code-on-remote-ipython-ker...
- https://kozikow.com/2016/05/21/very-powerful-data-analysis-e...
1. variables have to be explicitly output
The most important tool for programming, for me, is that window that shows you the current state of all the variables. When I step through a program, I look at the state. 90% of my debugging solutions come from seeing that variable doesn't have the right state.
2. Intellisense
For the love of god, I do not want to remember if it is len(), length(), .len(), .length(), .size(), size(1) or whatever.
That's it. But those two are so big that I have to code and debug in Spyder and then paste the code into notebook. I feel sorry for people who are new who think that all the debugging is happening in the notebook.
Also planning to add some nice debug features, plus hopefully integration into the inbuilt VSCode debugger!
Personally, I like VSCode more than Atom, so this was one of the reasons I started working on this extension!
I want to combine, but Neuron has stated they will not be accepting any PRs until December.
It offers some nifty things including, well, observables where cells of the scratchpad can automatically update by observing changes from other cells.
The real question is when (and whether) new social scientist stats courses will start teaching Python stats toolchains, rather than R. That seemed to be an inflection point for R (as folks moved away from SAS), and could be for stats-centric Python too.
I do like the SAS dev tools (especially Enterprise Guide). I'd really love if R had some sort of GUI front end for non technical people. Eg. I know the finance analysts in our org wouldn't have a clue how to configure their own ODBC sources -which you need to do with R Studio until it's as easy for them as SAS convincing them to switch won't get any traction.
Python definitely has more mindshare for machine learning, and particularly deep learning. However, that’s not all of statistics. For things like mixed-effects modeling, I think R still has a clear lead. There are some python packages (e.g., statsmodels) but R’s lme4 has more features, like custom covariance structures, and virtually every textbook and tutorial currently uses R. I’m actually not sure if I’ve ever actually encountered statsmodels in the wild. PyMCMC is relatively popular, but I think bugs/jags are also more common.
Don’t get me wrong, I like Python well enough, and knew it before I coded R. But Python is really behind R in stats support. I’d also add the tidyverse in there for general data munging.
If I want libraries I’ll use R; if I want a programming language I love I'll use Racket or maybe Clojure; if I want some libraries and an okay programming language I’ll use Python, I guess.
I don't expect either R or Python to go away either time soon, nor would I want them to, but I would like to see people moving to things like Julia and Nim, which have the same level of expressivity, but are much more performant. I have difficulty imagining many people saying "I love programming in R and Python, but don't like Julia or Nim."
I like Python but at least with stats/numerics there isn't a big reason to move away from R except for specific libraries (especially DL stuff) or front-end integration with web-land (and even then things like Jupyter mitigate against that).
In theory, there are Python and Julia equivalents to RStudio (JupyterLab, Spyder, PyCharm, Juno, whatever) but RStudio is just so, so, so good. A truly great piece of software.
And of course if you have a data pipeline type workflow, and it fits into the Hadleyverse paradigm and isn't too performance intensive, there's nothing better.
It's still in development (mutate and select have PRs) but it's almost there.
> and as more people realize that there's a benefit to simultaneously training researchers to run code as well as stats, I suspect we'll start to see an exodus from pure R solutions
I'm not sure what you are saying here.
I would actually argue that most of the Python data science toolchain is years behind what is available in R.
I do not want to litigate this on HN, but the problem with R is the toolchain around your data science work.
You've fit a model in R, and that's great! Now how do you get it into a real-time system? Or how do you test the software you wrote to train the model?
Not for everybody, e.g., the Swirl R package (https://swirlstats.com/) doesn't work in Jupyter, since Jupyter has limited support for R's many ways of getting interactive input from users.
Being interactive is what makes it good for teaching. But there are plenty of interactive options. And for a certain class of teaching, it is not "on rails" enough such that people will have to have a ramp up period first on Jupyter before they can really get into their topic.
How do people cope with this? Do you supplement it with other tools? I spend a lot of my time in an IDE and then just paste some of the code in to cells. That seems easier.
putting models (in a sense of more complicated models, not just a SVM), data-pipelines, shared visualization-code in a src folders and experimenting in the notebook divides stuff that's interactive by nature from "real" coding. I don't context-switch that much to be honest.
I don't really copy code into cells, because I only experiment there.
Also, what happens if you need to share code between notebooks?
I think notebooks should be simple and explain the experiments and the reasoning behind them to your coworkers. Otherise it's hard to coordinate and learn from each others insights into the data.
If you're writing a lot of code in them, it's probably better to put that code into libraries that get imported and reused.
And I do agree that default code environment is unbearable. Particularly the auto insertion of completing quotation marks, which has me continually fighting with the editor to get correct code into a tiny web text box.
What I'm specifically talking about is even that kinda hacky experiment code you end up writing. I don't try to implement whole projects in there, but even just "train this model" type code ends up being a hassle because of how bad the editors are.
My above comment was more referencing wishing I could spend more time writing experiment code in jupyter without copying and pasting all the time.
Work (and often debug) in jupyter -> open the notebook from pycharm when it's got some completed thoughts and write into a python module + test module, tidying up and adding type annotations.
Sometimes doing that multiple times so that the notebook is importing from modules which were originally pulled out of the notebook.
It sucks having to use two tools but I don't think there's any one tool that can do both as well as pycharm/jupyter, short of me getting a lot better at emacs or writing a lot of custom Atom extensions (I think).
(Relevant issue: https://github.com/jupyterlab/jupyterlab/issues/2163)
P.S. Disclaimer: I lead this project at JetBrains, Inc.
Is it possible to use it as what seems like a drop-in replacement for jupyter notebooks?
We have more data then I think would make sense to transfer out of our clusters/datacenter and privacy issues would probably be raised but I would love to use something like this.
We are seriously considering on premises version.
>Is it possible to use it as what seems like a drop-in replacement for jupyter notebooks?
Jupyter import/export will be released soon.
Folks that try to do all programming in notebooks typically drown in complexity and suffer.
Maybe due to often importing and naming (something you don't do in Java.)
E.g
Import matplotlib as plot
Vs
Import java.util.This is a trade-off between how much code you're writing and how much data you're processing. If you're writing maybe 20 lines of code but you have enough input that it takes several minutes to run, the notebook becomes a clear win for your development process.
I also notice that developing in this way encourages me to create smaller, more testable functions that i can easily work with inside a single notebook cell.
One thing I haven't figure out how to do is to generate fully styled LaTeX manuscripts from notebooks (like papaja for RStudio). Is there a way to do this with pandoc?
However, the output of this isn’t nearly as well formed as a hand written latex document is.
My brother in law wants something like this for structural analysis reports where the code, data and report are all one thing that can be pulled out and examined.
Watching the new iPad announcement today I think this is something that would make an excellent iPad app as well.
i'm pretty sure this is a usecase which others have and curious what people use to solve it. i've heard, variously, that some options are to use tableau/similar or email a csv and ask the end user to import into google sheets/excel
It's easy to share data between the R and python session, and calling python from R, or R from python is straightforward.
And it's meant to work with the RStudio IDE, so I get a much more seamless experience going between regular code and notebooks (although this is admittedly a more R-centric benefit, at least until and unless RStudio adds Python support outside of notebooks).
Frankly I find that all programming environments for scientific computing are deficient in some way or another. If you look at the set of features in Visual Studio, R Studio and Jupyter notebooks, you will see that the Union of useful features is large, and the intersection is almost empty.
The results of a notebook can be shared more easily than a plain repository(via nbviwer or binder) and more importantly the science there it's reproducible.
Obviously a great part of why we use Lightroom/darktable is because of the speed with which the recipe-processing occurs. Plus a smooth UI, a catalog-viewing feature, and a well vetted choice of image operations. The appeal of moving this work to a notebook would be that an actively maintained Jupyter ecosystem could supplant lock-in to a specific software, and open up the underlying math magic.
At the very least, this could be an interesting platform for experimenting with image processing methods. And the reordering of cells could become a virtue, to run an image processing pipeline out of the standard order.
I'm curious if anyone has already worked along these lines. I find through a quick web search that people are doing some image processing, but more in the face detection or ML for medical imaging aspects. I see as a basic toolkit that http://scikit-image.org/docs/dev/auto_examples/ is something, though this isn't the whole range of operations needed for, say, fine art image tuning.
Still, the fact remains that JuptyerHub is powerful, but difficult to install and manage if you're not a university IT dept. Any SMB solutions?
If so, I’d run the notebook on the remote server and just teach them whatever command they need to make a ssh tunnel there. Something like what is described here: https://techtalktone.wordpress.com/2017/03/28/running-jupyte...
So they would utter the unknowable incantation and then point their browser at localhost:8000 or whatever and then use their version of the notebook.
Disclaimer, I work on Polyaxon.
Sharing articles with team members and letting them run them is trivial. We automatically version the article, the data and the environment (docker image) and you can remix (fork) other articles. `xoxo` is a signup code you can you if you want to give it a try.
We're not far out (~two weeks) from launching our beta for private research. Here, you'll get your own private data store and docker registry as well as secrets management (stored securely in hashicorps vault).
Notebooks aren't ideal for creating functions (standard text editor features are lacking and testing is impossible).
Notebooks encourage an "order dependent variable assignment" programming style without abstractions. Here's what you'll commonly see in a notebook:
val df = spark.read.csv("some_data")
df2 = df.withColumn("clean_name", trim("name"))
df3 = df2.filter("clean_name" === "Mark")
I've found that notebooks are very useful if you write all the complicated code in separate GitHub repos and attach binary executables to the cluster. If you try to write all your logic in notebooks, you'll quickly struggle with order dependent, messy code.
[1] Github: https://github.com/janushendersonassetallocation/loman [2] Quickstart/Docs: https://loman.readthedocs.io/en/latest/user/quickstart.html
https://docs.google.com/presentation/d/1n2RlMdmv1p25Xy5thJUh...
https://www.youtube.com/watch?v=7jiPeIFXb6U&feature=youtu.be
“In many cases, it’s much easier to move the computer to the data than the data to the computer,” says Pérez of Jupyter’s cloud-based capabilities. “What this architecture helps to do is to say, you tell me where your data is, and I’ll give you a computer right there.”
Installing R packages through anaconda is like pulling teeth and the docker images for my Jupyter notebooks push past 6GB and take multiple cups of tea to build.
Is there a good solution I'm missing? A good hosted solution perhaps?
The Docker image is 6GB.
We archive full-stack reproducibility by allowing you to install arbitrary software and version these environments using docker. You can reuse these in other articles or pull and use them locally. `xoxo` is a signup code you can you if you want to give it a try.
In addition to our own runtime protocol, we support Jupyter kernels and you can import Jupyter and (R)markdown documents.
Also: https://www.theatlantic.com/science/archive/2018/04/the-scie...
“The notebook interface was the brainchild of Theodore Gray, who was inspired while working with an old Apple code editor. Where most programming environments either had you run code one line at a time, or all at once as a big blob, the Apple editor let you highlight any part of your code and run just that part. Gray brought the same basic concept to Mathematica, with help refining the design from none other than Steve Jobs.”
[1] https://rmarkdown.rstudio.com/ [2] https://orgmode.org/worg/org-contrib/babel/
I do my task management and note-taking in Org-mode, and recently I found myself doing things like jotting in the middle of my notes[0]:
#+BEGIN_SRC http
GET address.to.api:123/sth
#+END_SRC
and tapping CTRL+C twice, to get the actual response of the API I was debugging.Or, the other day I was making notes about gravity batteries, and was wondering how efficient is one startup's solution. I briefly thought about firing up Jupyter, but then simply wrote the following[1]:
these guys power a LED (or three?) with a 0.1W, generated through dropping
a 12kg weight down 1.8 meters over 20 minutes.
Doing some basic math on that:
#+BEGIN_SRC elisp
(let* ((m 12)
(g 9.81)
(h 1.8)
(_t (* 20 60))
(E (* m g h)) ; E = m*g*h
(P (/ E _t)) ; P = E/t
(efficiency (/ 0.1 P)) ; efficiency = Pout/Pin
)
`("ideal power [W]" ,P
"efficiency [1]" ,efficiency))
#+END_SRC
Typing CTRL+C twice, out pops: #+RESULTS:
| ideal power [W] | 0.17658000000000001 | efficiency [1] | 0.5663155510250312 |
(which is automatically rendered as an org-mode table I can operate on, or even reference in other code snippets).Point being, note-taking in org mode makes it ridiculously easy to invoke any programming language you hooked up to Emacs without breaking your flow, and you get to edit the code in the mode specific to that programming language - so everything from autocomplete to linters work.
I know Emacs is niche, but I can't recommend it enough.
--
[0] - BEGIN/END_SRC block is under convenient autocomplete of "<s TAB".
[1] - this is a real note, so if I got the physics wrong, I just made a fool of myself publicly -.-
I don’t think it quite makes sense to compare these notebooks to Knuth’s literate programming. The whole point of that was that you could present things out of order, which is impossible and actually a huge pain point for notebooks.
The first version of MathCad (for DOS) came in 1986, but it's difficult to find info on how it looked like. Did it already have the notebook interface? This is how MathCad looked in 1989:
https://en.wikipedia.org/wiki/File:Mathcad_252_screenshot.pn...
Mathematica 1.0 came in 1988, and it definitely had the notebook-interface.
https://reference.wolfram.com/legacy/v1/contents/whatis.html
Wikipedia says Maple got its first graphical interface in 1989.
> 10 PRINT "HELLO WORLD"
> RUN
HELLO WORLD
>
Just being a smart-aleck, of course, but scientists have been using interactive programming tools since they became available, and they have only grown in sophistication.Most editors can't open a terminal that you can use VIM keybindings on to search/navigate history and treate like any other buffer.
VSCode -> not currently possible because they wrote it in a restrictive way with Panel as a special case very different to code window Atom -> probably possible but I don't think terminal-plus is quite it. Any IDE I've tried -> not possible. Emacs -> possible.
Not that this is the be all end all feature but it is useful as hell and kind of a litmus test for whether you can program your environment.
edit: LightTable seemed kind of cool but became abandonware like the author's other projects
I don't think it supports cell folding (?) as an example missing feature
Minor point but it also can't/shouldn't support widgets, which we use at work. Any extension to jupyter is going to be written in javascript so there's an element I'd be locking myself out of the ecosystem.
I didn't mean this as a criticism of emacs, my post said that the best thing I know is emacs, just that it's (probably) not the future of editors imo so I'm reluctant to throw 100 hours into it.
If I understand what you mean correctly, that would be handled by built-in outline-minor-mode, or by a third-party Emacs module like yafolding or fold-this.el. Emacs packages tend to be made to compose well with other packages (it's a requirement given how everyone's Emacs is a special snowflake, unlike any other Emacs).
As for widgets/Jupyter extensions, then yes. Emacs can't really help you there AFAIK.
Would this work for other languages? Maybe JavaScript or Powershell instead of Python?
Notable examples...
R: https://github.com/IRkernel/IRkernel
node: https://github.com/notablemind/jupyter-nodejs
In my mind, one of this big advantages of jupyter is its extendability. I was able to quickly modify notebooks to run unit tests, so we could use them for projects at DataCamp: https://www.datacamp.com/projects.
I am curious people's thoughts on using Jupyter for long-running code. Having a totally self-contained experiment in one notebook, even if it long-running, is very useful for reproducibility. It works fine on my local laptop and a remote server, but not with SageMaker.
My experience as a data engineer/architect/application developer attached to data science teams for a while now is that most really good data scientists are very good at what they do, write somewhat competent code, and do not--in any way--care about writing good software or good application code.
Jupyter is a bane of my existence because people who use it want to use it for everything. Oh, it can have a web interface? Okay. The app is done. DEPLOY TO WEB USERS! NOW!!
It's a great tool. A lot of the people who use it are not software engineers, and they don't want to be. For a lot of people it's the straight line from point a to point b.
But in my experience, legit data scientists are pretty smart and are willing to learn a little if you're willing to give a little. This is a good exercise because they are typically skeptical about everything. So you have to be really secure about why you want certain things done certain ways, and why you definitely don't want things done other ways.
It's a good exercise for everyone involved if you have the right team dynamic and mutual, healthy respect for each other.
If you don't . . . well, then Jupyter notebooks completely suck.
However, I cannot stand typing any text into a web browser window. Is there any way to edit a jupyter notebook with a text editor and then run it in the browser? The native json is not really human-editable.
This would be possible today if the notebook file was python code with comments, for example, instead of an uneditable json.
I personally look forward to trying this out, as it means that I can use Jupyter in a way that doesn't mean adapting my workflow to the tool so much.
[0] https://towardsdatascience.com/introducing-jupytext-9234fdff...
If I am just poking around I use a Jupiter notebook. If I have to do a lot of prototyping, I use the emacs plugin. So the muscle memory typing works.
I find the whole "notebooks are a revelation!" thing kind of amusing, given that we have had REPLs for a long time. emacs is just a big REPL if you know elisp. But, yeah, ein is great.
https://nbviewer.jupyter.org/gist/ontouchstart/58c62c8248540...
Python is a good "glue" for optimized C++ or fortran libraries which the core of things like tensorflow or numpy. Everything fits together nicely.
maybe some part of it, but the Lisp Machine UI has a full window system, many different applications based on it with different UIs (font editor, file system browser, process overview, chat program, terminal, Zmacs editor, debugger, documentation browser, documentation editor, drawing program, ...)