Jupyter Notebook 5.0
blog.jupyter.org
blog.jupyter.org
I'd like to see some things ported into Jupyter from R Notebooks, like JavaScript data tables and the separation of code and output, making it easy to version control only the code. (Atleast this 5.0 release makes tables nonugly)
It's in alpha right now but they're making crazy fast progress on it. I've been using it and it's awesome
Thanks so much for all your hard work!
Or do you just mean that R Notebooks are specifically better than the R support in Jupyter?
Here's an example of one of my notebooks with all 3 things: http://minimaxir.com/notebooks/breach-network/
It's definitely a better fit for R.
https://github.com/bloomberg/bqplot
They even have gamepad/controller support, which allows dynamic interactions with widgets.
Edit: Live demo of the project from PyData London 2016: https://youtu.be/eVET9IYgbao?t=27m45s
These widgets support embedding in other contexts like static HTML pages or Sphinx docs [1]. An example can be seen on the docs [2]
[1]:http://ipywidgets.readthedocs.io/en/latest/embedding.html
[2]:http://ipywidgets.readthedocs.io/en/latest/examples/Widget%2...
Currently, Jupyter works well for development, but the result is often hard to read because code gets in the way. (unless you use nbconvert, but that sometimes defeats the purpose)
That said, both Jupyter and R Markdown Notebooks are but a pale shadow of the support offered by Org-mode (seriously!).
Know of any good resources that show this for folks that don't know what we are talking about?
Yes, but the large number of excellent notebooks available for Jupyter (and R as well) all over the web as well as support for CUDA and all kinds of extremely powerful libraries such as tensorflow) give those a serious edge over Org-mode, even though Org-mode is super powerful by itself.
If you setup a source block, you can use TRAMP to actually execute the command in a source block on a remote machine. So:
#+BEGIN_SRC sh :dir /user@remotemachine.com:~/remotedir
ls
#+END_SRC
The above will run ls on "remotemachine.com" and put the results in an output block below it.Edit: Meant to add that you can do the same with docker. Just use "/docker:dockerId:" as the dir, and it will execute in a docker instance locally. Using multihop addresses, this can get extreme.
Org mode is one of the few real literate programming tools that I'm aware of, the other one is 'Leo'.
I used to try complicated examples, which led to pages such as http://taeric.github.io/Sudoku.html. I'm now much more into writeups such as http://www.howardism.org/Technical/Emacs/literate-devops.htm.... (Note, I did not write the second one.)
I'm really intrigued by it, but Jupyter is much more clearly documented (org docs are downright sprawling), so I've always gone that path.
There are R packages which allow communicating with a server, mostly big data packages. (sparklyr allows you to connect to a remote Spark cluster)
Other than that the plan is to keep this (or something similar) running as long as there is sufficient interest and so far there seems to be quite a bit!
Note that it does run on docker which means ultimately it's not fully "secure", but we hope to switch to hyperv-linux when it's available.
[smortaz at msft]
And there's nothing wrong with knowing and using both languages.
I'm definitely using this for my next project.
from rpy2 import robjects
my_dictionary_results = some_method()
names_dict = robjects.ListVector(my_dictionary_results)
%load_ext rpy2.ipython
%R library(some_lib)
%R -i my_dictionary_results doStuff(my_dictionary_results)
Image display etc. works fine
I'm not sure how others work with this stack, so there is likely tooling I don't know about. I'd love to hear suggestions.
! git commit add ./my_file.ipnb
http://ipython.readthedocs.io/en/stable/interactive/python-i...From the docs: "nbdime provides tools for diffing and merging Jupyter notebooks."
It includes both graphical diff and merge tools, command line diff and merge tools, VCS integration for git (so git uses the nbdime diff and merge for notebooks), etc.
Might be worth looking into if that is what you are looking for.
To be fair, wasn't Jupyter designed to be an IDE - experiment with code, make tweaks and get the final version out in a text editor - rather than a repository of production-ready code?
(You used to be able to run ipython notebook --script)
The obvious problem is that accessing Jupyter is technically similar to allowing full shell-access and you have to deal with local privilege escalation, but I wonder if there has been any progress. I evaluated to use it as a UI for domain-specific applications that give users some kind of graphical shell, but in the end I decided against it, because of security concerns.
Or https://github.com/jupyter/docker-stacks/blob/master/r-noteb...
Unauthenticated, sandboxed notebooks in Docker containers. There are various limits.
If you manage to break it, please let me know!
(Regret incoming in 3... 2...)
Any learning from multi-user deployment ? We are trying to do this internally inside our company and jupyter hub is a little hard to grok.
I know of a lot of people who would pay for a faster jupyter that can also be run as a standalone dashboard/script - without the heavy duty interactive kernels. Basically reduce the prototype-deploy loop.
Check some of the comments here - https://news.ycombinator.com/item?id=14033129
It convinced me that there really exist viable open alternatives to the proprietary closed systems and I am extremely thankful for your efforts - I will definitely check out sagemathcloud.
Some form of isolation - be it containers, or jails, or VMs - is going to be part of any solution precisely because of this.
I would actually dare say that FreeBSD jails are probably the best (most stable and secure) candidate of those available currently.
Does anyone know if there's a plan to introduce multi-kernel support in single notebooks like what Zeppelin does? Not that I have a strong preference for it but it appears to hold a lot of appeal in Spark-like environments where not all packages are available in Pyspark and you need to move between native Scala/spark and Pyspark.
What exactly do you mean here? Are you referring to the parts of Spark which don't currently have a Python API? Because those are becoming smaller and smaller.
There is also Jupyter magics to let you change languages within a notebook. See %Rpush and %RPull from [1]. Not sure if there is a way to have a Scala kernel running and sharing the same Spark context though.
I think IBM is working on something in this area.
[1] https://blog.dominodatalab.com/lesser-known-ways-of-using-no...
But, the ease with which you can load interpreters on Zeppelin (apart from the pre-loaded ones) is impressive. I imagine it comes at a cost of some instability because it hangs more often than Jupyter.
[2] show you haw Hydrogen (based on Jupyter as well), does it.
And [3] (mine), show you how in the the same notebook to use Python, R, C, Rust, Fortran, Cython and Julia with data sharing and sending functions back and forth between languages. I was definitively lazy and did not include things like SQL, javascript and a few others. I haven't used spark in a while, and definitively never from scala directly, but I doubt it would be much harder to do as the C/Rust/Fortran/Cython took me an afternoon to write.
[1]: https://github.com/Calysto/metakernel [2]: https://www.google.com/url?hl=en&q=https://medium.com/nterac... [3]: http://carreau.github.io/posts/23-Cross-Language-Integration...
We are building dashboards in Jupiter and really would love it to be multi user... Without getting into the hub and stuff (way too complex to set it up)
Sorry to hear you found the hub too complex. We're working on making easier-to-use hub setups that fit different use cases. Can you tell us a little more about what your use case was and (optionally) which parts of the hub setup you found too complex?
Thanks!
So it's a bunch of different technologies - nodejs,etc. I'm kinda wondering if it can be built in Python itself. Make it part of a normal jupyter install, so just a "jupyter hub start " will work ?
EDIT: adding to that, you have built a nodejs based http proxy - can you not build it within Python (for uniformity) or nginx (for performance as well as mind share) ? Do you even need to mandate a http proxy ?
Second question is that can it run in a multiprocess - I don't want to run it in interactive mode, but just straight top to bottom. Perhaps there's huge memory savings there.
So it becomes a traditional webapp use case. Do you need all the proxy/websocket, stuff to do this ? Your nbconvert command still needs every user to spawn his own kernel right ?
About the first part - it would be great to have a simpler jupyterhub. One of the steps is to have everything in Python.
There's ongoing work on formalizing the proxy better (https://github.com/jupyterhub/jupyterhub/issues/848) - someone will probably write a pure python proxy when that gets merged :)
All the tools already exist in jupyter - except one: lightweight multiuser. I would argue that building this is going to be a fairly trivial thing for you guys (as compared to other features you build), but the end user benefit is immense.
Jupyter becomes much more than an interactive scratchpad - it becomes a full blown prototyping environment for data science and reporting. I would say, you would even go against tableau in a lot of use cases.
Please do think about it. My company will be happy to contribute to a gofundme on this.
I'm talking specifically about the usecase of code-execution (especially dashboards).
Here's a small point from me - perhaps you are overcomplicating the usecase for 90% of us. Give a proper ssl/tls+bcrypted password setup and roles: Editor and User.
I dont think you should be worried here in the context of people wanting to run a full on sagemath cloud kind of a thing.
If you can give me a low resource way of letting 100 "Users" on a dashboard form and one "Editor" (who can actually edit the underlying notebook), I'm golden. And I'm willing to bet that so will 90% of your audience.
Jupyter extensions are incredibly powerful, and I don't know how I used to do data analysis before finding them.
[1] http://jupyter-contrib-nbextensions.readthedocs.io/en/latest...
The new style right aligns everything which is good for numbers but bad for text. Also, it might just be the screenshot, but the contrast is poorer and the column sizes aren't as fitted.
It's a pity, I would rather have a decent default than to customize every notebook.
It'll be second to Github source code management with Jupyter support.