How Powerful Are Microsoft Azure’s Free Jupyter Notebooks?
walkingrandomly.com
walkingrandomly.com
Regarding the perf #'s - just a heads up that you were probably on that VM all by yourself :). Though once enough people sign in or cpu threshold gets to a certain level, new VMs are allocated.
PS On a related note, we are wrapping up at PyCon in Portland, and one of the hit swags was the "Jupyter Notebook Notebook" :). See pics here:
https://goo.gl/photos/L9C4fq6AsPxU7bfq5
/disclaimer: team lead/
How well does it behave when you've multiple VMs running on the same hardware?
Also, is it possible to setup prebuilt environments on your platform?
* It's possible - Anaconda comes with an MKL enhanced version of its math libs.
* There are multiple VMs, and each VMs hold multiple docker containers, one for each Library/user (library == collection of notebooks)
* Right now, no, but we are working on it. You can have a "prep" notebook where it readies your environment with !pip, !wget, ... etc. and then actual work notebooks. We'll soon have an initial "install.sh" that will be run upon start to run any prep steps you might have.
thanks!
- Weird behavior when disconnecting / reconnecting to sessions, especially from multiple computers.
- Tendency to flake out on long running jobs, i.e. 2 hours of the way through a 4 hour algorithm something dies and I have to restart or run from terminal.
Unfortunately this relegates it to exploratory viz for me, but maybe that's the intended use case anyway. But when I've wanted to build semi-persistent dashboards or check in on running jobs I've had better luck with ssh+screen and then dumping pdfs of results with matplotlib to files that I serve from a webserver with a little auto-refresh javascript wrapper.
That said, I always made sure to save before disconnect and refresh on reconnect.
But what if you add a library where you can only load data from Azure ML Studio? then you cannot share your notebook anymore. Your notebook got tainted with proprietary stuff from a specific vendor...
Science is about being able to universally reproduce experiments, and vendor lock-in prevents that. We already have enough problems in scientific publishing with journals.
So if you like Jupyter, and you want it to become the standard science needs, avoid proprietary extensions. Let's not go back to share stuff in paper or its modern equivalent, PDFs.
In my workplace, Jupyter needs special approval and can only be installed in limited environments because the company is afraid some developer will inadvertently do something that contaminates our code with GPL3.
IANAL - but note that we also have Microsoft R (and enhanced version of CRAN R) which provides local multi-threading, and cluster level parallelization + distributed memory support. It and the stock R interpreter have lots of GPL code. And given that both R & Python are now integrated with SQL Server, it seems that the lawyers have become comfortable regarding the separation lines. SQL Server for example calls an external script for R/Python.
https://www.microsoft.com/en-us/sql-server/sql-server-r-serv...
https://blogs.technet.microsoft.com/dataplatforminsider/2017...
Thanks!