Azure Jupyter Notebooks
notebooks.azure.com
notebooks.azure.com
What is it?
It's an offshoot of Azure ML Studio which has Jupyter support. We asked the powers that be if we could also instantiate it also as a free service - "yes" was the answer.
Who is it aimed at?
Students, faculty, casual users, folks that want to give webinars, classrooms, etc & want to skip software install headaches.
What does it support?
Currently R, Py2, Py3, F#. Python is backed by the Anaconda distro. More kernels will be added based on user feedback. Environment runs on Linux/docker.
Is it free?
Yes.
I need an account?
Not to view content (a la nbviewer). To run, create notebooks, etc. you need a Microsoft account (xbox, outlook, hotmail, ...). Sample notebook to view (click on the eyeball):
https://notebooks.azure.com/library/LIGOOpenScienceCenter
I don't like the UI!
We are admittedly not UI people and are grateful to our summer intern for the current UI! Please send feedback to nbhelp@microsoft.com and we'll improve it!
[EDIT: additions:]
Can I get a bash prompt, install linux pkgs?
Yes! In Jupyter, you can click on Terminal & you are in bash (ubuntu).
Can I use pip, install.packages(), nuget, ... ?
In Python/R/F#, you can use each environment's pkg mgr to install pkgs. EG "!pip install pkg", install.packages("ggplot2"), etc. See Py examples here:
https://notebooks.azure.com/library/Intro_To_CNTK
Is my environments saved?
Currently, your notebooks are saved based on your login. Proper data, load github repo, etc. support is coming.
For a bit more info, please view the faq:
https://notebooks.azure.com/faq
Thanks! [msft]
No thank you! I have been running a server on my own server at home, but this has so many awesome possibilities. I can see this as my go to for introducing Python or R in the near future.
Giving away access to a public Jupyter Server (open source project)
Running it in Linux (Ubuntu) and letting users have a play with it.
I for one am super happy to play with it however long this lasts.
Some info so far.
!free -h
total used free shared buff/cache available
Mem: 55G 7.5G 26G 70M 21G 44G
Swap: 99G 0B 99G
!uname -a Linux nbserver 4.4.0-51-generic #72-Ubuntu SMP Thu Nov 24 18:29:54 UTC 2016 x86_64 x86_64 x86_64 GNU/Linux
Look! I can even install stuff I want from pip...!pip install terminaltables
Collecting terminaltables Downloading terminaltables-3.1.0.tar.gz Building wheels for collected packages: terminaltables Running setup.py bdist_wheel for terminaltables ... - \ done Stored in directory: /home/nbuser/.cache/pip/wheels/96/0c/9a/0ec2bcad2ac1fb1d0e4695879386460cec1a947cd7413d1b17 Successfully built terminaltables Installing collected packages: terminaltables Successfully installed terminaltables-3.1.0 You are using pip version 8.1.2, however version 9.0.1 is available. You should consider upgrading via the 'pip install --upgrade pip' command.
Data point - our team (Python), wanted to hold a "Python day @ msft" event to teach/inform internal teams on the language/ecosystem. We were hoping maybe 50..100 would show up, if lucky. A thousand and twenty one people showed up!
Another data point - the official Azure CLI was just rewritten in Python (win/macos/linux):
https://github.com/Azure/azure-cli
Thanks!
RE Jupyterhub on Windows - if enough customers want it, I'm sure it can be done. Not convinced the demand is there yet. If you have data I'd love to see it!
Thanks.
https://github.com/jupyterhub/jupyterhub/issues/703
Also original issue here and "Wontfix" comment:
I'm a performance guy at my work responsible for making sure our Filesystem (Huge codebase of C, and bits of C++), making use of Python to quickly write scripts for Performance data analysis. Right tool for the job.
It is interesting to think about how many use cases there are in this ecosystem where that hard attention to performance is "wanted" but not needed. Could I write my scripts in a complied language and save two seconds of execution time? Yes, but I finished two more tasks in the time it would've taken to do it.
https://notebooks.azure.com/library/puzzles
I had to host the images in the notebook elsewhere rather than upload them with the notebook. It looks like that is possible but then I'd have to grapple with Azure blobs...
TableXPlorer (for Azure Table Storage is decent too)
(i.e. it's not stripping them out and making you manually re-embed them each time if they aren't an Azure blob?)
[msft]
It's really lovely when a tool written in another tech (Python there) can be used/extended with others stack/tech, because was written as language agnostic. That should be a rule for good design in oss tooling. Not starting using language specific communication, but extensibile from day0 in design (obv default language can be bundled)
Chart.Line(data) |> Chart.WithTitle("smth") |> ...
...makes anyone say "yuck" and avoiding F# for any exploratory work. And even after you understand how the pipe operator works and how useful it can be in other contexts. I mean, ugh... Even putting up with superfluous verbosity like `List.Map(myList)` instead of a `map(myList, ...)` or `myList.map(...)` because "that's how F#/OCaml does it, stfu" is ok, but this seems like purposeful obsfucation.No wonder R and Python are the only languages popular in ML. They are the only ones leading to sane readable code by people having other stuff to keep loaded in their head than language details. I mean, yeah, syntax that doesn't matter, unless it's so f annoying and ugly that you simply can't put up with it as much as you try!
This is very, very easy stuff for functional programming. Most of the people I know in ML, talented ML folks doing data science within the org I work for to do exciting things... they actually aren't huge fans of the way Python does things. They're huge fans of the libraries Python has to speed up ML work. F# has a real lack there (surprisingly!).
But like, saying that you like:
Chart.Line(data).WithTitle("smth").render()
as opposed to: Chart.Line(data) |> Chart.WithTitle("smdh")
seems to me like 6 of one or half a dozen of the other. Especially since when you're reusing any given part of the pipeline, the ML way is alot more terse.And it's not like Python is devoid of hideous examples of poor language design. For example, self parameters. All we're told is some new age mysticism about how "explicit is better than implicit", but somehow that doesn't apply when we need higher order functions and we're using private function with lambda names to try and signal to readers that, "Sorry I needed 2 statements in my list comprehension and Python talks down to me like that guy from Timecube because of it." Why is explicitness valuable there but not on tuple construction?
And it's not like Python doesn't make you arbitrarily choose between List.function and list.method with no seeming decisions to make.
Which is not to say Python is especially bad. It is to say that you may have normalized Python so much that you forget all computer languages are somewhat arbitrary.
Lists a bunch though the R interop is probably the best/most fully formed IMO.
No, avoiding unnecessary repetition (god damn DRY at the syntax level) and avoiding visual noise is not arbitrary. It's just good practical taste. I don't live in houses where doors have 2 knobs that need to be pressed at the same time to open a door. And I don't eat steak with a fork that has a knife as a handle. Both of those could work and make sense for some, but overall they are bad taste and awkward for 99% of people.
Python gets it right. Ruby gets it right. Julia gets it right. Most lisps get it right etc. (Even Java and C# get them "as right as they could" considering all their history and past choices that seriously limited what kinds of languages they could be.)
And my issue is that I like the semantics of languages like F# and OCaml and Haskell after having played a bit with them but by god, they couldn't have chosen more infuriating syntaxes and name resolution systems or module systems or tools for them... like they tried as hard as they could to piss off "the plebs in the industry" who actually care about syntax and other such little details, because, ya know, when what you're developing is not that interesting, you should at least have the pleasure of writing code that you aesthetically enjoy to read! And it's hard to convince fellow industry plebs of the usefulness of advanced type systems when the first code sample they see elicits an "ugh" or "yuck" reaction. Most of us programmers are shallow and lazy and we should be proud of this and build tools that cater to our "virtues".
Or maybe I'm in the minority by liking dense and non-repetitive notations and finding them easier to read too...
What. I ship Clojure all day, got CL in my past. Was a full time ruby dev. I got a list of ugh and eck for all of them. I can name 30 more issues with Python. Truth be told, I think Python is a rancid language and I think people who love it are basically eating barf pancakes every day and thanking people who look down on then for it. GVR doesn't do that bs "I don't get lambdas" garbage in the company of other language designers, that's for sure.
What you consider visual noise is arbitrary. It's like what color paper you prefer to note take on. Every language has issues. Python is riddled with syntactic noise and artifacts that you've normalized. Ruby's syntax is better, but still full of quirks and surprises. Don't even get me started on Scala.
"I am used to this" and "this is objectively better unless you are some academic" is a classic example of industry insecurity. You shouldn't avoid Haskell because you are irritated with a bit of syntax, you should avoid it because you can't ship or maintain the kind of deliverables you need to write.
> Most of us programmers are shallow and lazy and we should be proud of this and build tools that cater to our "virtues".
Pride is one of the ugliest sins of our industry, I agree. Too many developers refuse to accept that there might be progress in the industry outside of what they experienced in their first 2 years in the industry.
Take this example of Haskell from their docs:
thenP :: P a -> (a -> P b) -> P b
m `thenP` k = \s ->
case m s of
Ok a -> k a s
Failed e -> Failed e
It doesn't matter how good you are in [insert almost any common language], you'll really struggle to understand that code.It's my same objection to things like Coffeescript, if you think you are so much more productive not typing semantic tokens like ()s or {}s, then you need to lay off the coffee and get some sleep as you're clearly dillusional.
Perhaps I've been using Haskell too long but that code looks very clear to me.
thenP :: P a -> (a -> P b) -> P b
m `thenP` k = \s ->
case m s of
Ok a -> k a s
Failed e -> Failed e
But this is not a fair test. There needs to be a control. Implement the same functionality in Python and then we'll talk about which language is clearer.As to the comment re: control / python, I'm sorry but it's so obvious to me the syntax would be more legible to most developers (since "most" use Algol descendent syntax languages) even if done somewhat poorly, I don't feel like taking the time. I welcome you to prove me wrong, I'll gladly stand corrected if so!
To the extent that true, it's because of the weird historic moment between the mid 1990s and now where essentially every industrially popular general purpose language is from the C branch of the Algol family.
It's not really the whitespace (which is clear in semantics), and other than the lambda slash and the infix ticks, the whole thing is pretty clear from a Lisp background. (Well, I'm familiar enough with Haskell now that I'm not filtering it through some other language, but the similarity to Lisp is what helped it start to click early on for me, even though it's an ML descendant and not a Lisp descendant.)
We're actually starting to see more gains for non-Algol descended general purpose languages, so maybe the syntax shyness of the last generation or so of programmers will soon be a thing of the past, because knowing multiple languages won't so invariably be knowing multiple members of the same syntax family.
I can ship and maintain deliverables I need to write in Haskell more effectively than other languages.
And some domains are ill-suited to Haskell due to constraints on the VM.
Is any of this a surprise?
To clarify, you were saying "you should avoid Haskell in the case you can't ship or maintain the kind of deliverables you need to write" and not "you should avoid Haskell because Haskell won't let you write or maintain the kind of deliverables you need to write"...
Right?
R is a curious counter-example, given that the '%>%'[1] operator in dplyr is nearly identical to the '|>' operator in F#. Given dplyr's ubiquity, it would seem that people coming from R would have no issue with F#.
[1] https://cran.rstudio.com/web/packages/dplyr/vignettes/introd...
If there are other languages you would like to see, or other features or issues, reach out to us on https://github.com/microsoft/azurenotebooks or nbhelp@microsoft.com
But, Jupyter faces most of it's challenges in the word-editor part of their product rather than in the code part.
I'd love to see a partnership between them and the Eve programming language (http://witheve.com/) who have absolutely mastered the IDE interface in early renditions of their product.
I know a lot of people are against the concept of Eve from a purist perspective of code not needing to be humanized. But, one of the common goals in data is to communicate insights and solutions through visually crafted and accessible stories. I think there is still a long way we can come in that.
I'd wish the VSCode team could somehow integrate the concept. They seem to be excellent at execution.
(Or, if they finish their work on "html zones" (block decorators in atom), I'll start doing it myself)
...but perhaps software developers aren't the target audience in the first place. I tried to use Jupyter a few times, first when it was still only IPython, and it never seemed to fit in my code-execute-fix workflow that you have when developing scripts. In particular having to always reset the kernel to reexecute everything from scratch drove me crazy.
Do you know of a better tool for such a flow, combined with the ability to have markdown+latex docs intermingled with the code? 'cause this is my workflow, but jupyter's code editor and kernell restart drives me crazy and anything else will have me keep the notes/docs separate from the code...
And please don't suggest Mathematica :) I absolutely love it's ux/i, but nobody uses it in ML/AI...
You can use whatever editor you want, but I quite like rstudio (I'm generally not a huge R person, so an environment with more help is useful, whereas with python I'd prefer just my own setup).
Edit - Importantly though, you actually don't need to use R, you can use python. I'm not sure how well that works with caching, as I've never tried it, but it's probably worth a go.
(Right now I use jupyter for some things, ipython gui for others, and pycharm for "real coding" tasks. Tried Spyder, but something about it makes it neither a good IDE nor a good notes/documentation system... though I can understand its appeal for Matlab folks).
Admittedly, most of my experience is with ipython but it is perfect for that. If there is a new algorithm/method I want to explore or I am figuring out an interface/structure, it is great. And I tend to not have to reset the kernel too often.
In that sense, jupyter was a great idea as you can now integrate documentation with code in a nice format.
Unfortunately, it also led to a strong focus on treating them as containers for the purpose of deploying code (more traditional software development).
The restarting the kernel biz is painful, of course. But the interspersed plots and code are really useful.
Or had access to papers and videos showing how the development environments at Xerox(Lisp, Mesa/Cedar, Smalltalk) and Genera (Lisp) worked.
Hence why I am not found of having a graphics workstation full of xterms.
We need to make these workflows mainstream, not something that our descendants are reading about in paper and videos.
The only thing that's really missing for me is a more persistent data store in between kernel restarts. If it took more than 5 minutes to run something to transform or process my data, I don't want to have to redo it when I restart the kernel. I think there are a couple of plugins that handle this for you, but it would really be nice if it was implemented natively. The solution right now just seems to be producing a bunch of intermediate files that you reload when you restart the kernel.
I haven't played with it yet, but another HN user pointed it out to me recently.
With that in, now the Monaco integration is being worked on: https://github.com/jupyterlab/jupyterlab/pull/1382.
We all want stronger editing capabilities, but it doesn't make sense for the Jupyter team to get into the business of writing text editors (plenty of better folks doing a great job on that already). So we're just trying to make it easier to integrate other text editors into the everyday workflow.
Thanks Fernando, Brian, Min & team for everything you've done with Jupyter. Looking fwd to using Jlab soon!
s
But yes, JLab is shaping up quite nicely, opening up a lot of interesting possibilities. For advanced users/early adopters I think it's time to start playing with it (and filing issues for anything that's broken/sub-optimal, we really want to provide a great user experience with it once we hit 1.0).
RStudio's new feature to R is Notebooks and it is available in RStudio version 1.0+. It takes what is great about Jupyter Notebooks and adds easier version control and much easier to batch process your reports. Which are both huge wins for me.
http://rmarkdown.rstudio.com/r_notebooks.html
Blog Post: https://blog.rstudio.org/2016/10/05/r-notebooks/
"Interactive R Markdown
As an authoring format, R Markdown bears many similarities to traditional notebooks like Jupyter and Beaker. However, code in notebooks is typically executed interactively, one cell at a time, whereas code in R Markdown documents is typically executed in batch.
R Notebooks bring the interactive model of execution to your R Markdown documents, giving you the capability to work quickly and iteratively in a notebook interface without leaving behind the plain-text tools and production-quality output you’ve come to rely on from R Markdown."
Afaik, hydrogen plugin for atom does this [2] And there is another plugin that just wraps the notebook to live inside of atom [3]
[1] https://talkpython.fm/episodes/show/44/project-jupyter-and-i... [2] https://atom.io/packages/hydrogen [3] https://github.com/jupyter/atom-notebook
Admittedly, I don't use a code editor, perhaps because I'm just too old, and got used to living without one. I remember when being able to edit a program in full screen mode was a big deal. But I certainly wouldn't turn down better editing for Jupyter.
> "Internet Explorer is not supported by Jupyter: For best results use Microsoft Edge, Chrome, Firefox, or another modern browser."
This makes me so happy - MS products leaving old browsers behind gives all web developers a ball and bat in the same fight.
so kind of a learning environment, because you can have people fork something like a ML notebook and play around with the different parameters
> Currently, yeah, just JS.
How do they plan to make it support multiple languages without having a backend service?Not claiming this is a good solution, but definitely would be one possible route.
A professionally maintained, sufficiently sandboxed Jupyter environment could be awesome for people who want to work on and share Jupyter notebooks, but not be responsible for servers.
Sharing is basic now (ie unlisted-URL), but proper support (public/private/ro/etc) is coming soon.
(Disclaimer: service is run by my employer)
For a sample course that was taught for ~400 students this fall, see:
https://notebooks.azure.com/library/CUED-IA-Computing-Michae...
Click on the eye icon. You won't be able to run the notebook code, which would be a surprising thing to allow without some sort of authentication.
Also, how many middle school and high schools (especially title 1 schools) do you know have office365 for the teachers, let alone for each and every student?
Please don't regard this reply as being critical of the offering, it's certainly not and I applaud what you're doing, I'm more responding to the person who couldn't understand why login was a barrier.
Perhaps www.code.org is a better offering?
Thanks! Azure Notebooks Dev
Awesome! Cambridge University (Dr Wells) just did that. Here's their notebook:
https://notebooks.azure.com/library/CUED-IA-Computing-Michae...
Ok, let's use OneDrive to load data. Where is that option?
Nvm, let's use Dropbox instead. Does not work "Something went wrong. Please refresh the page."
Maybe Azure storage? Trying to sign-up, filled sign-up form, verified account via SMS, entered CC data. Nope! "Cannot proceed with signup due to issue with your account. Contact support". (tried multiple times)
Contacting support.. creating incident, describing the issue. When trying to attach screenshot I get "The file upload service server is not available at this time. Wait a few minutes and then try it again."
Seriously, Microsoft??
RE Storage/CC sign up. I'll locate someone in that group & reply back.
"Usage should be limited to learning, research, general computing, etc."
- Azure Notebooks Dev
I'm doing something similar with exploit development learning, but with a javascript based terminal and linux containers, and a markdown writeup (https://exploit.courses for anyone interested). But the close interaction of code and text in Jupyter is much more advanced, and useful :-)
I'm not sure I understand this bit though — by "data" do they mean they'll delete a notebook not accesssed for 60 days?
> Storage: We reserve the right to remove your data from our storage after 60 days of inactivity to avoid storing unused/abandoned user data
As we are still new, we haven't started deleting files yet, but we wanted to be sure to have a policy for handling this upfront and not decide when there was a problem later and then start destroying content without any sort of notice.
Thanks! -Azure Notebooks Dev
And it's really slow.
It took me 55 seconds to load the page (42 of which went into downloading a 6 MB large GIF). 40+ requests. And it doesn't seem like a slow Internet problem at all.
The UX was not done in any way to feel like something else. We tried to design something that worked well for Jupyter users as well as looked nice and was usable.
That being said, our teams strengths are not in UX by far. We are starting to work with designers to improve the "dev-UI"" we currently have.
If you have any more feedback you would like to share we would love to hear it. There is a public tracker at https://github.com/microsoft/azurenotebooks and an email that goes to the team at nbhelp@microsoft.com
"The Jupyter Notebook is a web application that allows you to create and share documents that contain live code, equations, visualizations and explanatory text."
And for some reason the back button doesn't work once you encounter the login page. That's enough to make me want to stay away.
Any notebook can be opened in an html view: https://notebooks.azure.com/library/fsharp/html/FSharp%20for...
It sounds like you would like being able to view those more easily without logging in?
Thanks! Chris
IMO a design that might be less surprising would be for clicking on the name of the notebook to take you to the read-only view of the notebook. From there you could click on something that would promote it to an executable/editable view and login if necessary.
But the broken back button on login page is a real bummer.
EDIT either it was my imagination to begin with or it's already fixed because I tried it again and the back button appears to work correctly.
You can set these things up in AWS and GCE, but it requires a lot more work. Here is my top google hit for setting things up on AWS, from a relatively recent and maintained repository: https://gist.github.com/iamatypeofwalrus/5183133
Many data scientists just run their notebooks locally. Few have the luxury of Data Science Engineers on hand to help them through the complex world of provisioning and maintaining servers, even in this canned AWS world where spark clusters unfold themselves in EMR.
Azure is absolutely full of "here is a canned solution for _insert common engineer task with open source focus here_" buttons and it's really quite great as a result.
Also: Azure is doing a very good job of making it difficult to have a default-insecure instance. You don't make choices about encryption-at-rest, for example. That's forced on you. There is a lot more tooling around sharing accounts and giving users minimal ACLs rather than just doing the simple thing in AWS or GCE and giving out power user accounts.
It's simple stuff, of course, but this simple stuff is the simple stuff AWS has elected not to do, so...
Our 1st interaction with the awesome Jupyter team was enabling it in Visual Studio / Python:
https://github.com/Microsoft/PTVS/wiki/Using-IPython-with-PT...
After that we loved the project (and the team) so much that we donated a $100,000 for furthering development. They used the funds to hire a dev & help improve the Jupyter for all platforms.
Since then we've put Jupyter in Azure ML studio which combines a drag/drop ML canvas with that of a notebook. You can use it to analyze your data / build models /deploy to Azure.
As an offshoot of the above project, we wanted to give something back to the community, so we figured we'd instantiate it as a separate notebook-centric offering for dabblers, students, faculty, demoers, webinar/seminar speakers, etc.
And by "we", I mean a small group in Microsoft who cooked up the idea - if it turns out to be useful, we'll keep it going. If it's not, we'll shut it off. So far thought it looks like it's hit a chord - thousands of notebooks have been created just in the past month. Various schools are using it to teach cs101, eng101 courses, people are using it to demo their products/services (an executable doc page vs a static one), ...
Beyond this initial offering, we hope to add a suite of features both around and inside Jupyter. For example, we hope to be working w the core team to add debugging/profiling support soon. We also hope that the VSCode editor will be integrated for the JupyterLab environment. We'd love to hear your ideas!
https://en.wikipedia.org/wiki/Dynabook
Is that what it basically is?
See:
Sample notebooks (html captures of runnable notebooks):
Smalltalk is somewhat of a primitive browser and very much a REPL on steroids. Instead of being polyglot from the standpoint of supporting multiple languages, Smalltalk was designed to be usable by people from grade school students up through researchers. (It largely succeeded in that goal, but failed in that it didn't attain enough popularity with mainstream engineers/programmers.)
Some of the educational uses of Squeak Smalltalk are very much in the same vein.
If this is blocking you, there are libs to use dropbox throuhg python: https://pypi.python.org/pypi/dropbox
-Azure Notebooks Dev
https://github.com/Microsoft/AzureNotebooks
(user request > our wishes)
1: Click "sign in". 2: Choose account I want to signin with 3: Click "cancel". 4: I get "500, Runtime Error" without custom error page.
Anyway, looks interresting