Announcing .NET Jupyter Notebooks
hanselman.com
hanselman.com
I think the only real drawback is no self hosting.
A pretty healthy example:
https://observablehq.com/@rreusser/2d-n-body-gravity-with-po...
I haven’t gotten properly stuck into observable - in what ways is it better than using a jupyter notebook?
But yeah I'd have to agree that Bostock is a beast.
Observable does some things quite well for that use case:
-- Easier dependencies: no need for 'npm install', just 'require(...)'
-- URL publishing
-- Collaborative merge flow
However, Jupyter made some good decisions that make it win over Observable for our day-to-day data work:
-- Manual reexecution vs. automatic: when working with big data, outside APIs, etc., Observable's automatic reexecution is a non-starter
-- Fully open source, embeddable, and successful history of non-vc funding: Jupyter is organized and provided in a way companies ( who aren't Amazon ;-)) can rally around, evidenced by contributions by Bloomberg etc.
-- Access to underlying unix/windows env and multiple environments (multiple Python versions, ...). As much as I wish Observable's choice of JS could provide a viable data environment, and I've personally invested in making it so and the path to making it first-class is clear, we need way more gov/google/nvidia/etc. support.
Ultimately, for not-too-sensitive collaborative data work, we use Google Colab, then offline and sensitive commercial work via Jupyter (Graphistry ships with it preloaded for GPU dataframe & GPU visual graph analytics goodness!), and if we did more in JS tutorial land, Observable would be my first choice.
Yeah automatic re-execution sounds like a weird design choice. Think I'll stick with Jupyter notebooks for now...
The Jupyter documentation includes an example for importing other notebooks here: https://jupyter-notebook.readthedocs.io/en/stable/examples/N...
while you're browsing: this is equally great :)
https://archive.nytimes.com/www.nytimes.com/interactive/2013...
Especially crazy that these work as well and quickly on mobile!
The (reactive JavaScript) Observable runtime is open source, and every notebook is available compiled as an ES Module.
You can embed notebooks in their entirety on any webpage, or just grab the embed code for an individual cell, if the notebook produces a single visualization.
For the details in all their nitty gritty, see: https://observablehq.com/@observablehq/downloading-and-embed...
A comment from May last year [1] says "But that won’t include the editable notebook interface or the rest of observablehq.com - unlike Jupyter, we’re really aiming to create a community and cross-pollination of concepts and code rather than individual installations. We might eventually launch something that enables offline editing, but that’s a bit further out in the future."
[1] https://talk.observablehq.com/t/noob-can-you-run-your-own-ob...
> Polynote is an experimental polyglot notebook environment. Currently, it supports Scala and Python (with or without Spark), SQL, and Vega.
It supports javascript and typescript relatively new but capable.
With C# it might be a different story, since the development cycle is fundamentally edit-compile-run an environment like Jupyter might give me an extra platform for in-between testing. I already use the Interactive built into VS quite a bit to figure out, for example, just the right format string for Datetime.Now.ToString() or stuff like that.
It isn’t really aiming to replace a dev environment like VSCode.
We’re seeing a rise of notebooks in particular because they encourage writing a lot of plain text around your code, and that’s really good both for teaching, and promoting good documentation skills for learners.
>encourage writing a lot of plain text around your code
Is this targeted at non techies? Teachers?
Can someone provide some good example of this? Not having used it much trying to understand what the advantages are over good quality documentation/tutorial pages on say msdn.
The use case is basically a scientific journal. You collect data. Analyze it. Make plots. Write notes (e.g. with equations, or diagrams). All in one place.
Then you can share the notebook with anyone else. If they have similar data (in the same format), they can execute the notebook on their data and get the plots all in one place.
It's great for data oriented work. It was never meant as an IDE replacement, nor as a general purpose development environment.
In a previous job I had set up a notebook to analyze some of the models our team produced - it was essentially a Q/A notebook that generated data from our models, algorithmically looked for unphysicalities in our models, and plotted any it found.
The rest of my team used it for the models they were working on. The alternative/old way was just too painful (lots of manual steps).
But although reuse by others is a touted feature, it's not really that important. It's incredibly useful for one's own workflow. Think of the convoluted Excel sheets people often have at engineering companies. They get new data, copy and paste it into Excel, and get new plots. This is no different.
Can't imagine trying to pump my data into someone elses pipeline and crossing my fingers for results!
Sure! Here are some ways I use notebooks:
Imagine any scenario in which you'd rather make a video to demonstrate how something works. For a lot of those cases, a notebook is a great way to show how the code works and how it all comes together.
Another set of use cases is you've done some research, need to call a few api's, post process the results, and share that with your team or have it ready for later to do something similar. You could think of this as a "Super Postman". For example, you need to run some api's in a loop, filter the results, and accumulate some things and print totals at the end.
Another use case is you're troubleshooting an issue, and want to keep notes of what you've tried, what the results were, and be ready to go back and run things again.
You can even have a notebook run in a cluster.
Well, that's exactly what Databricks does.
https://en.wikipedia.org/wiki/Literate_programming#Contrast_...
It's a convenient way to present both data and narrative to whoever might need to see it.
I also like being able to run snippets when writing scripts, at least when it comes to building/testing. I'm more on the data science side than the dev side though, so that's probably why Jupyter Notebooks appeal to me more than other IDEs.
Edit: if you think “oh that’s nice, just like Java/C++”, HECK NO. It goes beyond this, where complex modeling can easily be represented and things like pattern matching and other features not known in other langs
That said, improvements to the language can only go so far. We're trying to figure out what a good set of libraries for the whole range of ML tasks looks like for F# and .NET.
The key, IMHO, is going to be getting some influential people and projects on F# generating code and blog content. The user experience of getting setup for them and the audience they reach needs to be streamlined. From there if it JustWorks™ F# should do the rest of the selling :)
And to digress, because I always mention this when talking about F#, it deserves a batteries included web framework story as good as Elixir's Phoenix.
I'd check out Saturn for that: https://saturnframework.org/
Since C# is my daily language, this is pretty cool though. I just really wish that I had a use-case strong enough for it to overcome the initial pains of learning it and the ongoing mental maintenance of staying usefully fluent in yet another set of libraries.
That being said, I've found the only way I can get interactive programming to stick in my mind is to pick a sample dataset, wonder some open-ended questions about it, and then answer those questions in a Jupyter notebook with the help of Google, documenting your thought process as you go. A great starting point to find real-world sample data is the UCI dataset repository: https://archive.ics.uci.edu/ml/datasets.php
And to ”close the loop”, you can of course access vscode using browser with Visual Studio Online [2].
[1] https://code.visualstudio.com/docs/python/jupyter-support
[2] https://visualstudio.microsoft.com/services/visual-studio-on...
This is me_irl
like others said already, jupyter isn't a development environment and trying to use it like one is asking for trouble, but it is a great REPL on some awesome steroids.
Even better if it works well with Deedle and FSharp.Charting. But I'm not going to demand everything on release day.
There are a couple of efforts to build .NET wrappers around Pandas, or to reimplement Pandas in .NET; see http://scisharp.github.io
I am really excited about the new Data Frame from corefxlab to complement what Deedle lacks. Since Deedle's original core engineering team has left and moved on, some core technical debt are beyond my capability to address. But Deedle is still the best data frame library in .Net ecosystem for now if you are able to get around some of its odd edges.
What kinds of nontrivial samples would you be interested in seeing?
(hint to Jupyter newbies: alt-enter re-evaluates the code block after editing it)
The best way to work with DUs is pattern matching. Using the previously-defined type definitions, we can model withdrawing money from an account:
Then a sample that has nothing to do with withdrawing money from an account.Plus some of those APIs that made the cut into .NET Core, like the UI ones, are Windows only.
Things that work when compiled, worked about a decade ago, but die on FSI preventing multiple data interaction & scripting scenarios. Ref: https://github.com/dotnet/fsharp/issues/3309 , https://github.com/dotnet/fsharp/pull/5850 , https://github.com/fsprojects/IfSharp/issues/206 , etc etc
Specifically, referencing a library that uses an FSharp.Data type provider, and calling that library in FSI in Visual Studio.
Updated to the latest and greatest, 5 y.o. code looks like this when we execute it:
> System.MissingMethodException: Method not found: 'System.String FSharp.Data.Http.RequestString(System.String ....