Apache Zeppelin
zeppelin.apache.org
zeppelin.apache.org
bqplot [1] for example is great for 2D dataviz, very responsive and updates real-time. Based on D3 I believe. Usually I can do what I want with base widgets and bqplot and the result is pretty.
ipyleaflet is another popular library for maps.
I especially enjoy using them with voila [2] to create an app, or voici [3] for a pure-frontend (wasm) version.
If you want to develop a widget, the new-ish anywidget library can reveal handy [4].
For an example, see this demo [5] I made with bqplot and voici, that visualizes a log-normal distribution.
[0] https://ipywidgets.readthedocs.io/en/stable/
[1] https://github.com/bqplot/bqplot
[2] https://voila.readthedocs.io/en/stable/
[3] https://voici.readthedocs.io/en/latest/
[5] https://horaceg.github.io/long-tail/voici/render/long_tail.h...
I would add two more:
1. VizHub [1]: for D3 based visualizations. I have not tried it, but I have watched some D3 videos [2] by its creator Curran Kelleher who uses it quite a bit (oh, and a shout out to the great D3 content he has!).
2. This is slightly unusual but I have recently been using svelte's REPL notebooks [3] to try out ideas. Yes this is for svelte scripts, but you can do D3 stuff too. And on that note, svelte (which is normally seen as a UI framework) can be used for pretty interesting visualizations too, because how it can bind variables with SVG elements in HTML (you can get similar results with React as well). For ex., here's a notebook I wrote for trying out k-means using pure svelte [4]. Be warned: fairly unoptimized code, because this was supposed to be an instructive example! On a related note, Mathias Stahl has some content specifically for utilizing svelte with D3 [5].
[2] https://www.youtube.com/watch?v=_ByiP7KM0So
[4] https://svelte.dev/repl/1689f5c3699640ff86d9bd6a04ac8272?ver... Note that the "Iterate!" button iterates once; keep clicking it to move things along.
any idea what "BQ" stands for in BQplot? I find that I am able to remember and recall tools and terms that I actually understand the full forms of :)
Here's our repo: https://github.com/marimo-team/marimo
Edit: looks like Mercury (A jupyter extension) has them: https://runmercury.com/docs/input-widgets/
It is quite telling how long industry takes to adopt cool ideas, while rebooting some bad ones all the time.
Didn't work out all that well for a number of reasons.
The most important thing is, users are used to Jupyter. Zeppelin's ui is very different, and most people are not willing to jump on yet another learning adventure just for the sake of it.
Then, it's not as widely adopted and supported as JupyterHub- with JupyterHub you can easily integrate whatever you want to. Want several simultaneous jupyters for each user? Sure. Want separate quotas, different k8s namespaces for user groups? Easy. A shitton of plugins? Here you go. A selection of different images for each user, depending on the tooling required? Welcome.
Third thing is really unfortunate, but Zeppelin proved to have a less than stellar stability and performance, at least in my experience. People are wary of something that's often unreliable.
So I've finally decided to just go with JupyterHub, and users can't be happier. Everything's fully customized, things are smooth and familiar to a non-dev crowd.
Another, and in some ways, better solution would be to go with vscode, but I doubt a typical analyst/ds would prefer vscode, at least for now.
All in all, I don't see a place for Zeppelin- it can't compete with what's already on the market and yet doesn't bring anything new and worthwhile.
Toree is mostly dead but might also get a Scala 2.13 release now that Spark 4.0 is approaching.
P.S. I was committer there until changed job.
Not all of them get that much love, but often they have pretty nice functionality.
I still remember that setting up Apache Skywalking was one of the easier ways of getting some APM and tracing in place, compared to the other options out there.
And, of course, the likes of Apache2 and Apache Tomcat are also quite useful in some circumstances.
Sometimes I do worry about the long term survival of the ASF. Many projects are largely supported by 1 person. A lot of projects are mostly abandoned (but not yet moved to the attic). Many others suffer from the blight of "what the hell is this for?", where their website is so vague that it might as well not exist.
The community was especially helpful, responsive, and patient with my limited understanding of their tool.
In the end it was the most stable part of the overall project operationally speaking.
Zeppelin does make it easier to run Scala Spark, I find, but Scala Spark usage has declined rapidly.
I agree that Jupyter for PySpark makes more sense in almost every use case. We made the switch as an org about 2 years ago and haven’t looked back. Jupyter has its own issues but does feel much usable by just about every metric.