Interesting ideas in Observable Framework
simonwillison.net
simonwillison.net
It brings together d3, Observable, Observable Plot, HTL and layers on a bunch of new ideas as well.
Today I was beginning to look at how to host a static Jupyter Notebook, or hosting it interactive with WASM.
But actually I think that for most of my purposes Observable Framework will be a better fit.
And it's not like d3 is easy to use so that you can use it without examples, specially considering that changes between versions are often incompatible.
But apart from this, there's a lot of incredible graphics to find on the site.
I've had this gripe more with ObservableHQ notebooks -- great examples and a pointless resource all at the same time.
This framework effort seems to be a bit more open though (at least you can self-host), so I'm keeping an eye on it.
i, embarrassingly, haven't tried it
[0]: https://observablehq.com/@bumbeishvili/convert-observable-co...
In this case, Observable Framework is an open-source static-site generator that runs on your machine and uses standard Javascript, so really don't understand what's the connection to specific hosted Observable notebooks.
(But I guess it's of some interest/disappointment to the Observable folks that at least some people are encountering their Observable project mainly in the context of looking for d3 examples.)
I tried out Observable Framework and built a little interactive plot (https://github.com/willmeyers/observable-ssta). It was incredibly easy to setup and get data plotted.
My only gripe is that I wish you could configure Python data loaders to use virtualenvs.
I just created a python project and then instead of `yarn run dev` to start the dev server, just run `poetry run yarn run dev` so the python is executed within the virtualenv.
This setup also lets you use a custom python package to define reusable and unit-testable code for the dataloaders that you can import into the *.json.py files to keep those really simple.
But if you want dynamically generated data at build time and want to make use of Observable’s dataloader automatic execution of data/*.json.py, for instance, while still maintaining a custom virtualenv for the project rather than your system python, you’ll need some way to specify that virtualenv’s interpreter while observable executes the build for the dev server or the full dist/ output.
So for both options it’s largely a matter of taste. I personally like using the poetry virtualenv because it’s simple to manage dependencies and the venvs in one tool, while letting me use observable’s dataloaders with third-party or custom python packages. It sounded like the parent comment wanted to use this type of approach so I focused it to that scenario specifically. I like the simplicity of the single command to generate the data and build the site.
It's honestly been really wonderful. Learning all of those tools has taken some significant energy, and I'm missing some functionality I'd love around parameterizing my data generator, but the final notebook is beautiful and functional.
Using markdown and reactivity makes notebooks like this actually feel usable. Jupyter's custom format made version control a giant pain and without reactivity your iteratively designed notebook easily becomes a write-only, stateful mess. I've also tried making this work using Quarto and their Observable integration and it was hacky and piecemeal.
Genuinely, this was the first time I've been pleasantly surprised and excited to write a notebook and share it with others. I'm sure there will be more sharp edges, but it's become my first choice notebook tool after this project.
Asking between Python and R tends to get people to throw around opinions.
> Everything in a code block with the js content hint will be executed in the users browser immediately. If you want to show the code you have to hint 'js echo'
Am I the only one thinking that it would have been better for backwards compatibility if it where the other way around? I.e: having an opt-in code hint like 'js exec' that runs code in a user's browser and leaving the widely used 'js' hint alone? The way this currently is set up, you cannot integrate that renderer in an existing app without having to manage where it is allowed to run.
Here is a mermaid example:
````
```mermaid
Code here
```
````
Copy that into a Markdown file to try itI love kotlin and tried creating a data loader for a kotlin script but that had some rough edges. Kotlin expects script files to be named foo.main.kts but observable expects executable shebang loaders to have a foo.exe extension. So I created a proxy exe script to call the kotlin script, but it then doesn't trigger auto reloads of the data.
A bit of friction compared to marimo or jupyter is using variables between data loaders and the notebook. For example, I want to use the date picker view component to change the range of data fetched by my loader. It's not clear how to do that, so exploratory analysis is slowed down a little. I'm aware this goes against the paradigm but just wanted to point it out. It ends up with you potentially moving a lot of the data munging to the notebook as you explore, which isn't ideal from a performance perspective.
One last thing is I wish you could define dataloaders inline. I'm a big fan of single files, so being able to just add a python code block and let Framework extract that as a file would be a nice little QoL improvement.
Still the early days, but Framework seems promising! I'd love to have my all my markdown notes running through it to get a sort of org-mode type situation without going full emacs.
As for inputs-driving-data-loaders, that does go against the grain a bit since Framework favors static data snapshots so that the built site is self-contained and performant. But a technique that works well is to generate Parquet files in data loaders representing the superset of data that you want to interact with, and then using DuckDB/SQL in the client to extract the subset you want to visualize. This tends to perform well, though obviously it’s dependent on the size of the superset you want to interact with.
But I didn't try the new Observable Framework - interesting to see similar examples where it queries a database live. I hope that preloading and caching all the data is not the only option because these types of apps should be interactive. Ideally, it should expose SQL for live editing.
It's of course the de-facto language for interactive display in browsers. The use case for dashboards and data visualisation is clear. But it's an awful language for data science and data analysis, compared to Python or R.
This is it, more or less.
It is far, far easier to build an app like this where you want a plethora of users as a web application than a native one, for instance.
For anything JavaScript as a runtime / language is missing, WASM can boost as well. For math and data science, WASM is a natural choice for any missing pieces
So you absolutely can do the data processing step in R or Python and have that output JSON or CSV which is then visualized at the end using JavaScript.
Not a small feature, but I bet it would be possible to use WebAssembly to add support for Markdown blocks that get executed in other languages as well, using Pyodide for Python for example.
I myself like the self-contained aspect of it, since I can publish static files. Also, D3 is the pioneering library for data viz on the web. Especially with maps, which is what I've used it for back in the day. Time to refresh those skills.
I wonder if one could combine this and reactive Python based Jupyter notebook alternative https://docs.marimo.io/guides/wasm.html
Perhaps with web component packaging it should be doable. Web component attributes might allow tying reactive events from one side to the other.
I don't see the advantage
Demo: https://simonw.github.io/observable-framework-experiments/pa...
Source code: https://github.com/simonw/observable-framework-experiments/b...
My only gripe is that data loaders don’t seem to support Parquet files, which is really annoying.
There’s an interesting possibility here where you can have large datasets in Parquet, exposed via HTTP, whilst being generated at build time with all the benefits that gives you (being able to read only specific columns, filtering via row group statistics etc). Not dissimilar to the “SQLite over http” WASM demo I guess.
Because right now I need to take my large, nicely compressed dataset and export it as either a CSV or a zip file? And then the browser needs the entire thing, even if I’m just viewing a subset of the data? Which is much bigger and much slower than it needs to be.
https://github.com/observablehq/framework/blob/main/examples...
Using data from here: https://github.com/observablehq/framework/tree/main/examples...
Rendered version here: https://observablehq.com/framework/examples/api/
If you want to run a Data loader that outputs to parquet there are plenty of ways to do that - I would suggest a Bash or Python script that wraps DuckDB.
That would be a pity, because Quarto is really good. I haven't tried Observable yet, but in outline they have some similarities:
1. Documents written in Markdown
2. Ability to embed code blocks in the Markdown, with code executed when the document is rendered.
3. Ability to embed output of the code blocks in the rendered result (e.g. tables, charts).
4. Ability to render to multiple formats (pdf, static site, ...).
Quarto supports Python and R as languages in the code blocks (maybe more, not sure). I personally prefer it to Jupyter notebooks because the source is plain text so (1) there's a choice of editor and (2) moving between text and code blocks is seamless.
I can't say Quarto is better than Observable but it is good. It has depth from its history in RMarkdown (like rendering mathematical equations, naming & cross-referencing).
It's certainly worth consideration for anyone looking for a "code notebook" solution.
I don't see Framework changing things there - if anything the ISC license should make it a better partner for D3.