A Tiny Grammar of Graphics
observablehq.com
observablehq.com
https://byrneslab.net/classes/biol607/readings/wickham_layer...
Wickham is the Chief Scientist at RStudio and created R packages such as ggplot2 and the tidyverse.
The original implementations go back to SYSTAT and SPSS GPL (Graphics Production Language).
GPL especially, with its statement-based approach, has arguably better ergonomics for interactively and iteratively producing plots compared to function-based approaches.
(It's been a few years, maybe I should take another look at the book.)
> ggplot2 is a system for declaratively creating graphics, based on The Grammar of Graphics.
Springer runs occasional sales up to 40% discount about once a year, but a don't recall if "The Grammar of Graphics" was eligible last time.
Leland Wilkinson (GoG inventor) and I designed it together a couple of years back.
The function for creating marks (a layer) tries to be as "flat" as possible, in the sense that it should be possible to render most common kinds of plots without having to pass nested/hierarchical options: https://wave.h2o.ai/docs/api/ui#mark
Grammar of graphics always was this North Star for me. It is very helpful to go through papers and books and search for inspiration how to organize your system. But direct implementations are finicky to work with. And in my hubris I attempted to write yet another one implementation of grammar of graphics and it resulted in exactly the same problems! With complex marks it is ambiguous what is a data point and what is a series. Tuning looks require this configuration objects scattered around chart definition and composition sometimes require to inject something in two different parts of definition.
Now I treat grammar of graphics as this collection of patterns and good practices. But surrender to pragmatic solutions when necessary.
Anyway I think I owe big part of my career to works of Wickham and Wilkinson.
I sort of agree with you. I've implemented the Grammar of Graphics from the ground up four times professionally (!), twice in collaboration with Leland, all in different products.
The main reason why it might be finicky to work with directly is that the point vs. series vs. series-of-series distinction can run arbitrarily deep, so there's some mental gymnastics involved on the part of the library's user how to refactor the data and present it correctly so that the library can do its thing.
Tableau, which is also a GoG system, sort of deals with this by having slots for "Dimensions", "Pages", "Color", etc. as proxies for multi-level aggregation ("group/slice/dice" in BI terms). So even though it's not immediately apparent to new users how to present data correctly to the rendering system to get the kind of vis they want, at least it's pretty low-friction UX to shuffle variables between those slots till satisfied.
With programmatic use, that shuffling-around gets cumbersome because now you have write code to munge data into submission.
Tableau introduced the "Show Me" feature precisely for this reason - most new users would rather get stuff done quickly than figure out how the GoG can best solve their vis problem.
Javascript data products fall into a weird gap of having the best visualization tooling and the worst data manipulation tooling. I know there are efforts like arquero and DuckDB which make data more accessible, but there’s no really strong scipy/numpy/statsmodels/scikit-learn equivalent.
For your situation, is it possible to separate the tasks: Using Python to handle manipulating the data, then JavaScript to display it?
https://observablehq.com/collection/@skybrian/digital-signal...
That said, I’m weird. My blog is artisinally hand crafted HTML, JS, and CSS for the same reason.
I built an experimental system to do just that - design the layout and all data visualizations in a single Sketch/Figma document, export to SVG, then map data to the SVG elements in the browser. It's all 100% declarative.
Here's an over-the-top example: https://youtu.be/S9cmi89fvT8
There are no limits to what you can achieve with this approach, but obviously needs Sketch/Figma skills (or a colleague who does).
edit: typo
There are a few primitives that d3 uses, and once you implement those you can produce easy SVG results. `Scale` is the most important, for mapping your x,y plane into SVG pixel coordinates. Then some of the ticks helpers can be handy too.
Writing the SVG directly is just the fastest way to get what you need.
http://www.geo-informatie.nl/courses/grs60312/visualisation/...
This applies to any charting library that forces you to provide both spec and unaggregated data to memory/cpu constrained clients (e.g. Javascript in the browser). This is done for implementation-simplicity (Vega, for example), but obviously doesn't scale to larger datasets.
I've implemented a system where the data part of the spec is munged in-database, and aggregated data is provided to the browser, along with hints for axes, scales, legends, etc. It requires a part of the GoG interpreter to be resident on the server-side.
It looks like with later versions they switched to kind of a hybrid approach (part-remote, part-local) with Hyper to reduce latency for interactivity.
> there is no sharing of the underlying data set for multiple projections across the same large data set
But that would require some kind of open standard for portability, no?
I like the approach AGGrid uses - they provide a viewport based interface that the grid uses to display data, and you can implement that interface on top of your data model - https://www.ag-grid.com/javascript-data-grid/viewport/. Unfortunately it's only available in their enterprise version, but this approach scales to both grid and chart based UIs. D3 has a bit of that flavor as well, since you can map visual attributes into your underlying data any way you'd like.
The Ag Grid approach makes sense if data and vis need to be wired together programmatically.
Luckily for me, the main purpose of the lessons was not so much about how to build a charting tool, but rather concentrated on how to break the code into modules in the hope that some of the modules could be reused in other, similar projects.
If I'm making obvious mistakes in the approach, or code, that I set out in the lessons then feedback is always welcome so corrections/improvements can be made to them!
[1] - Building the chart frame, code management, etc - https://scrawl-v8.rikweb.org.uk/learn/eighth-lesson/
[2] - Generate bar charts and line charts from crime data - https://scrawl-v8.rikweb.org.uk/learn/ninth-lesson/
[3] - demo of the final code - https://scrawl-v8.rikweb.org.uk/demo/modules-001.html
https://github.com/h2oai/lightning/blob/master/src/lightning...
It's used in H2O: https://github.com/h2oai/h2o-3
For me GoG is basically "use ggplot2 instead of base R plotting libraries when in R".
Which is fine and all, but graphing utilities of matlab/octave are far simpler/flexible.
ggplot2 to me seems like it feels it has to add a layer of complexity to achieve the "grammar" part, without a real practical benefit over matlab etc.