Time Series Density Plot
observablehq.com
observablehq.com
I'm not sure it's not quite right either, because the weight of each series in a bin should depend on the derivative but it doesn't look like it does.
This sort of plot always has a conflict between showing the distribution of each time domain sample and showing the time behavior of each curve. For the former, showing quantile contours may be clearer (though not as pretty!). For the latter different colors with transparency might be better, though it becomes unintelligible at some point.
(This is mentioned in the intro post for this library: https://observablehq.com/@twitter/density-plot-introduction?...)
I suppose one can incorporate uncertainty on f(t) by having g(y,t) = gausy(0,\sigma(t)) * \delta(y-f(t)) (which would also effectively antialias in the few-lines case).
Also I guess you have to be careful not to plot a Weierstrass Function!
So in a traditional 2D heatmap, we discretize the space into squares or bins, then color the bin "proportional" to the number of points that fall inside the bin.
If we're making a 2D heatmap of a timeseries, though, we don't really want to count up the discrete points that make up the timeseries because our binning may not line up with our timeseries sampling frequency. Especially not if we want a very high-resolution heatmap, where each bin is sized to be 1 pixel.
So the solution isn't to count the number of points (from the timeseries data series) in each bin, but rather to count the number of lines that get projected/drawn into each bin (pixel).
But we need to go a step further... if we think of each bin as a hypothetical square, the shortest line through a square is a vertical or horizontal line. And the longest line through a square is a diagonal line. So if we want to represent "how much" of the line goes through a square, we need to measure its slope. Hence, derivative.
So by 2D heatmapping lines instead of points, we're not just counting how many lines fall into each bin, but we also need to weigh each bin-line occurrence by it's arc length, which we use the derivative as a proxy for.
Neat! (I think...)
Imagine a point moving along the curve, depositing a constant amount of density/ink onto the canvas per unit of time. When the point is moving quickly it deposits less density, and when the point is moving slowly then it deposits more, since the density deposited per time unit is constant and the point traverses less space when it moves slowly.
You can think of each vertical strip of bins as representing a unit of time. The discrete approximation to arc-length normalization means, for each time series, making sure that it contributes a single unit of density per vertical strip: if a curve goes through one bin in that strip, then that bin has its density increased by 1. If it goes through 3 bins, then the density in each is increased by 1/3.
That makes sense. Thanks!
Expected a bit more.
If you're interested in the technique behind the plot you can read more about how that works in this paper: https://arxiv.org/abs/1808.06019
I wish more (hosted) monitoring systems would shift to this type of visualization, since computers are fast now (I'm told!)
Traditional graphics pipelines aren't built to handle more than 256 incremental levels of opacity since it's an 8-bit channel, and when you have a thousand time series you can no longer distinguish between levels of density. Conditional information (following a single line) becomes difficult as well since there are so many lines. Density plots solve the overdraw problem by accumulating density into an offscreen buffer before rendering, which allows the color scale to be tuned to the amount of overlap in the specific dataset. Interactive selection/hover techniques can be used to recover conditional information.
This example is part of a density plot library published earlier this week (https://observablehq.com/@twitter/density-plot-introduction?...) and one of the references there is to an opacity-based approach I've seen used for 2D point clouds, though I haven't seen an analogous demo for curves: https://observablehq.com/@rreusser/selecting-the-right-opaci...
The reduceOp caught my eye in your API. What uses have you imagined/tried for changing the density accumulation function from x + y?
When the number of the bins in the data is not an exact integer multiple of the number of bins in the histogram, adjacent data bins can get mapped to the same histogram bin, resulting in e.g. 2x the data volume in some rows/columns of the histogram.
When a bit of loss of fidelity is acceptable the two solutions I've used are to render the histogram at an exact factor of the number of data bins then set the canvas dimensions to the desired size (relying on the browser to downsample the resulting image), or to use `max` as the reduceOp and render the histogram directly at the intended size.