(This is mentioned in the intro post for this library: https://observablehq.com/@twitter/density-plot-introduction?...)
(This is mentioned in the intro post for this library: https://observablehq.com/@twitter/density-plot-introduction?...)
So in a traditional 2D heatmap, we discretize the space into squares or bins, then color the bin "proportional" to the number of points that fall inside the bin.
If we're making a 2D heatmap of a timeseries, though, we don't really want to count up the discrete points that make up the timeseries because our binning may not line up with our timeseries sampling frequency. Especially not if we want a very high-resolution heatmap, where each bin is sized to be 1 pixel.
So the solution isn't to count the number of points (from the timeseries data series) in each bin, but rather to count the number of lines that get projected/drawn into each bin (pixel).
But we need to go a step further... if we think of each bin as a hypothetical square, the shortest line through a square is a vertical or horizontal line. And the longest line through a square is a diagonal line. So if we want to represent "how much" of the line goes through a square, we need to measure its slope. Hence, derivative.
So by 2D heatmapping lines instead of points, we're not just counting how many lines fall into each bin, but we also need to weigh each bin-line occurrence by it's arc length, which we use the derivative as a proxy for.
Neat! (I think...)
Imagine a point moving along the curve, depositing a constant amount of density/ink onto the canvas per unit of time. When the point is moving quickly it deposits less density, and when the point is moving slowly then it deposits more, since the density deposited per time unit is constant and the point traverses less space when it moves slowly.
You can think of each vertical strip of bins as representing a unit of time. The discrete approximation to arc-length normalization means, for each time series, making sure that it contributes a single unit of density per vertical strip: if a curve goes through one bin in that strip, then that bin has its density increased by 1. If it goes through 3 bins, then the density in each is increased by 1/3.
That makes sense. Thanks!
I suppose one can incorporate uncertainty on f(t) by having g(y,t) = gausy(0,\sigma(t)) * \delta(y-f(t)) (which would also effectively antialias in the few-lines case).
Also I guess you have to be careful not to plot a Weierstrass Function!