So in a traditional 2D heatmap, we discretize the space into squares or bins, then color the bin "proportional" to the number of points that fall inside the bin.
If we're making a 2D heatmap of a timeseries, though, we don't really want to count up the discrete points that make up the timeseries because our binning may not line up with our timeseries sampling frequency. Especially not if we want a very high-resolution heatmap, where each bin is sized to be 1 pixel.
So the solution isn't to count the number of points (from the timeseries data series) in each bin, but rather to count the number of lines that get projected/drawn into each bin (pixel).
But we need to go a step further... if we think of each bin as a hypothetical square, the shortest line through a square is a vertical or horizontal line. And the longest line through a square is a diagonal line. So if we want to represent "how much" of the line goes through a square, we need to measure its slope. Hence, derivative.
So by 2D heatmapping lines instead of points, we're not just counting how many lines fall into each bin, but we also need to weigh each bin-line occurrence by it's arc length, which we use the derivative as a proxy for.
Neat! (I think...)