Histogram vs. ECDF
brooker.co.za
brooker.co.za
It's basically saying "yes, given this data I am very confident that there is an underlying difference that is not just an artefact of random sampling".
Still, if you have 5-50 samples from an unknown distribution I think the KS test can be a very valuable protection against false-positive i.e. seeing patterns in random noise.
Even more so if you are making multiple comparisons, e.g. comparing multiple response variables, since you can easily make a Bonferroni correction to avoid "the xkcd 882 problem" which is hard to do by eyeball.
So I see a lot of utility in the KS test (and Wilcox, etc) when measuring computer systems on multiple metrics with unknown distributions (often not Gaussian) and where samples cost ~$0.10-$1.0 each. I see a lot of false positives in the absence of such tests, or otherwise generally low confidence in inferences based on e.g. eyeball/mean/median.
If I care about 'were these two datasets generated by the same distribution' then I'll step back and ask 'why do I care?' and 'what will I do with the answer?' and then use a more specific test.
I mention this because histograms, especially HDR histograms, are a very compact way of measuring distributions, and it's nice that you can keep those benefits and still convert to an eCDF.
Histograms are nice in that they effectively compress non-trivial datasets (at least those that have a reasonable bounded domain) to something quite manageable.
I guess there is nothing stopping you from doing the same thing here, but it does kind of discount the author's claim of not being able to go between histogram and eCDF.
Am I missing something?
What is a completely different thing from the arbitrary bucketing for histograms. CDF doesn't go to zero or becomes misleading if you bucket it wrong. You just lose the details.
If there are duplicate sample values, you can still store a sorted list of (sample,count) here and generate either a histogram OR an eCDF, or any other plot really.
Effectively it is not a fair comparison to compare the two methods since they both have storage tradeoffs that are not really discussed.
The entire discussion is about the quality of the information the plot communicates to you. Histograms can be completely misleading, CDFs can't. Finding a bucketing so that an histogram communicates the background data is a non-trivial problem, for CDFs, it's not a problem at all.