(Better yet, just give me the raw data so I can analyze it myself. I find it hard to blindly trust someone else's conclusions considering all of the p-hacking going on nowadays.)
(Better yet, just give me the raw data so I can analyze it myself. I find it hard to blindly trust someone else's conclusions considering all of the p-hacking going on nowadays.)
The system would then load the data from S3 back into the live system via Hadoop. Turns out it was pretty cheap to store highly compressible files in S3.
(Do NOT, for the love of your sanity, emulate this design exactly. Use Backblaze or Google's nearline storage. Do not touch Glacier at all if you can avoid it. When I wrote the Glacier integration, I did it because it was far cheaper than its competitors. That's no longer the case.)
FYI I don't think its right to say "forget means, medians, percentiles, etc." They are hugely useful.
You just have to note that they aren't everything. As you've noted, only the raw data is everything.
If you're saying that people should provide the full data along with their analysis, I fully agree! But if you're saying "skip the analysis", then unfortunately I, and most people, are not statistically educated enough to do all the analysis ourselves.
Always keep in mind that estimating distributions from samples is very hard, in particular for multidimensional and continuous variables. Sometimes even infinite samples do not suffice (I am not kidding).
The good news is that there are interesting and relevant tasks (for example classification) where you do not need to know the distribution to solve the task optimally, although if you somehow managed to do so (find the distribution) you would be able to solve the task optimally. Or more succinctly, knowing the distribution is sufficient but not necessary for such tasks (small mercies)