Do you happen to know which they are? This has intrigued me for a long time but when I asked Wheeler he was skeptical of doing it that way.
Do you happen to know which they are? This has intrigued me for a long time but when I asked Wheeler he was skeptical of doing it that way.
In the past I've also tried the trick of taking the logarithm of values to convert a power distribution into a more normal-ish distribution. My project was to see if I could alert on an increased rate of uploads to RubyGems, signaling a possible malware campaign. Time between uploads follows a power law. I used the logarithm trick to convert it -- though not successfully. There are actually multiple distributions -- the "random" one which is prominent, but also other peaks centered around 1 second, 2 seconds, 3 seconds etc. Basically the detection was made difficult because of mass uploads being rate-limited on the uploader's side.
My current simplistic (and very dumb!) solution that I've used for power-law type distributions — like HN virality, for instance — is to count the number of days between viral events, and then subject that to process control.[1] I basically take Wheeler's approach to chunky data and use that for J-curve type data, which tells me if the behaviour of my 'HN virality process' has changed.
I'd be very interested to learn of other approaches.
[1] HN traffic for commoncog.com displays routine variation most weeks with an Upper Process Limit of 192 and a Lower Process Limit of 0, unless one of my articles hit the front page, at which point I get 11-16k additional uniques).
I did forget to bring up the Poisson approximation you mention though. I'll include that too.
One possible way out is to look for measurements that contribute to running time but which are not affected by other factors. I remember the YJIT folks talking about using CPU instruction counters, but I can't find it on the benchmark website.