How odd is a cluster of plane crashes?
bbc.com
bbc.com
For example, you could mean: picking tracks with `putting back'. You could shuffle tracks like you shuffle a deck of cards, ie picking without putting back. It's still random, but with a different distribution.
Even picking the first track with 90% probability, and the rest with some other smaller probability is random. It's just not a uniform distribution.
So I would anticipate the clumping to even out as the sample size increases. Or maybe the evening out is a large collection of clumps. Anyone know if this is the correct way to think about it?
Clumping won't even out. A useful way of thinking about it is in terms of the distance between samples (e.g. a radial distribution function). If you don't have clumping, that distribution of distances is decidedly non-random and has peaks. Additionally, there's no specific length-scale for the clumping to happen on, it's essentially fractal in nature.
(Happy to be corrected if I'm wrong on this.)
Are you talking about a random distribution of points on a line / square, or about rolling dice?
Right, exactly. I made my comment in reply to the grandparent problem talking about dots on a screen. For a discrete distribution like a dice roll in the parent post it'd be a different matter. (I was going to edit my comment to clarify, but the edit window had expired.) Which serves to illustrate why it's important to be precise when talking about what distribution you're sampling.
Also, trying to run this experiment in practice is also problem even for dots on a screen, because eventually the finite display size of your dots will become an issue, as will floating point arithmetic (since you only really have a finite number of floats in a given range).
> Also, trying to run this experiment in practice is also problem even for dots on a screen, because eventually the finite display size of your dots will become an issue, as will floating point arithmetic (since you only really have a finite number of floats in a given range).
Yes, but that's not a problem for the theoretical analysis of fractal dimension.
Uniformly placed dots do have a characteristic density and a relatively simple distribution of nearest neighbour distances. Even the clumpiness doesn't lead to scale invariant lumps.
The lumps get absolutely bigger on bigger scales, but relatively smaller and smoother.
It would be useful to do the analysis properly, though.
It would probably true for that too but as another comment says, we don't want to have an increased sample size here!
Are you saying we need to crash more planes? That seems a little harsh just to get crashes coming at a consistent rate. ;-)
Do the experiment you decribe and look at the numbers.
var results = [0,0,0,0,0,0];
for (var i = 0; i < 6000000; i += 1) {
results[Math.floor((Math.random() * 6))] += 1;
}
console.log(results);
The results are: [1002140, 999355, 1000009, 1000401, 1000014, 998081]
Basically, they are all within 1% tolerance of each number getting 1 million occurrences.So when you look at the clumping on random dots on a screen, given a large enough sample size, I would anticipate that the probability of keeping a clumped visual distribution would become quite improbable.
For example, a naive me would generate: [3,4,5,1,2,4,2,1,2,6] whereas a true random distribution might generate: [3,4,4,4,4,1,2,2,2]. When you generate 6 million samples you would indeed have about 1 million per number. However you would still have subsequences of the same number.
The problem is when you get less samples than slots. For example if you roll 3 dices, you should get "1/2" hits for each result. So the question is: How probable is that one result is repeated? (i.e. you get two equal dices out of the three)
[spoiler alert]
All equal: 3%
Two equal and one different: 42%
Three different: 55%
"Intuitively", it's expected that with a uniform distribution you get something like 1% + 9% + 90% instead (this numbers are completely made up). Most people don't expect that the probability of getting a repeated dice is so close to the probability of getting all of them different.
Just wanted to add that a few years ago (when you wrote your first computer program) a randomly selected random number generator was more likely to be less random than a randomly selected random number generator today.
Edit: I wrote a program[1] to fill random pixels based on rand/modulo and mersenne twister/uniform distribution and I could not tell the difference in the generates images.
http://bl.ocks.org/mbostock/fe3f75700e70416e37cd
>Uniform random is pretty terrible. There is both severe under- and oversampling: many samples are densely-packed, even overlapping, leading to large empty areas. (Uniform random sampling also represents the lower bound of quality for the best-candidate algorithm, as when the number of candidates per sample is set to one.)
Taken from: http://bost.ocks.org/mike/algorithms/
Is this not the gambler's fallacy? Shouldn't the odds be unaffected by a previous crash?
We can reduce the rate of crashes by taking planes out of the sky, and one way to do that is to crash them.
The gambler fallacy would be "I have seen 9 days without crash, on the 10th day there is higher odds of crash". But you are not asking this, you are asking: "What is the chance that in the next 10 days, exactly first 9 days would be no-crash and on the 10th day there is a crash".
EDIT: Also, another view:
If you have seen 9 blacks in a row, there is an even chance 1/2 (roulette without zero) that the next will be black or red. Because in all sequences of length 10 starting with 9 blacks (this is given, you observed this), there is equal number of those with 9 consecutive blacks and a red and of those with 9 consecutive blacks and a black. This is gambler fallacy, believing that after repeated colour, you have higher chance of the other colour.
Here, you are asking: how many of all sequences of length 10 have exactly 9 blacks and 1 red. And there are all kinds of sequences with 1 red in first 9 places, 2 reds, 3 reds, etc., so the odds of observing the 9 blacks in a row in the first place before asking about the 10th spin is lower, but it does not change the odds of the 10th spin itself.
I had tried the Algolia search box from my tablet and got no results. Using Google (constrained to HN) brought up today's discussion and the one from last year.
site:news.ycombinator.com plane crash cluster