Sure there will of course be some unlucky articles, but does that actually matter?
Sure there will of course be some unlucky articles, but does that actually matter?
Quote: "The least viewed article in the sample, Erygia sigillata, has a page_random value of 0.500764585777. The article Katherine Hanley is right on its tail with a value of 0.500764582314, which is just 0.000000003 less, or 3e-9 in scientific notation. This is 98% smaller than the average random gap. In other words, Erygia sigillata is an extremely unlucky article as far as the “Random article” button is concerned! It’s 50 times less likely to be landed on than an average article."
It’s only once your supply of numbers runs low that differences will start to equalize, reaching 1 when you have exhausted your supply.
Besides, the author isn’t really making an argument. They are giving you actual data showing the differences to the next lowest number. It’s hard to argue with that.
> the gaps will differ because there is no mechanism that would make some number with close neighbors less or more likely to be drawn than any other remaining number.
When a new article is inserted, there is a higher probability it will be inserted in a large gap than a small gap, so it should balance out.
I suppose you're right,i am responding to an implied criticism to the randomness method that the author didn't make. They just offered it as explanation.
That’s a good point and I’m not entirely sure why it (appears to) not work that way. Maybe it’s because that interval has a higher likelihood, but there is no preference for numbers towards the middle, that would dissect it into (roughly) equal parts?
More interesting question would be what is the standard deviation (of gap size), not what is the worst outlier
Differences around zero is actually preferred: https://mathworld.wolfram.com/UniformDifferenceDistribution....
Now think about a random-gap list with five million entries. With that many entries, will the gaps balance out? In the best case, five of them end up in the range from 0 to 1 millionth, five of them end up in the range from 1 millionth to 2 millionths, etc. But we've already seen what it looks like to have five uniform random numbers in a range (whether it's big or small doesn't matter); the gaps tend to be really varied. So we're going to get this sort of imbalance between gap sizes (viewed as a ratio) no matter how many entries we insert.
One other thought experiment: the article mentioned a page with a gap size of one billionth. How many pages will it take for that to balance out so that page doesn't have an unusually small gap any more? How many pages does Wikipedia have?
(This is similar to the reason that infinite space packed with marbles has the same packing density as infinite space packed with bowling balls.)