In search of the least viewed article on Wikipedia
colinmorris.github.io
colinmorris.github.io
So this sentence made me wonder why he didn't actually just go do it, 6M pages isn't really all that big of a data set. Turns out that it's a problem of how the data is arranged. The raw files are divided by year/month/day/second[0], and then each of those seconds is a zipped file of about 500MB in size, where pages are listed like this:
en.m Alcibiades_(character) 1 0
en.m Alcibiades_DeBlanc 2 0
en.m Alcibiades_the_Schoolboy 1 0
en.m Alcide_De_Gasperi 2 0
en.m Alcide_Herveaux 1 0
en.m Alcide_Laurin 1 0
en.m Alcide_de_Gasperi 1 0
en.m Alcides_Escobar 3 0
en.m Alcimus_(mythology) 1 0
with en.m getting a separate listing from the desktop em, and all the other country codes getting their own listings too. So just collating the data would be a huge job.
The API also doesn't offer a specific list of all pages by time, so you'd have to go and make a separate call for each of the 6M pages for a given year and then collate that data too.
[0]for example, https://dumps.wikimedia.org/other/pageviews/2019/2019-01/
1. Reject the intuitive definition of "interesting"ness as being ill-formed. We need an objective definition non-self-referential definition of the word interesting.
2. Accept that there are interesting or uninteresting numbers, but their status as objects of interest is tied to a specific time. For example, 1729 might have been an uninteresting number until the conversation between Ramanujan and Hardy. Several of the examples on the Wikipedia page have such temporal limits on when they were / were not interesting.
As an odd extension of this idea (I don’t know of an English equivalent), if this were Spanish we would probably use “estar” instead of “ser” to affix the description.
Example: https://trends.google.com/trends/explore?date=2022-01-01%202...
There was interest in "Ukraine" at a certain background level until events in January 2022 began an increase of interest. Then the invasion in February drove it through the roof. Interest in Google Trends peaked specifically on 24 Feb 2022. However, after the initial invasion, interest steadily and rapidly waned. By 24 Mar, interest had waned to 10% of its invasion peak. By 24 April, 5% of that peak. And 3% by 24 May.
So "Ukraine" interest is still 3-4x times its pre-war baseline (2021 trend line), but it is 1/25th to 1/33rd of peak interest near the start of the invasion.
And, due to a specific awards show incident, interest in Ukraine was eclipsed by searches on "Will Smith" centered around 27-28 Mar 2022. That interest did not peak as high as the peak interest in "Ukraine," and was even more ephemeral. By 02 Apr 2022 interest in Will Smith fell below that of interest in "Ukraine." And now interest in that individual has basically resumed to its pre-stochastic event baseline.
Interest in topics is not only temporal, it is also geographical and social. Interest may rise (or fall) in certain countries or states or even cities. Or its locus may be based on virtual communities other social reference groups. "Marvel" or "Star Wars" fans might be hyped by certain news, and interest can spike amongst them, while only having peripheral interest in other social reference groups or amongst the general public and popular culture.
Cult-followings of music, art, or fashion, mass-produced luxery fashion labels (where market growth is the kiss of death, though a new brand can of course, be easily launched, and is), a secret vacation spot (where travel costs are largely the same as to any other point on Earth), etc., etc.
Although this same paper says that "more species of beetles (>350,000) have been described than any other order of animal, insect or otherwise"
[1] https://www.biorxiv.org/content/10.1101/274431v1.full.pdf
Just speculation...
Regarding the number of beetle species, JBS Haldane (probably) said that "God is incredibly fond of beetles" (or words to that effect).
https://en.wiktionary.org/wiki/mother does not mention that fact, though.
Nevertheless, what a nice new ambiguous word to spice up the context!
Furthermore, the user Ruigeroeland has contributed to three of the insect pages.
So I guess these users have the distinction of having contributed to multiple of the least-seen articles on Wikipedia. (They probably also contributed to many widely-seen articles!)
> For example, the 12-word stub Pottallinda (5 views last year) was created on 18 January 2011 by User:Ser Amantio di Nicolao, who happens to be the most active editor in all of Wikipedia (as measured by number of edits). Within 60 seconds of creating this page, the same editor also created Polmalagama, Polommana, Polpitiya, Polwatta, and dozens of other substantially identical articles.
"the smallest uninteresting number is itself interesting because it is the smallest uninteresting number, thus producing a contradiction."
https://en.wikipedia.org/wiki/Interesting_number_paradox
Incidentally, the last line of that article is:
> The mathematician and philosopher Alex Bellos suggested in 2014 that a candidate for the lowest uninteresting number would be 247 because it was, at the time, "the lowest number not to have its own page on Wikipedia".
(I didn't see this linked in the post. Apologies if I missed it)
[0] example: https://en.wikipedia.org/w/index.php?title=Weimer_Township&a...
To insert a node, generate a random bit string: descend the tree and when the counts are equal then take the branch corresponding to that position in the bit string. When the counts are unequal, take the branch with the smallest count.
To remove a node just remove it from the tree and update the counts up along its path to the root. This assumes that articles are added as often as they are removed.
To sample, just generate another random bit string and traverse the tree according to it.
But I agree with the sibling comment that the current non-uniform algorithm is fine for a cute “Random article” button.
It could even be maintained entirely outside the Wikimedia servers, relying on database dumps.
However there was a brief time period where randomPage used elasticsearch to get the random article instead.
Maybe a hashing function could work.
Sure there will of course be some unlucky articles, but does that actually matter?
Quote: "The least viewed article in the sample, Erygia sigillata, has a page_random value of 0.500764585777. The article Katherine Hanley is right on its tail with a value of 0.500764582314, which is just 0.000000003 less, or 3e-9 in scientific notation. This is 98% smaller than the average random gap. In other words, Erygia sigillata is an extremely unlucky article as far as the “Random article” button is concerned! It’s 50 times less likely to be landed on than an average article."
It’s only once your supply of numbers runs low that differences will start to equalize, reaching 1 when you have exhausted your supply.
Besides, the author isn’t really making an argument. They are giving you actual data showing the differences to the next lowest number. It’s hard to argue with that.
> the gaps will differ because there is no mechanism that would make some number with close neighbors less or more likely to be drawn than any other remaining number.
When a new article is inserted, there is a higher probability it will be inserted in a large gap than a small gap, so it should balance out.
I suppose you're right,i am responding to an implied criticism to the randomness method that the author didn't make. They just offered it as explanation.
Now think about a random-gap list with five million entries. With that many entries, will the gaps balance out? In the best case, five of them end up in the range from 0 to 1 millionth, five of them end up in the range from 1 millionth to 2 millionths, etc. But we've already seen what it looks like to have five uniform random numbers in a range (whether it's big or small doesn't matter); the gaps tend to be really varied. So we're going to get this sort of imbalance between gap sizes (viewed as a ratio) no matter how many entries we insert.
One other thought experiment: the article mentioned a page with a gap size of one billionth. How many pages will it take for that to balance out so that page doesn't have an unusually small gap any more? How many pages does Wikipedia have?
(This is similar to the reason that infinite space packed with marbles has the same packing density as infinite space packed with bowling balls.)
Differences around zero is actually preferred: https://mathworld.wolfram.com/UniformDifferenceDistribution....
That’s a good point and I’m not entirely sure why it (appears to) not work that way. Maybe it’s because that interval has a higher likelihood, but there is no preference for numbers towards the middle, that would dissect it into (roughly) equal parts?
More interesting question would be what is the standard deviation (of gap size), not what is the worst outlier
https://wap.business-standard.com/article/pti-stories/wikipe...
Besides finding hidden gems or undervalued articles, I would really like to see traction tracked. Now that rarely viewed articles are somewhat in the spotlight and get promoted, a before and after comparison would be interesting.
Fun anyway.
https://pageviews.wmcloud.org/?project=en.wikipedia.org&plat...
I've come across plenty of articles which were low in views and only have a handful of lines of content. I would love to expand these articles, but the "culture" of Wikipedia actively works against this.
Can you expand on this? I've only had positive experiences editing for Wikipedia as a newbie.