In search of the least viewed article on Wikipedia (2022)
colinmorris.github.io
colinmorris.github.io
This is not why this happens on Wikipedia; sourcing is a factor, but notability derived from the nature and quality of the cited sources is the most commonly used criteria to delete articles.
For example, Wikipedia once had a notability guideline that a person who completed in a sport at the international level was probably notable, so a weakly sourced article on, say, a footballer with a single international cap could be notable enough to avoid deletion.
But then editors removed this guideline and fell back onto the general notability guideline (GNG), which is much stricter for people: substantial biographical coverage from reliable mainstream sources at the national level. Match reports don't count (too routine), no local interviews (too primary), no enthusiast publications (not reliable enough), etc.
Of course, this meant that hundreds of stubs and dozens of full articles about women's international footballers would never meet GNG due to a relative lack of mainstream media coverage, even if they won a World Cup. So all it took was one editor to decide to do almost nothing but flag hundreds such articles since the opening of the Women's World Cup this summer, almost all of which were deleted.
So it's not this reason that keeps moths and places:
> On the other hand, no-one has yet come up with a way to monetize a topic like Pseudoneuroterus mazandarani or use it to push a contentious point of view.
But rather, very different standards of notability are why there are more low-traffic articles on items and places that can't sue Wikipedia editors for libel.
And does anyone actually believe this is applied consistently?
I'm willing to bet that there are countless Wikipedia articles about people in technical, academic, local political, etc., etc. roles/positions that wouldn't meet that criterion.
Classic lawful evil behavior.
I would go there a few times a day, and noticed that almost every single article it gave me was some random amateur athlete, with only a results page as a source.
I was at an early Wikipedia meetup and a woman editor told me that women porn stars were more likely to get (and keep) a profile page than a woman sci-fi author. It's pretty clear that Wikipedia's quality has been diminished by this level of misogyny.
Once upon a time, subject-specific notability guidelines carved out exceptions by arbitrarily defining typical notability for a given subject at a lower standard than general notability. Those exceptions are being rolled back toward the higher bar of general notability.
When a women's sport gets 1/10th the mainstream coverage of the men's equivalent, there's no policy justification to have articles under general notability. Even if Wikipedia was capable of banning every editor whose focus is in deleting articles about women (and it's sincerely not capable of banning any of them), the core policies themselves would justify not creating articles — especially for almost all who played in the decades before the recent boom in coverage and investment — because the coverage wasn't and isn't there.
It’s that time pressure gives a great deal of power to anyone with an agenda of any kind.
So for example, per this policy, it's impossible to rewrite https://en.wikipedia.org/wiki/Dana_Rivers to point out the angle of male violence against women, despite there being sources that exist discussing this, because it's forbidden by policy to refer to this male as a male or use sources that describe him as a man. Instead, the pretence that this is the crime of a woman has to be maintained.
https://en.wikipedia.org/wiki/Wikipedia:WikiProject_U.S._Roa...
Because there's very little effort required to flag articles for deletion, and the burden to keep is often on the contributors doing research on a 7-day timeframe if there are even just one or two editors supporting deletion, most times the contributors eventually give up running the research treadmill and the content gets deleted anyway.
And WP:GEOROAD is a subject-specific notability guideline, like WP:NSPORT is. Those aren't immutable; NSPORT was fundamentally changed in 2022,[1] and those changes now justify all these deletions of international footballer articles created before. If the Roads editors have all left, there'll be little or no opposition to changing WP:GEOROAD,[2] and deletions of that content from WP could get done even faster.
1: https://en.wikipedia.org/wiki/Wikipedia:Village_pump_(policy...
2: https://en.wikipedia.org/wiki/Wikipedia_talk:Notability_(geo... - and I see some of the editors who participated in NSPORT-related deletions this summer also in there advocating to change GEOROAD and make it more explicitly defer to GNG, and therefore easier to delete articles.
Pretty much sums up the problem with every community-run moderated site. It directly leads to the downfall because contributors get discouraged and leave, and eventually all that's left is people with the same view as the heavy-handed moderators. Having hours of work wiped out from editors/moderators who've spent less than a minute on it sucks.
Stackoverflow is another notable example of this.
It's often a bad thing in practice, much like how splitting off entirely separate social networks for specific subjects is often a bad idea. If there aren't enough people or activity, interest wanes. If the project relies on one or two heavily active editors, it collapses when they're unavailable. If you never hit critical mass for SEO or can't/don't know how to promote your content, all the work is done in a vacuum and everyone sees the Wikipedia stubs over your stuff anyway.
Fandom (as in the company), as awful as it is, largely exists by virtue of participation in the subjects where there's enough interest to facilitate it; the mass wiki-farm infrastructure gives them enough of a SEO boost to dominate even some of the wikis that fork off of Fandom to escape its policies.
> you can always import back the articles into Wikipedia. So maybe it can be a useful way for communities to "incubate" articles?
Developing an article on a separate wiki has additional technical and administrative overhead.
CC BY-SA requires attribution. Wikipedia attributes contributions by article history. The only way to import article history is via Special:Import. You have to be an admin or have import rights to use Special:Import on Wikipedia.
If you can clear all those hurdles, then in theory you can import changes with the required attribution from another wiki. But if the imported changes conflict with changes made in parallel on Wikipedia, resolving them can be very contentious, especially if any of the editors involved disagree on the resolution. Or the import might fail, because Special:Import isn't particularly robust.
And as the forked wiki admin, do you decide to set up scheduled imports _from_ Wikipedia to keep up with upstream? These forks often split because Wikipedia is deficient in some way beyond just editing the text. AARoads' fork's use of improved maps, for instance, would probably break the article on Special:Import because it looks like they use different wiki templates and modules.
It's odd but I come back to Cantrill's "Fork Yeah" talk here.
So the basic problem is one of resources: forked wikis still have hosting costs, so they need either a patron community (often just one person) or a for-profit company (with lots of crappy ads), while Wikipedia by all accounts is rolling in the dough with very little accountability (i.e. when professors and teachers remind students "don't believe 100% of what you read on Wikipedia," the-Wiki-community is blamed moreso than Wikipedia-the-nonprofit). If you fork Wikipedia then you lose out on the resources.
Without that, you do have a "forkophobic" culture which is why you get this "governance orgy" -- the notability guidelines and so forth. But the difference is, Cantrill's software examples expect the software to have some sort of editorial control, so if the Linux or Apache foundations take on some project it's because it's used by thousands or millions of people and they don't really take kindly to "oh yeah upstream was vandalized by someone who came in and just made every request to the Apache server return 'HTTP 499 BUTTS BUTTS BUTTS'."
In Wikipedia you get this strange direct democracy by "whoever happened to show up." Deletion votes are often done with like, 20 votes or less of just random passersby. Worse, those random passersby are usually the people who visited the article in the first place and saw that it was up for deletion, so they'll say things that are nonsensical like "oh, he's a very notable figure in the XYZ community, Googling him turns up 40,000 results so clearly he is notable."
20! That'd be a luxury.
Here are some of those footballer AfDs from this month that passed with 1 vote on top of the nomination. Without passing any judgment on whether the subjects are notable, see if you notice any participation trends:
https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...
https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...
https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...
https://en.wikipedia.org/wiki/Wikipedia:Articles_for_deletio...
You don't even need that. IIRC, the "PROD" process can get an article deleted with no votes at all. All you have to do is tag the article and if no one removes it for a week, it will be deleted with no further discussion.
The caveats are you can only PROD something once, and I believe no discussion is required to delete a PROD'ed article (if anyone remembers it to want to bring it back).
Migrating all that cruft to a separate pokemon wiki was an improvement for everyone, no matter how "notable" you think Beedrill might be, it doesn't need it's own separate wikipedia article. With a separate wiki, Pokemon fans can go into as much depth and lore as they like.
Personally I think the criteria for fictional things should be even stronger. There ought to be an article about the work of fiction itself, but not articles about fictional characters or events unless they're notable outside of the work of fiction.
This keeps wikipedia about facts, not fictional canon.
e.g. Pikachu derves a separate article, Charmander does not.
To elaborate further, the Charizard article has a "Physical characteristics" section.
It's a anime / computer game. It's not a physical being, so any "physical characteristics" is not factual information, it's fictional information. Wikipedia does a poor job at separation of fact and fiction in articles about fictional beings.
At a guess, it's probably something like practical restraints around "not enough people monitoring for quality", or the fact that the hard drive space to save all of this information is not free and unlimited, or that simply, it might be better served by niche communities who will be devoted to caring far more about such specific topics?
I just find it frustrating that there isn't a kind of...ultra, super mega colossal set of all human knowledge of everything stored under a single digital roof. I suppose that ideal itself probably isn't practical for the reasons mentioned...it just seems so neat in concept. Just one place for everything.
- some editor probably
In my handful of encounters trying to contribute to wikipedia it's always been such a frustrating experience.
"not enough people monitoring for quality" is one way to put it, but I've often found it to be one very zealous person monitoring for their idea of quality. It ends up quite frustrating, especially if you're a domain expert.
I've corrected articles where things I've written have been cited and had the changes reverted. It was enough to just give up.
I've heard about this happening enough that I stopped treating WP with any credibility whatsoever even for what should be cut & dry fact (aka: non-controversial/political topics). I've heard of people who were being quoted updating the context to more accurately reflect what they were saying and having the changes reverted. As if the person who said the thing being quoted doesn't know what they meant. Often because it didn't meet some guideline or another but more often than not because one overzealous editor has decided that the page being edited is "their page".
I learn the truth more from perusing the edit history or talk pages than from ever reading the page itself. Also despite claims of neutrality it's amazing how often pro-communist articles are heavily maintained almost exclusively by diehard self-proclaimed Marxists making politically biased edits.
Exhibit A: https://en.wikipedia.org/wiki/Talk:Holodomor
What they meant at the time they said it and what they want it to mean later upon reflection, certainly could be two completely different things and should be scrutinized.
Granted, this is also a problem with the obscure moth species articles. I think the moth species articles survive because no one really cares enough to start the crusade against them. When we used to have every Pokemon, it was a common line in deletion discussions to say "well if every Pokemon has an article, why can't <my obscure topic> have one too?" -- I think some people eventually got fed up and decided it was worth putting in the work of figuring out what the notability standards should be. What I've learned from a long time of editing on Wikipedia is that often, things are the way they are not because it's the best way, but because the project has a lot of inertia -- it's a lot of work to make a big change happen.
The concept you long for sounds a little bit like Wikidata. It's much less in-depth than Wikipedia, and just describes its subjects as structured data instead of with prose, but the notability bar on Wikidata is much much lower. Every Pokemon, every scientific article, every book, every village, every athlete, etc is generally in scope.
This has always been my sticking point with deletionist thinking. Why doesn't it need it's own separate wikipedia article? What's the harm in it? Are we worried that people will start treating Charizard as a real creature?
Where I see the value of notability criteria it is mostly in preventing vanity articles. Beedrill, presumably, is of general enough interest that people are willing to contribute and reference the information. Why isn't that enough?
The reason that there are wikis dedicated to "everything Fallout" is because of these deletionist sentiments. Most of these things started on wikipedia and had to migrate off because of the constant barrage of deletion fights.
I'll bite, though -- why the passive voice? What is the argument? The first blush here is that these topics are non-serious and make Wikipedia seem less serious. That's clearly a strawman though -- what's the deeper argument? I mean, for Brittanica, you only have so much print space you can use, and an article on Beedrill is a waste of paper. But Wikipedia is not printed; and while space is scarce in theory we shouldn't be rationing until the need it apparent.
So at the very least that's not actively motivating the continued state of the notability guidelines, though I don't know if it's possible the original motivation was as you suggest.
"Jimmy Wales also hasn't exerted control over Wikipedia policy in a long time, and as of this year no longer even nominally has the right to overrule ArbCom."
Except these two facts niggle.
He's on the board, and will always be the founder. He is SV aristocracy in a town where the money you earn is less important than the people you know and the things you have built. He will always have influence and it's naive to believe he wouldn't. So, it is not entirely accurate to say he has no control, just less control than he did. Is it too little control? Perhaps, perhaps not.
If I wanted to construct a set of policies that drove traffic to another site, I would make them quite like Wikipedia has now e.g. Wikipedia is not a gamers manual. Then I would target for deletion much of the nerd content, quite like now. This is quite a coincidence.
I'm not attached to this theory. Also, I don't like spreading negativity. But to me it seems at least somewhat plausible, and that possibility disturbs me a little.
https://en.wikipedia.org/wiki/The_Mexican_Runner
I had to fight a bit with other editors to convince them that this guy was notable. He had coverage in Polygon and Kotaku, so that seemed enough to me, but it was a difficult task at first.
Is some dude who pushes buttons very quickly notable?
If someone gets significant coverage from multiple reliable secondary sources for taking a shit, they can have a Wikipedia article. They don't have to win an award for it (though it'd help) or contribute to society (though it'd help).
The action itself might be broadly, by common sense, not notable. Only the notoriety of the person _in published sources_ is relevant.
A bit weird that the article includes details of his mother's medical condition though!
https://imgur.com/gallery/ILp6TtA
I get why they have to be collapsed by default, but I use a userjs to move them to the top of the page, uncollapsed, so I can use them to navigate (I set their CSS zoom to 0.3 so they don't take up too much space)
I cannot see it, not at all.
This makes me wonder if there isn’t some intentional weighting going on, e.g. certain favored pages get space in front of them cleared out so they’re more likely to pop up.
I agree that rerolling is probably a fine strategy, but it's not obvious that everyone would be willing to just trust probability like that. In fact, if you set up your own wiki with 99 bad pages and 1 good page, hitting the random article button won't work most of the time.
You could of course pick a random number from 1 to M, where M is the number of valid articles. But that runs into the same problem of how do you map the 1-to-M number back to an actual article id?
This method may be biased, but it's constant performance.
(I still prefer just keeping all valid articles in memory and choosing a random one of those. There are only 7M as of October, so including the deleted pages, there's probably ~70M, which is on the same order as HN's total item count. Heck, even creating one file per valid article on disk and doing a random search with "find . -type f | shuf -n 1" probably wouldn't be terrible if you cache the result every few seconds, though that'd also be biased in its own way.)
I'd have 1) a remote cache for deleted articles 2) the application servers adding to the remote cache whenever they roll a deleted article. The remote cache can be data-center local and backed by durable storage, with a long TTL. It's possible to not roll un-deleted articles until the cache TTL, but that's ok. Data-center locality ensures some level of fault tolerance and low latency. Durable storage means restarts don't have as much of a cold-cache problem. The ID range maximum water-mark can be updated asynchronously. It's ok if new articles don't show up for a short while.
I suppose current implementation makes sense since the weights would be roughly equal as the number of articles increases, and any individual user is unlikely to care about fairness as long as they get some random article.
If you do something naive like "ORDER BY RAND() LIMIT 1" performance is worse than abysmal. Similarly, doing a "LIMIT 0 OFFSET RAND() * row_count" is comparably abysmal. While if you do something performant like "WHERE id >= RAND() * max_id ORDER BY id LIMIT 1", you encounter the same problem where the id gaps in deleted articles make certain articles more likely to be chosen.
There are only two "correct" solutions. The first is to pick id's at random and then retry when an id isn't a valid article, but the worry here is with how sparse the id's might be, and the fact it might have to occasionally retry 20 times -- unpredictable performance characteristics like that aren't ideal. While the second is to maintain a separate column/table that includes all valid articles in a sequential integer sequence, and then pick a random integer guaranteed to be valid. But the problem with performance is now that every time you delete an article, you have to recalculate and rewrite, on average, half the entire column of integers.
So for something that's really just a toy feature, the way Wikipedia implements it is kind of "good enough". And downsides could be mitigated by recalculating each article's random number every so often, e.g. as a daily or weekly or slowly-but-constantly-crawling batch job, or every time a page gets edited.
You could also dramatically improve it (but without making it quite perfect) by using the process as it exists, but rather than picking the article with the closest random number, select the ~50 immediately lower and higher than it, and then pick randomly from those ~100. It's a complete and total hack, but in practice it'll still be extremely performant.
My comment covered your suggestion already -- that's why I wrote "you encounter the same problem where the id gaps in deleted articles make certain articles more likely to be chosen".
And you've still got to decide how you're going to pick random ID's from an array of tens of millions of elements that are constantly having elements deleted from the middle. Once you've figured out how to do that efficiently, you might as well skip all the trouble and just use that algorithm on the database itself.
How so? you have array of only valid IDs [1,2,3]
Add a column to the article database for "random_index". This gives each article a unique integer index within a consecutive range of 1 to N. To pick a random article, just pick a random number in that range using your favorite random number generator and then look up the row with that random_index value.
The tricky part, obviously, is maintaining the invariant that random index values are (1) randomly assigned and (2) contiguous. To maintain those, I believe it's enough to:
1. Whenever a new article is created, give it random_index N+1. Then choose a random number in the range (1, N+1). Swap the random_index of the new article and the article whose random_index is the number you just rolled. This is the Fisher-Yates step that gives you property (1).
2. Whenever an article is deleted or otherwise removed from the set of articles that can be randomly selected, look up the article with random_index N. Set that article's random_index to the random_index of the article about to be deleted. This gives you property (2).
Yes, this should definitely considered to be the correct, definitive, performant solution then. Thank you!
I think you'll need a full table lock in the case of a race condition trying to do a remove and add at the same time, but that's probably not usually going to be an issue?
> I think you'll need a full table lock in case you're ever trying to do a remove and add at the same time
Yeah, I think you need some sort of locking to ensure N can't be concurrently modified but I'm not a DB person, so I don't know how to go about that.
Also I liked your book.
I believe that for the operation of randomly sampling a single row, only the consecutiveness invariant is needed. The random permutation invariant is unnecessary, because the choice from the range (1, N+1) is random anyway.
However, I can imagine the random permutation invariant being useful for other operations (although I can't immediately think of one).
EDIT: the random permutation invariant could maybe be useful for a list that can be displayed to the user. I can imagine some cases where it is nice to have an order that doesn't change a lot while you don't want it to depend on any property in particular (like the order in which the rows were added).
Is this really a big deal? Surely for a SQL database, looking up the existence of a few IDs, just about the simplest possible set operation, is an operation measured in a fraction of a millisecond. Wikipedia is not that sparse, and a lookup of a few IDs in a call would easily bound it to a worst case of <10 calls and still just a millisecond or two.
Edit: what would be really cool is if I could leave a note on a page saying "hey, I visited this page and I'm looking to connect with other people interested in the Peruvian Foovius Barivius Moth"... ping me
Stop the destruction! End this dumb rule and make Wikipedia truly open to all.
Same game in a one language can be extremely notable and well known to everyone, while in other it'd get deleted without a trace.
Sometimes it feels like "racism" based on countries (countrism?)
I'd call it national chauvinism.
Generally not. There is a separate process called "oversight" [1] which behaves the way you're describing, but it can only be invoked by a very small number of privileged users, and is only used in extreme cases.
There's a place called Ksar Ouled Soltane [0] and I couldn't quite figure out if it was a film site or not. The locals assured me it was, but there were conflicting reports online. I saw that it was listed on Wikipedia, but when I went through the history, I saw that someone tried to remove it [1]. Unfortunately their removal was reverted and their source [2] wasn't considered good enough.
That said, I read their article and I'm much more inclined to believe them than the wikipedia editors.
It was the first time I ever encountered something like this in the wild.
[0] https://en.wikipedia.org/wiki/Ksar_Ouled_Soltane
[1] https://en.wikipedia.org/wiki/User_talk:Dbavandy
[2] https://galaxytours.com/research/ksar-ouled-soltane-debunkin...
It is telling that he never responded to Davin Dbavandy question about why it is not a reliable source.
One thing that always stood out to me was how till 2014 you could find many references on wikipedia to pre-2014 attempts from Crimea to secede from Ukraine.
"Crimea" (quoted, as it may refer to different groups/leaderships in Crimea) tried to split from Ukraine twice already during Soviet times.
But this whole thing got edited away in years across dozens of articles.
I think there's a significant nuance between the two narratives.
Unfortunately, it seems so ossified that that work can only be done within Wikipedia (notable article, standards, etc).
A far more reasonable approach would have been to have CommonPedia... and then regularly pull qualifying information into Wikipedia.
This is called "The Web". You’re free to write whatever you want, it’s just not on Wikipedia.
> it's not like they're limited by the size of a bound encyclopedia.
No, but articles still require maintainance, so it’s false to say that just because this is "in the cloud" you can have an infinite number of articles.
Her legal name had appeared in some public court filings, and some user had added it to her Wikipedia page with the reference.
Someone else (purportedly one of her co-stars) had requested removal.
A Wikipedia editor then got involved, and refused to budge from the position the original actress was notable enough to merit inclusion of her legal name.
Cue months of back-and-forth deletions and reversions.
Finally, I think a year later, someone quietly edited the legal name out and it stuck.
Sometimes people with a little power are the greatest assholes. Probably because they sought the little power.
I don't know about this case, but it seems entirely reasonable that to argue for including the name, or not.
Plenty of public personas have aliases they wish to be known by, but their Wikipedia article prominently notes their real name.
So by definition it's not exposing any information you can't find in other public sources.
Whether it's notable and something that should be included is another matter.
In any case, I'm saying that if you look at the articles for Madonna, Bono, Jenna Jameson, Marilyn Manson etc. you'll find their real name displayed prominently.
We can't form our own opinion in this case, as the GP is deliberately keeping the specific article in question hidden.
But in general this seems like a thing two editors could reasonably disagree about.
For the section in question, see: https://en.wikipedia.org/wiki/Wikipedia:BLPPRIMARY
Miloš Moth (Jan 3, 2023–May 17, 2023) was a Peruvian moth,
Are you seriously suggesting that before the internet no one published or wrote about highly niche or uninteresting subjects? That you couldn't find mundane facts from history or nature written down somewhere?
> Edit: what would be really cool is if I could leave a note on a page saying "hey, I visited this page and I'm looking to connect with other people interested in the Peruvian Foovius Barivius Moth"... ping me
No, that wouldn't be cool at all. What we'd then have is a social network of people competing for some kind of points or popularity, or to prove who was the most expert or most interested in the subject. And we would see the quality of the articles degrade, and suddenly no one wants to work on the unpopular stuff. And people like you would be left wondering, "gee, why is it no one wants to write unpopular articles for this encyclopedia?"
The standards came much later.
It also already exists, but in a better form. If you find something is interesting, write about it. Then to connect with others you can find others who have written or researched it and then write them a letter/email/message. They will likely not be an island and you can get connected to a group of people who are genuinely interested in a subject without the degenerative downsides of a social media platform.
Yes and no. I'm suggesting that before Wikipedia, no encyclopedia could have published an article about Jack Wade, Alaska, an incorporated community of probably < 100 people (I picked this place somewhat randomly off of google maps). If you put even one sentence about every town similar to that in an encyclopedia, you'd have an encyclopedia set that no one would be able to afford.
> What we'd then have is a social network of people competing for some kind of points or popularity, or to prove who was the most expert or most interested in the subject.
This really depends. For example, maybe I'm an American with a grandparent from Volosianka, Lviv Oblast, Ukraine (again, I picked this semi-randomly). Population is about 1500 today. There's probably a handful of people in the world who speak English and who have a particular interest in that town. Maybe my grandpa is from there and i'm looking to learn more about his life and i'd love to chat with someone else who knows more about the place. I think that there are a lot of articles on the long tail of wikipedia where there's probably only a few people in the world who care about the topic and if I'm one of them, I'd be interested in meeting with others. Obviously, there's plenty of topics where this isn't the case. I don't know how you'd implement this feature, because I don't think you want to turn wikipedia into a meeting room for the Israeli-Palestinian conflict [1]. But for these very very niche things, it could serve as an interesting meeting place.
[1] In other words, I am stating a problem here... not a solution
The moth would be described in the original scientific publication (from when it was first discovered) and then in books listing the Moths of Peru etc.
There's a feature to write a note and leave it on the page for others (with the extension installed) to see and maybe have luck someone contacting you.
> There's a feature to write a note and leave it on the page for others
Your site says that "nothing is stored." Does writing a note require everyone to be on the site at the same time? At first I thought it was sort of like the old "Dissenter" extension that added comments sections to every page, but the description seems to indicate a more live chat style.
Start with ranking all articles by their interestingness. Then order the list and look at the lowest values. There must be a least interesting article. But the very fact that this is the least interesting one in the entirety of Wikipedia makes it interesting.
In the same way, I assume the least viewed articles in this blog post will have lost their status by now.
No matter how many views they get now, they would still be the least viewed article in the year 2021.
You didn't directly say it, but as the status given them in the blog post was "least read article of 2021" the statement is wrong. It's more because you worded it ambiguously, but that's why you got a pedantic reply.
I don't get why you're so reluctant to admit you could have worded something better. "I assume the least read articles of 2021 mentioned in the blog post are no longer the least read articles this year" would have said the same thing without the ambiguity. It's not that big of a deal.
Unless articles become interesting if they're in the wikipedia category "Uninteresting".
* https://www.youtube.com/playlist?list=PLt4q5oaptyI9U2zddss8d...
It's a great way to delete uncomfortable historic facts. Now it's just a funny little story about a rootkit with no specific political intrigue at all.
That's not how I'd describe the page today
[1] On Google maps view, i can't maybe 8 houses there? https://www.google.com/maps/place/Jack+Wade,+AK+99732/@64.15.... But it's on wikipedia here: https://en.wikipedia.org/wiki/Jack_Wade,_Alaska
It strikes me strange that those are considered "notable enough" mainly because they've been in Wikipedia long enough, but new articles have to fight.
Petty bureaucrats can only show their power by saying no, if they say yes it's as if they weren't there.
Spoiler for the interested, and it could be an artifact of the dataset the author used, but one of the commonalities between the least viewed articles is that their subject matter falls into a category that isn't usually eligible for deletion under Wikipedia content guidelines.
The edits show there was disagreement earlier this year about the wingspan of the Scrobipalpula crustaria.
Is it 11-13mm or 10-13mm?
People feel strongly about these things.
Personally, I think the reflexivity is part of the fun if you do a followup analysis. For example, I recently scraped WP to find 'the first unused acronym on Wikipedia': https://gwern.net/tla - it turns out to be 'CQK', and I'm looking forward to checking back in 10 years or so to see if anyone wound up using 'CQK' for a company or something, precisely because I wrote it up. We'll see!
TBF, he's analyzing a 2021 dataset, and that will of course not be affected.
6.0e6 been a brute force number for decades.
A linear search might have taken less time than reading the article. Almost certainly less time than writing it.
It would not have been as entertaining, of course or as clever.
Good engineering is like that nearly all the time.
In search of the least viewed article on Wikipedia - https://news.ycombinator.com/item?id=31524943 - May 2022 (68 comments)
And further, there's good arguments that they shouldn't be equally likely. Instead, rank articles by size, or popularity, so people are more likely to see articles that are reasonably interesting and full of content.
https://huggingface.co/NeuML/txtai-wikipedia
Example query.
# Find wikipedia pages for a topic in the bottom 5th percentile of page views
SELECT id, text, score, percentile FROM txtai WHERE similar('topic') AND percentile <= 0.05
As someone who was already impressed by Wikipedia, wow! Talk about toiling in obscurity
The power of slashdot/HN
What about smallest number of views V, and word count C.
What is smallest V/C?
LOL
https://en.wikipedia.org/wiki/Draft:Playwright_(software)
It is highly demotivating to try to write a quality article for Wikipedia because someone can just reject your days of work in seconds and then leave it in draft forever.
I hate when things get mature and enshittify.
I don't create articles these days; it's too much like hard work.
You used to be able to create an article by making a redlink wikilink somewhere, and then clicking it - that would take you to an editor. The new article would appear in mainspace. Do you have to "volunteer" to appear in mainspace? Are all new articles deemed to be "draft", and subject to review nowadays? Or are some users allowed to create new articles in mainspace, and others not? If so, what are the criteria?
The previous system of doing post-publish review of new articles had an even worse backlog because it let a huge volume of shit into the front door.
The root cause of this is Wikipedia’s huge influence in Google SEO and knowledge graph. As the open web is dying, it’s one of the few reliable ways to dump information straight to the top of Google SERPs. When I write new Wikipedia articles it is often indexed to the first page of results in minutes.
[Edit] I wonder if Curb Safe Charmer automated his review using something like this?
I don't know how Wikipedia works -- is there not a way to submit a smaller/stub page or something, and then once it passes review for notability, then you fill out the rest?
Does Wikipedia expect people to put in days of work before an article gets approved? I hope that's not the case.
I've written dozens of articles starting from a few sentences and slowly expanding it over the course of a few days. I simply make sure the topic is notable enough and that every sentence is cited.
Generally the hardest and most annoying part about a case like this is wikipedia's desire for reliable sources which generally means newspapers or academic sources. You can have something super well known, around forever, and often mentioned but it can be hard to find sources writing about it directly. For example, I wrote the Wikipedia article for high touch (https://en.m.wikipedia.org/wiki/High-touch) and it was quite annoying to try to find sources. Similarly I wrote one for "smell training" (https://en.m.wikipedia.org/wiki/Smell_training) and the requirements are even higher for anything related to medicine.
Anyway, I'm not gonna say I advise you to do this. But if you're really confident that the article meets the notability guidelines you can also create the article in mainspace (instead of draft). If someone tries to delete it then the guidelines are a little different (vs reviewing drafts) because they have to be sure that there aren't sources to make it notable (not just that you didn't use them). Often times people will actually find sources and add them to the article rather than just complain you don't have good enough ones.
PS: I'm sure someone will come and explain why I'm wrong here but this has generally been my experience.
Which is one of the main ironies associated with Wikipedia notability. Something or someone written about in some small-town newspaper or an obscure journal that five people have read can be considered more notable than something with a fairly big online footprint but nothing canonical.
For that matter, there's a ton of pre-web information that just doesn't really much exist outside of primary sources.
The current version looks fine. Unfortunately reviewing drafts is a very thankless job (~90% of drafts are worthless) so there are never enough reviewers (speaking as an Wikipedia admin who used to review a lot of drafts). The backlog is definitely not something people are happy about, but it isn't easy to solve.
Also, having your drafts reviewed is actually not required. Once you have made 10 edits you can move the draft to the main article space yourself (or directly create articles). The reason why brand new users can't directly create articles is that whey they used to be able to, a ~third of all new articles ended up being deleted immediately because they were spam/gibberish/vandalism, which ends up both being a lot of work for reviewers, and very discouraging to those new users.
[1] https://en.wikipedia.org/w/index.php?title=Draft:Playwright_...
What's interesting is this is more a phenomenon on the English wikipedia, so if you're disappointed that a search for information on a FOSS project is turning up nothing on en.wikipedia.org, try just switching en for fr or de (both relatively large) and then using google translate or your browser's built-in translate [firefox] to get the info you need.
Sad. And yeah, I feel deletionists are out of control on english wikipedia.
https://en.wikipedia.org/wiki/Playwright_(software)
Thanks!
It only gets worse. Try getting into a content dispute sometime.