Wikipedia's in Trouble (2019)
blog.spencermounta.in
blog.spencermounta.in
Or the category thing - yes, it has circular references which breaks the nice theory behind it, but what practical issues to the readers or editors does this present?
The syntax is messy, yes and it's a problem. But converting it to a markdown/HTML is downright absurd. I think the current approach of improving it step by step is the most reasonable. Wikipedia is a project which can afford to sustain its own custom syntax (because of history, compatibility concerns and wikipedia's special needs).
I think the biggest issue isn't technical in nature - it's the actual task of building the encyclopedia - to decide together on its scope, decide what is notable and what isn't. To fight against brigading from interest groups, fight against attempts to game their rules. All of these are intrinsically difficult tasks and of course wikipedia has made many mistakes. But even then it still fares pretty well in my book.
Difficulties are what they are. Part of why things are difficult is because perfect, mechanical solutions aren't feasible. If they were, the problem would be solveable, commoditized and something else would be difficult instead.
This article does a good job presenting difficulties, and they are all important areas where doing a good job is important.
Eg "Or the category thing - yes, it has circular references which breaks the nice theory behind it, but what practical issues to the readers or editors does this present?"
There probably isn't a good, legible solution that will "fix" this. It's an epistemological problem about knowledge itself. Wikipedia's approach is an approach. It isn't perfect because there are no perfect solutions to knowledge problems. It is however, a workable approach given the problem space.
None of these need 100% solutions. They just need workable solutions.
I feel like Wikipedia (at least English Wikipedia) has long past this, and pretty much keep everything that has even the slightest of notability.
And I have no issue with it. I see this as one of the (big) advantage of Wikipedia over traditional encyclopedia.
(Don't get me wrong, I knew https://en.wikipedia.org/wiki/Wikipedia:Notability still exists.)
Like you, I have no issue with most articles at some level of verifiability sticking around. Most things are notable to some degree if you get local enough (whether geo, field of interest, etc.).
Wikipedia keeps deleting plenty of pages. There's a lot of self promo, a lot of pages about not sufficiently notable topics etc. You could think that e.g. mayor of your local 5K town will have "slightest notability". Well if that's their only claim to fame, then their nice Wikipedia page is going to be deleted.
> And I have no issue with it. I see this as one of the (big) advantage of Wikipedia over traditional encyclopedia.
Wikipedia is not limited by the physical constraints, but other limits are still present - e.g. the ability to check the validity of claims of the original article and subsequent edits is limited by the available manpower.
When I was replying I was thinking of some common "complaints" about Wikipedia like "why the hell we need an article for every single team's 1928 season in 3rd tier Polish football" etc.
(Unless you specifically go looking for it, in discussions about what was deleted or whatever.)
Adding a new page is harder than ever, because of WP:N.
I have peers in my field, executives at major companies, journalists who have bylines pretty much every days, book authors, etc. who are quite well known in the context of my field or at least a subset of it and are easy to verify--but I'm guessing that many articles about those subjects would be somewhat arbitrarily deleted for lack of notability.
The really serious problems are organizational in nature: the overgrowth of bureaucracy, the jungle of policies and "sort-of" policies that can be weaponized, the emergence of an inner clique skilled in navigating the quagmire, and the resulting impossibly high barrier to entry for new contributors.
The experience of contributing to Wikipedia these days is inherently adversarial. People come and write something genuinely useful or at least do so in good faith, and then realize they unleashed on themselves a whole lengthy process when they have to defend what they did on numerous talk pages over an extended period of time, or succumb to deletionists. Even if they prevail, the experience is hardly pleasant, so they eventually leave and never look back.
An artefact of the above is also that the editorial base suffers from an overrepresentation of people with vested interests, as they will always be the most highly-motivated to stay. There are of course many great editors too but the way things are going, the project is not sustainable, and its quality is being compromised due to an organizational failure.
What Wikipedia needs is not Markdown (not that I have anything against it) but courage, fresh air, and less of the siege mentality.
Some years ago I served as a CTO for a major news media org.
Google approached us and asked to implement some 'structured web' stuff on our pages, so I dug in and realised there is a mandatory field for a link to a book or movie reviewers personal page in wikipedia.
So I've found out that one of the most notable Russian book reviewers does not have an wikipedia page. No big deal, I've created one with some very basic info and some links to prove she's 'notable enough'.
What ensued shortly after is kafkaesque story of some entitled shitheads who have power because they know some local lingo and have unlimited time resources to waste.
Luckily, someone poured too much time and effort and made that page 'standard wikipedia article about a person' with all those 'templates and boxes' and used the right lingo to defend it's notability.
I swear I tried to do it myself and failed. I am not stupid, the wikipedia editing is just too complicated.
Thanks for reading, a few years passed and apparently I still have not got it fully out of my system. That feeling of complete helplessness, when those bureaucrats are clearly toying with you and enjoying it. It sucks big time.
Whereas every major league Bangladeshi cricketer is allowed a page? It's exhausting. And my Wikipedia account is old enough to vote. The only place you can make articles easily is when everyone involved is dead.
I work on Wikipedia articles in niche worldwide arts and history topics. I just see few people caring enough to debate me about Uyghur modernist poetry.
I had a slight bit of trouble recently when I wrote a script to generate stub articles for Indonesian films. I got some flak for creating articles without enough extra detail (just the name, year, cast, genre, and awards), but only one was temporarily pulled down into the sandbox area.
I recently got someone who I wrote an article on to release their photo to the public domain. I had to have them email Wikimedia Commons a legalese template to do that. But the person on the other side said, "how can you release a photo of yourself"? (It wasn't a selfie.) So he had to get his daughter who took the photo to send them the legalese. It's a bit much, but I see how they have to be careful with legal things. It wasn't the smoothest experience. It's much easier to release the copyright on your own works.
Is getting more contributors involved a limiting factor for wikipedia? Wikipedia's goal isn't necessarily to maximize the number of contributors, or articles. They aren't facebook.
Ultimately, wikipedia is still pioneering its space. There isn't really anything else like it, and all considered, it works remarkably well. I don't see how any version of wikipedia could really avoid adversersarial dynamics. It's those adversarial dynamics that make decisions, even though clear policies or authorities don't really exist. Sure, many will find it alienating. Most people don't need to edit wikipedia for it to work though.
I think what we need is other wikipedia-like things. The ideal solution to a lot of deletionists problem would be more highly active sites that work in different ways.
Wikipedia predates markdown by several years. The only truly “universal” markup language I know of at the time was bbcode.
Even if markdown has been designed, it’s lacking support for templates and tables that wikis make heavy use of.
Every wiki software I’ve used has had a custom language for one reason or another. Usually good reasons.
I stopped reading the article early on.
https://en.wikipedia.org/wiki/The_purpose_of_a_system_is_wha...
"Parsing ... is the process of analyzing a string of symbols ... conforming to a formal grammar." [1]
Wikipedia content is most definitely not a formal grammar in any usable meaning of the terms. This is why it's a nightmare to work with, and why it's not parseable using the correct meaning of parseable.
The hard part is in things that are impossible in markdown.
Yeah it would be cool if you could 'inject' arbitrary HTML into markdown to make up for its lack of support for tables
https://riptutorial.com/markdown/example/1741/creating-a-tab...
Once you start extending markdown to support all of the features Wikimedia supports, you end up with the same problem.
If mediawiki syntax was as feature-limited as markdown, it would be equally trivial to parse.
-wikitext unparsable - wikitext is a bit insane but there exists a parser called parsoid. If we couldn't parse it it would be impossible to make a visual editor
- most pages are redirects: not sure what the problem is.
-boilerplate text - there is a template system to have repeated text in only one place. Not sure what the issue is.
- he goes on a rant about no search without giving much context. There is a search feature based on elastic search (older versions were based on lucene directly). In my opinion its a pretty decent search engine (especially compared to most sites that make their own search). I'm not sure what the actual complaint is.
-complaints about wikidata - this is more political than technical, however "If wikidata was a company, it would not exist anymore, and you wouldn't have heard of it." seems patently false. Wikidata is pretty popular even outside of academia, and is used quite extensively.
- category tree being a graph not a tree - that's kind of unfortunate but what exactly is the problem here. Its a problem on commons, but i've never really seen how its an issue in practise on wikipedia. Complex categorization is being taken over by wikidata anyways.
Template ecosystem is complex - it could certainly be better, but the complexity here is a trade-off allowing more flexibility and allowing the system to evolve.
Inclusionist vs deletionist: no comment
UI design: perhaps a fair point here, although i do kind of like the stability of the current design. Most of the modern web sucks imho.
Moral failure: i dont want wikipedia to fix the world. That's not its role. Its job is to document, not to partake.
Viz editor not being enabled by default: i agree, although i think it was pushed too hard in early days when there was still kinks, but its long past time now.
To be clear, i definitely dont think wikipedia is perfect, i just disagree with some of these specific criticisms.
With respect to this point, I think the author is saying that it's a dated system that introduces a ton of potentially unnecessary link rot. Basically, there's a ton of unnecessary technical debt because there isn't a more automatic way to set up/maintain redirects and instead often multiple manual entries are added for e.g. common misspellings of article names.
I probably bias myself by being too used to the status quo, and thinking its better than it really is. As a computer programmer, i probably underestimate the problems with the complexity of wikitext - since after all, a user facing markup language should not be anywhere near as complex as a programming language but there are times where it does feels like wikitext is basically spaghetti code.
Anyways, cheers.
Freebase, DBpedia, and many others (me) have tried, but the reality is that the markup language is poorly defined and the only path that is really tested is the one that ends up rendering HTML.
If you feed HTML from Wikipedia into a web parser that supports the DOM (say Beautiful Soup) you can generally parse out what you want pretty effectively. Once I switched from the "markup rabbithole" to "parsing standard HTML" I was able to turn my MediaWiki extractor into a Flickr extractor in about 15 minutes.
https://en.wikipedia.org/wiki/Thomas_Sowell
I went and used this
https://addons.mozilla.org/en-US/firefox/addon/openlink-stru...
and the JSON-LD data comprised of "wrapper metadata" for the document, a photograph of that person, and a link to a concept in Wikidata where you would find some of the facts that you'd be looking for if you mistook the wikipedia page for the person.
The page at Wikidata for that person
https://www.wikidata.org/wiki/Q553449
has the birth date, gender and other facts such as he taught at both Cornell and Stanford and also links to identifiers for that person in many bibliographic databases.
What matters to Wikipedia visitors is content quality (and managed depth) and reliability. All the rest is froth. Layout? Are you kidding?
Could it be better? Hell yes. It'd be great to see many articles, written by committee, visited by a professional editor with years of proven experience -in that category-. The Foundation's got the money to pay one per category. Each has to visit 5 articles per day. Each they finish is locked and marke 'pro-edited in (the year)'. Make that fact a search filter. ('Administrators' choosing? Brrrr.)
Then there's Wikibestia.org. Every year, copy a limit of 1 million articles there. Chosen how? Poll the visitors, it's theirs! 'Wikibestia ... the people's choice!' 50 net downvotes? Guillotine! Big contest, media advertising, whatever.
Articles that can only be read by subject-matter experts completely miss the point. They're just there for vanity or whatever. Flag them, give them one year to move to Wikiexpertia, while versions for non-experts are prepared (or not), then delete.
Map resembles terrain!
https://en.wikipedia.org/wiki/Alfred_Korzybski
I knew the expression (and was obviously alluding to it), not its origin.
Thanks!
I think the biggest problem he mentioned is the "exclusionists." There's no harm in including a page on every local elected official, every Nobel prize winner, even every band that has a reference or two somewhere. Most likely it will remain a little article, but so what? The history of small local bands, for example, would be very interesting 100, 200 years from now. (Wouldn't it be fun to read about some town's local musical groups in 18th century America?)
The counter to which is, of course, that Wikipedia doesn't have to be run by a relatively-tiny number of highly-dedicated editors. Over the hill somewhere just out of sight is a set of technologies, policies and social norms that would lead to still more people participating.
You can put bounties or some kind of incentive for people to take on the conversion, but I think if the new format is better, a number of people will feel strongly enough about it to want to convert each article when it's updated. And even if 12 years from now some articles aren't converted, so what?
But that's fine. It'll create more innovation as people ditch it and create better tools. The one part of the article I disagree with is where they say it's "unparseable", which obviously it's not. It's parsed millions of times a day. Even by third party tools that harvest it for semantic data.
That means it can be transformed and used to seed other databases using non-terrible description languages like Markdown.
It's untrue that they haven't added features over the years: they have, but they do it at a glacial pace. Consider one of the recent ones: when you hover over a link, a box pops up and shows a preview. They used to render math on the client-side, but now, they just convert to SVG: as a result, it loads instantly, and renders consistently. The WYSIWIG editor rollout was slow, because people felt that it would attract low-quality low-commitment edits. They first released it as an optional feature, and then turned it on by default, when they were confident that it worked as intended. Oh, and my favorite? Allowing an article to start with a lower-cased word (say iOS); I remember that there were a bunch of redirects just to correct for this deficiency.
Yes, it is a giant pain to edit some pages in that arcane syntax, but nothing else even comes close in terms of features.
Yes, there are an enormous number of templates, but in practice, an infrequent contributor just finds a page that uses a similar template, and copies it out.
Yes, there are lots of bots, and they try very hard to guard against spam, without making you sign up or even solve a CAPTCHA to edit. Plenty of bot edits are "good" edits: they revert rage-rewrites, rage-deletions, and all kinds of malicious user behavior.
What you don't understand about redirects is that, the good ones can't be automated. It's not a string-matching problem. Yes, they could automate /some/ of the redirects, and they try. I've personally never run into a typo-redirect in recent years.
Yes, it can get political at times, and it's _very_ difficult to have objective guidelines about which pages are worthy of existing. Politicians' pages often get locked, when there's an upcoming election, and this means that you need an account to edit. Again, MW has lots of great features.
Wikipedia is aging, and nobody can deny that, but who would want to do the thankless work of parsing the markup and porting it to another system, AND correct the breakages? What commercial value does it have, and who's going to fund it?
this reminds me of the library in the david brin uplift books, where the myth is that it contains perfect knowledge but in fact there are sinister memory holes
(not saying it's sinister in this case)
To be more accurate, those are the editors on the English Wikipedia (the largest one). The French Wikipedia, for instance, have this visual editor enabled.
Markdown is a decent language for basic stuff like Reddit, Stack Overflow, GitHub READMEs, etc. Stuff where the most important feature is making it easy for people to make content.
Places like Wikipedia or blogs are where Markdown doesn’t work so well. Sure, you want it to be relatively easy for people to edit Wikipedia pages, but it’s worth sacrificing ease of editing because you want to improve the reading experience.
Honestly, wikipedia would be a great proving ground for a hypothetical evolution of LaTeX
I do believe there exist people who never look up information online. Such people must exist somewhere and thus I should not exclude them from existence. For the rest of people I strongly believe they either use a search engine, a virtual assistant, or go directly to Wikipedia. The search engine scrape Wikipedia, the virtual assistant use the search engine to find answers and Wikipedia is Wikipedia. In either case it is highly likely that one is using Wikipedia directly or indirectly.
However, the reasons for the discrepancy are different. Many articles are defended by petty tyrants, against improvement by people who know better than they do, but have less free time to spend un-reverting work. Often enough the tyrants are self-righteous tenured or emeritus professors who imagine themselves holding back the deluge as they enforce interpretations they picked up decades ago and have not kept up on.