This is something they should work on. But as every club that get successful and was open (I attended many Wiki/Wikipedia meetups 20y ago) it gets more restricted and close to keep the masses out. What a shame.
"Trying to edit something" is a bad idea these days - the content is far too developed for that. Instead, post on the talk page with a description of your suggested edit, wait for people to object/complain/discuss, refine the text on talk if needed, then edit. It's pretty indistinguishable from a "pull request" workflow on github.
Yes, I know that a lot of novice-focused content advises people to edit directly. That's fine for fixing typos and bad grammer, not for substantive work.
Second if this is the way to go - and I don't agree [2] - the UI should reflect this workflow and not propose a different one that mostly fails.
[Edit] PRs also assume some kind of owner and committer relationship, which might be right for open source projects but is not the case of Wikipedia - even if many editors on Wikipedia consider the content "their own".
[1] https://en.wikipedia.org/wiki/WikiWikiWeb - "Cunningham's idea was to make WikiWikiWeb's pages quickly editable by its users, [...] . The usefulness of Wiki is in the freedom, simplicity, and power it offers."
[2] Wrote a quite successfull wiki 20y ago called SnipSnap, the engine of which started Atlassian Confluence, and did several years of research in using Wikis back then.
That's absolutely not how it's been presented for the last couple decades, though.
Wikipedia recommends doing things, not talking.
Articles are also graded, I don't know what article GP is talking about but I'd say anything below "good article" is fair game.
Neither is true for Wikipedia. There are no clear owners for each article so the editors are more or less hidden in the shadows (from the readers and casual contributors, at least), instead of being front-and-center of an article and explicitly taking responsibilities. It is also very unclear why they get to “own” the article like a GitHub repository; are they the initial author? Major contributors that took the baton from the original author? In many cases they are, but Wikipedia does not make that relation clear, so the editors look like self-appointed to an outside eye, which is why people have issues with gate-keeping (not the gate-keeping itself).
Making things worse, an “outside” contributor also has no chance but to work with those (seemingly) self-proclaimed maintainers; only one article can be under an entry on Wikipedia, your own modified copy has no chance competing with the original, even if you managed to make other people aware of it and agree it’s objectively better. That’s competitive imbalance, another trait of undesired gate-keeping.
I’m not saying what the Wikipedia community is doing is wrong. But if that’s how they want to do things, they need to fix Wikipedia to reflect what they want the process to work. The current system is broken.
Also, I know companies or celebrities who pay for their Wikipedia pages. You can find this service in 5 minutes by googling.
"Profit-seeking and mindless consumption" can be a motto for the current WMF.
For the smaller ones, that is warranted - for example, the Croatian Wikipedia is infamous for its bias towards nationalist positions (https://en.wikipedia.org/wiki/Croatian_Wikipedia), something that's only made worse by the aftermath of the Balkan wars decades ago... as a half-Croat myself, I can only say: let's keep it at "the situation is complex and sucks all around, Wikipedia is just one symptom of a much deeper problem".
For bigger projects such as English and German Wikipedia, sheer numbers make outright disinformation spreading much more difficult - there the problem is different: what you see is usually factually correct and neutral-ish, but there is a widespread consensus that Wikipedia editors skew towards white, male and somewhat privileged (https://en.wikipedia.org/wiki/Wikipedia#Coverage_of_topics_a...).
Personally, I trust scientific and pop culture related articles in most if not all versions of Wikipedia - but keep a healthy skepticism for everything political, since there it is the hardest to judge completeness and accuracy of any subject.
They do, particularly in educational contexts - getting students interested in Wiki projects and having them write review articles with good sources as "homework" that can then be posted to Wikipedia. But supporting that takes money as well!
Money won't change that.
Example 2, there was two mass shooting in March:
- https://en.wikipedia.org/wiki/2021_Boulder_shooting
- https://en.wikipedia.org/wiki/2021_Atlanta_spa_shootings
In the Boulder shooting, all victims where white and the shooter was targeting white people but they don't mention any of that.
In the Atlanta shootings, the article goes out of its way to mention 6 of 8 victims where asian (asian-hate narrative). In reality, many spa owners are asian.
Theres is definitely a bias in their editing.
https://www.usatoday.com/story/news/factcheck/2020/07/15/fac...
Where can I learn about this?
Roughly, English Wikipedia hit its editor peak around 2007 and has been very slowly declining since then. Various people argue about various causes. Theories I've heard include wikipedia becoming too rule heavy, to community just not scaling past that point, to the internet being a lot more interesting now then it was in 2007 so more distractions, community being mean/troll-filled now, everyone using mobile phones which are crap for writing long form content, all the easy to write articles are written, etc (Probably missed some).
It's probably less satisfying to change the wording in one sentence in an article about an Australian cheese than it is to write that article. So fewer people care.
https://slate.com/technology/2012/07/kate-middleton-s-weddin...
Having 10% of new editors stick around seems amazingly high.
Has anyone done the same for post 2015?
https://stats.wikimedia.org/#/en.wikipedia.org/contributing/...
I could be wrong but today it seems that most editor efforts in the English Wikipedia goes towards news, politics and media. It would not surprise me if those have a natural lower retention rate than other form of articles.
On the other hand, the grown amount of content means there is proportionately more content to maintain—which might take even more effort and attention to detail than writing new content.
That, as well as the reputation and pervasiveness of links to Wikipedia, together make it more vulnerable to and a more compelling target for instances of blatant or stealthy misinformation, censorship and vandalism.
On that topic, I suspect a resource that tracks change histories for sensitive Wikipedia articles[0], aggregating and correlating changes by IPs and usernames along with some smart NLP analysis on text contents of the diffs, could end up being insightful.
[0] Many of us know a certain date coming up, publicly remembering which starting this year is already punishable by a prison sentence in Hong Kong. Maintaining the article covering this one is hopefully taken care of for now as it is pretty conspicuous—but for how long, and how many lesser known articles are there…
It's a very sharp peak. What happened in 2007?
People donating to wikipedia only want wikipedia, they are unaware of the other stuff wikimedia does.
Those projects would crash and burn though, because frankly, no one cares about them. Stuff like wikidata and SPARQL endpoints are an academic curiosity with no real value.
The semantic web won't happen, articles do and will.
In addition to that, the whole deletion first and no primary information attitude makes sentences like "ecosystem of free knowledge" laughable.
If you just optimise for today, you will die tomorrow.
Gradual improvements won't get you out of a local peak when that local peak is swallowed by rising tides.
If your organisation isn't doing things that look pointless today but in a principled way, I don't think it has great prospects of long-term survival.
Heck, even the wiki concept itself started out as this long-shot experiment nobody at the time expected to be successful.
The semantic web is "tomorrow technology" for 20 years now. At what point does tomorrow technology become yesterdays vaporware to you?
Piling onto the glorified garbage dump (just follow the isA edges of any concept you'll see what I mean) is a strawman metric.
On the contrary, Wikipedia itself has a lot of reliance on the other projects. For example, Wikidata solves the problem of interlinking Wikipedia languages and other projects when they're hosting content about the same real-world entity. In fact, this was Wikidata's MVP - the reason it got funded in the first place. It then took off from there.
Wikimedia Commons is a similar story - in addition to the obvious languages issue, there are a lot of concerns about media (such as licensing. enhancement etc.) that really are best addressed in a dedicated venue, with its own committed contributors. This has been very beneficial to Wikipedia itself.
Wikidata isn't just used for interlanguage linking. Its also used to keep infoboxes synced across different languages, and a way to easily query that data (Depends a bit on the language, english wikipedia isn't fully on board with the system, but its been a big boon to smaller languages).
As a result, people can update these "facts" independently of language, keeping them altogether more up to date, and it allows the data to be queried independently (via https://query.wikidata.org - try some of the example queries (button top left) if you haven't, the level of power SPARQL gives for querying this type of data set really is very cool).
Part of Wikimedia's mission is to spread knowledge. Allowing querying of factual data with a query language like sparql helps advance that goal by letting people use extract the knowledge for new purposes.
Like I said, gross overengineering.
"You can save cents and seconds of time updating knowledge boxes, by spending millions and hours developing federated SPARQL query capabilities!"
And since the data happened to be there anyways, we added an RDF export feature, and threw that into a BlazeGraph DB so that people could do SPARQL queries if they want. The (not really federated) SPARQL queries have basically nothing to do with the central-place-to-store-facts that are used in multiple places, feature, but it was pretty cheap to do once everything was in a central place.
Edit: And thank you for some insight into the back end of wikipedia - I did not know multiple languages was such a big push and it is such a sensible direction . Keep up the good work and keep pestering me for cash each year.
There's a very old architecture doc from 2012: https://www.mediawiki.org/wiki/Manual:MediaWiki_architecture its kind of outdated and missing a lot of the newer pieces.
p.s. I used to work for the Wikimedia foundation about 1.5 years ago, but don't anymore. Prior to working there I was a volunteer for the FOSS project, and in theory I still am but I honestly don't do much anymore. My opinions are my own and don't represent WMF or anyone else.
How often is numerical (which is the only knowledge you can really automate this way) knowledge updated? How often is said knowledge dependent on cultural idiosyncracies? How much effort is the code maintenance of the specialised editing system? How much effort is maintainance of the centralised infrastructure. How much effort is it really to change the data in multiple places, given that enough eyes every bug is shallow? What are the oppirtunity costs of not being able to properly distribute and share articles because they are suddenly tied to a very specific codebase? How much infrastructure cost do I have because I can't get volunteers to host my data?
Overengineering solving the wrong problems at it's finest, no wonder the software has stayed this bad for decades now.
Numerical knowledge is not the only knowledge that is stored there.
I'm going to go with update rate of often. It has a higher edit rate than english wikipedia does.
> How often is said knowledge dependent on cultural idiosyncracies?
Sometimes that is true, other times it isn't. The system is used where it makes sense, and not used where it doesn't.
> How much effort is the code maintenance of the specialised editing system? How much effort is maintainance of the centralised infrastructure.
There is certainly some. I wouldn't say its dominating, but it certainly exists. There is plenty of centralized infrastructure for other things too.
Prior to the introduction of wikidata, more complex fragile systems existed to try and keep things in sync in a more manual fashion. They also had a cost.
> How much effort is it really to change the data in multiple places, given that enough eyes every bug is shallow?
A lot. And we don't have that many eyes in lots of languages. There's a lot after all: https://en.wikipedia.org/wiki/Special:SiteMatrix
> What are the oppirtunity costs of not being able to properly distribute and share articles because they are suddenly tied to a very specific codebase
Not sure what this has to do with it. Articles are written in a custom markup language. That's already tying it to a specific code base quite heavily. The wikidata dependency is pretty trivial, and there are open apis to get info out of wikidata, and regularly made dumps of all the data in wikidata if you want to copy it. If that's not good enough, just copy the rendered version instead of the source.
> How much infrastructure cost do I have because I can't get volunteers to host my data
I'm not sure what you're talking about. Wikipedia was never hosted in some magical peer-to-peer fashion on volunteer web servers, nor is there any situation where it is likely to be hosted that way. The goals of (near) instant update time, as well as some central quality control and user blocking, makes distributed hosting pretty impossible.
Everything wrong with wikipedia summed up in one paragraph, nice!
But disagreing with wikipedia's goals as a product is very different from a claim that the solution intended to meet those product goals is over-engineered.
How are centralization, real time editing, and user blocking properties toward the above goal?
You probably know full well that network effects are prohibiting any other collaborative knowledge base from from ever croping up again.
All big companies link against wikipedia because the network effect, all funding and donations go to wikipedia due to the network effect, all editors spend their time on wikipedia due to the network effect.
Wikipedia is eating it's own childen, and if we ever want to get something better, it either needs to grow up and start acting like a responsible adult (which we know it won't because at it's core is the failed web equivalent of the stanford prison expetiment), or die and make space for a new genration of tools.
And if you have some magically technical solution that is much better and much easier, then please share it. And don't just complain that everybody else is doing everything wrong.
It's not that easy to work with, but there aren't any good alternatives in my case. Freebase died. DBpedia doesn't compare. Domain-specific databases don't have the data I'm looking for (or it's woefully incomplete and/or hard to link to external entities) because it straddles multiple domains, not to mention they're usually not free. I for one am very thankful for Wikidata existing and seeing good activity.