Wikipedia is now drawing facts from the Wikidata repository
gigaom.com
gigaom.com
There's plenty of opinion of the deletion-happy (if that's the operative term) policy of the admins, especially as it pertains to notability, and I think this is a common complaint : http://www.highprogrammer.com/alan/rants/wikipedia-delete.ht...
I do something that may seem like trying to fix a leaky dam with chewing gum, but every time I see an article nominated for deletion, I copy the source to a private wiki I'm running. Sometimes, the deletion nomination goes away, other times it does get deleted, but at least then, I have a copy.
But we also have to keep in mind that there are plenty of other resources on the web specifically aimed at niche resources. E.G. Wikia. I've lost count of how many comic book related articles were deleted only to show up on the DC or Marvel wikis. ( http://dc.wikia.com/ http://marvel.wikia.com/ )
Likewise, it's not unreasonable that a lot of content gets missed by the editors, since it's a big place and most editors have one or two areas that they focus on. As the editing guidelines state, when in doubt engage in dispute resolution, not edit wars. If you can make a good case for why an article should be there in the first place, be persuasive in the talk pages.
So a few pointers for people getting angry at Wikipedia:
First, ask whether the article is a good fit for the wiki. Can it go on a blog or a niche wiki (like a dedicated Wikia) instead? The Pokémon example is a bit extreme, but I think many of those pages may get deleted or merged. There's always the Pokémon Wikia : http://pokemon.wikia.com
Second, notability is a very tricky thing. Reputable sources may be even trickier. Rather than debating notability, focus on reputable sources (since that's the biggest hiccup for references). If you can link NY Times articles or BBC or some other news source, rather than just community sites, other blogs (depending on popularity) etc... you'll have a better chance of getting the article/section through and staying there.
Third, well written articles have a better chance of surviving than those that give off the 2-3 paragraph stub vibe. The more reputable citations and best content structure you can give, the better the chances an article survives (this may partly explain the Pokémon pages too). If you have trivia, try to merge it into the content body more, rather than list at the end; that feels tacked on and superfluous, if not directly supporting the main content.
Fourth, try to be a bit more empathetic to the goals of the Wiki while being objective to the subject when asserting your views (especially on controversial articles). How you word things is a big hint as to whether that content will remain or get scrubbed the next day/hour/minute.
Now, I hope can go back to discussing Wikidata and how Wikipedia and everyone else will benefit.
http://archive.is/EOZqQ - The Imaginary Theatre ( post deletion obviously :( )
http://semantic-mediawiki.org/wiki/Help:Introduction_to_Sema...
If not, doing that should be possible using an NLP library that does NER. Along with heuristics, one could use the list of currently existing articles as a seed.
Edit: Of course, if all you're trying to do is link to existing pages, then you don't use the set of existing pages as a seed, you just use them as the list. But if you're trying to extract "entities" that don't have WP pages yet, then you'd still want to fall back to other NER techniques, which include various heuristics and what-not. Whether or not there would be an value in that is an open question, I suppose.
Also, if you're interested in that sort of thing, two other projects you might find interest are:
and
Both involve extracting semantic meaning from unstructured data. It's pretty cool stuff.
It would be a fun project to try and determine the correct link based on the context.
When a search term matches a tag, and none of the tagged pages have a clear "majority probability" of being correct, it would display a list of all pages with the tag, in order of popularity.
Both of these things would be amazingly annoying to the majority of Wikipedia users.
But they could be disambiguated by humans, which is my point. Humans understand context.
Whether or not something like that would be a net win for Wikipedia is up for debate I guess. That said, I think they already do have a bot that can do at least a limited amount of auto-linkification, but I can't swear to it.
It would take some AI to work it out, unless using the context around the link. "Sun" could refer to many things, its not really possible for wikipedia to know which one you're on about. So links are still done manually.
Edit: found the link (PDF) http://www.cs.mtu.edu/~nilufer/classes/cs5811/2003-fall/hilt... Here's the actual quote:
“Do you mean Anthrax (the heavy-metal band),
anthrax (the bacterium) or anthrax (the disease)?”
“The bacterium,” was the typed answer, followed by
the instruction, “Comment on its toxicity to people.”
“I assume you mean people (homo sapiens),” the
system responded, reasoning, as it informed its
programmer, that asking about People magazine
“would not make sense.”I'd be curious how good the results are. I've found a bunch of articles, but no live demo. If someone set up a small-scale version where you compare the auto-linked version of a few hundred Wikipedia articles with the existing manually-linked version, I think that could convince people it was worth adopting (if the results looked good).
Except, of course, any mention of the name of a listed company in a financial/business publication.
The Free Dictionary uses something like that - try double clicking words within the definition: http://www.thefreedictionary.com/link
You could seed the database with famous people's family trees from Wikidata. The Mormon church also has lots of genealogy data that (perhaps :) they might share for not-for-profit use.
The biggest challenge would be preventing trolls and spammers from uploading false data. I've sketched out some rough ideas where family links can be "thumbs up'd" bidirectionally by people on both sides of the connection, but not necessarily the immediate people.
There's no open source genealogy programmers because the young nerds that do open source dont care about genealogy.
Not to mention a security issue as well since most financial institutions ask for your mother's maiden name as a security question.
http://pauillac.inria.fr/~ddr/GeneWeb/en/
There are others. I guess there aren't many, but there are some.
This seems to be more "all wikipedias can now draw facts from wikidata", and certainly isn't "all facts in wikipedia come from wikidata." The former is cool, the latter would be mind-blowing - but I'm not sure how far along we are on the path to "{all|most|some|a few} facts come from wikidata"
It looks like the ability to pull Wikidata into infoboxes was just rolled out a week ago, so I guess it wouldn't actually be used yet in a lot of places.
Deletionism is the reason it's worth using to begin with. It prevents it from being flooded by crap and allows it to stay on mission. Not everything has to be all things to all people.
Wikipedia isn't one big giant article, where the existence of marginally noteworthy elements distracts you or diverts your attention. The articles that get deleted by the little notability hitlers are contributions with no cost incurred by anyone, save the trivial storage space.
It's like saying some guys tripod page that you never visit is somehow cluttering your experience of web browsing.
I wish they'd err on the side of leaving stuff alone in the 1% of cases where it's not black and white (like the obsolete BBS software example)
Godwin, and wrong on the merits, as well.
The stuff that gets deleted is the stuff that can't be verified, which means there's no way to fact-check it. It isn't about storage space or clutter: It's about not having stuff in there that can't be verified.
Wikipedia is nothing more than the biggest plagiarist / content farm on the Internet. It isn't scrutinized because it has been grandfathered in.
Wikipedia is the ebaumsworld of information. Completely unreliable. Steals credit, traffic, royalties from the content creators. Policies focused on self preservation rather than serving a public good or respecting creators.
The delicious irony of your hatred is that your point is so poorly argued, you must be a wikipedia editor.
I'm actually a pretty accomplished wikipedia editor with several original article credits. The articles are standing today. I have barnstars and everything. But I actually submit my articles to Encyclopedia Britannica now, because I realized the truth about wikipedia. It's just a really low quality content farm that can't be trusted on anything.
What is the point of getting your knowledge from unreliable losers? All the biggest wikipedia editors are no-life losers with zero respect in any real intellectual community. They are divorced from the community of experts and receive nothing but scorn from them.
What's the point of reading an encyclopedia written by people who you can't rely on? Getting 80% of the content right is not an achievement--every content farm on the internet does the same--from eHow to expertsexchange.
Wikipedia is just an ideology and volunteer driven low quality content farm. Due to Wikipedia's overpowering marketing/SEO, when Wikipedia writes an article on a topic, that Wikipedia article will now have higher visibility than the original information source that it scraped and now cites. The Wikipedia article will steal traffic from the original source.
The internet would be so much better without content farms. And Wikipedia is the worst of the worst.
When I feel like writing an encyclopedia article, I send it now to Encyclopedia Britannica. They've edited my writing and incorporated parts into their high quality encyclopedia.
By the way, that study that said that wikipedia was just as good as Britannica was complete bullshit--as flawed as your informal fallacy that you just shat out right above this comment.
[1]: http://www.gwern.net/In%20Defense%20Of%20Inclusionism#the-ed...
It's true the admins take notability pretty seriously, but to be fair, so does Encyclopaedia Britannica. Wikipedia folks freely admit that they're elitist in that regard, but I don't think that's necessarily a bad thing. They set it to that level for the sake of maintaining quality and relevancy (which is a measure of notability) to an acceptable degree.
The foundation's ultimate goal is to make a reliable resource for knowledge after all.
What we need to "balance it out" if you will, is maybe another entity(ies) with a broader scope and maybe a looser threshold for notability, but hopefully with more reliability with expert verification ( Citizendium comes to mind : http://en.citizendium.org , but it's no where near as comprehensive ). Still it's not easy to run something like Wikipedia, so whoever will do it would have to have deep pockets and seriously dedicated.
The complaint about wikipedia is that it fails to live up to its ideals. If you had ever been a wikipedia editor you wouldn't talk the way you do. Even as a domain expert it's intolerable to deal with the wikipedia amateur hour culture. It's far more rewarding to submit your articles to Encyclopedia Britannica.
wikipedia'd!
you might enjoy Deletionpedia: http://www.deletionpedia.dbatley.com/w/index.php?title=Main_...
I can see why most were deleted, I chose to view a few pages at random and they were:
- a really non-notable musician (probably self-promotion)
- a hoax ("The Independent City-State of Sonora")
- a game guide to Super Smash Brawl, this is probably something that should have gone on a blog / GameFAQs
- a witty bio about a non-notable person (probably self-bio, or a friend's)
I believe the four of them were deservedly well deleted. Two of the cases, the musician and the game guide, should be hosted elsewhere (personal blog or website).
Under a critereon of "notability" that works great for dead trees but is largely irrelevant to an online primarily text reference.
It'd be nice if the WikiMedia projects had a proper GitHub presence - it's hard to get a sense how plausible a self-hosted version of this is.
Insignificant revenue. The click-through rate might very well be so low that only a small amount of revenue would be brought in by the ads. It would not be worth barraging thousands of readers with ads for only a few pennies of revenue.
Ads cheapen the encyclopedia. By their very nature, ads are biased content intended to influence people. They are thus diametrically opposed to the goals of a neutral encyclopedia intended to inform people. They would cheapen the encyclopedia in the eyes of many readers, as evidenced by the numerous anti-ad comments received during every donation drive.
Contributors may leave. Many contributors vigorously oppose ads (see the forking of the Spanish Wikipedia, 1, 2, 3, 4), and in 2009 the Wikimedia Foundation promised to keep "Wikipedia. Ad-free forever." Since about 2002, Jimbo Wales has repeatedly stated that he opposes all advertising on Wikipedia as well. Based on these statements, some editors have probably contributed with the understanding that their content would not be diluted with ads. Changing the long-standing no-ads policy now could reasonably be perceived as a bait and switch tactic. Numerous contributors are likely to leave as a result and new ones are less likely to start. Contributor goodwill is Wikipedia's main asset and should not be gambled with.
Annoying and distracting. Readers come to us for encyclopedic information, not for ads. Ads have to be processed by the brain (if only subconsciously) and therefore distract and annoy. "The free encyclopedia" also means: free from distractions and annoyances.
Privacy violation. If an ad consolidator such as Google AdSense is used, the privacy of our readers is compromised. The consolidator will invariably learn which Wikipedia articles a given IP address reads or searches for; they can then correlate that information with other data they may have about that IP address (e.g. Gmail account).
et cetera et cetera et cetera.