Meta is building an AI to fact-check Wikipedia
singularityhub.com
singularityhub.com
Here is what the model is actually (attempting) to do:
>It calls attention to questionable citations, allowing human editors to evaluate the cases most likely to be flawed without having to sift through thousands of properly cited statements. If a citation seems irrelevant, our model will suggest a more applicable source, even pointing to the specific passage that supports the claim. Eventually, our goal is to build a platform to help Wikipedia editors systematically spot citation issues and quickly fix the citation or correct the content of the corresponding article at scale.
You can't use it to verify citations to books.
Edit: To clarify it could replace a citation to a book with a web page assuming that information exists on the internet.
Nothing that can't be done in principle, just hard/tedious engineering
Since it is open-source, I'd imagine some researchers with computer vision expertise and skilled with document/text digitization could greatly enhance this AI by digitizing these old books and including them in the training.
PDFs are a pain in the ass to read, especially on mobile
Ah, the renowned human editors/fact-checkers/moderators that Facebook employs, great.
I know, right? The halo effect is strong in this thread.
Ive got no love for FB or "fact-checking", but I knew much of this thread would be a dumpster fire.
Even a pretty poor model (and it probably will be) would be a huge boon to editors in an encyclopedia full of references which are missing, out of date, a bit stretchy or plain bad faith. Wikipedia already has bots on it doing grunt work, and plenty of editors that like doing grunt work of checking minor details more than writing new content.
We're doing similar things to find errors in our regular ML datasets - a large proportion of the examples the model can't predict are mislabeled. Those mislabeled examples have a big penalty on performance. Since Wikipedia is often used in ML, it was time to clean it up.
(The video's context is the sinking of the German battleship Bismark. Whether or not you've any interest in that subject, the video can still give you a far better sense of the problem than Singularly Hub's article.)
But. Compare to the rest of the internet. Compare to every single propaganda website. Text on Wikipedia has one of the highest chances of being true by default. If a random website contradicts Wikipedia one shouldn't trust that website.
I'm sick of people comparing Wikipedia to peer reviewed journals... When instead people get their knowledge from tabloids, random newspapers, individuals getting mad in youtube videos, and websites like Facebook.
If Facebook claims to know more about the world than Wikipedia, it should factcheck itself.
The lesson teachers should have been teaching was that you shouldn't cite the thing in a scholarly work as if it was a primary (or even secondary) source. Just like you shouldn't cite the Encyclopedia Britannica, or any other encyclopedia, in a paper.
Wikipedia, by contrast, is written by whoever can engage in the Wiki-world equivalent of getting the most likes. And its writers are whoever wins what is effectively a glorified Reddit-like upvote system. Ostensibly this is balanced against the fact that editors cannot themselves work as primary sources on Wikipedia: Einstein himself would not be allowed to write about physics on Wikipedia, but this also means Joe Editor also can't write as a primary source.
Instead everything is exclusively sourced from other sites. The problem is that this is where bias comes into play. If Joe Editor wants to publish that 2+2=5 in the mathematics section and he can find an article on The Verge claiming so, then he can do so. If somebody wants to correct him, and he doesn't want to change, then edit wars begin. And in these cases, the truth itself is secondary to more normal aspects of social media. This is where Wikipedia fails.
[1] - https://en.wikipedia.org/wiki/Encyclop%C3%A6dia_Britannica#C...
In spite of occasional quality concerns, Wikipedia is one of the most trustworthy sources on the internet. That does not mean you should blindly believe everything you read there, as for any other source. Also, as stated by others, it is not a scholarly source nor a scientific publication.
So most of Wikipedia isn't decided by Wikipedia's ritual, but just by whoever put in the effort to write it.
Which is about as bad. This can directly translate to "the only person that put effort into writing it, no matter how competent they are in the subject".
For as long as possible people held off any kind of voting, instead using a rough consensus system. This was chosen because voting would lead to too much inaccuracy.
I know about this because at one point I researched how and why the system was working, and helped write some of the documentation on it.
My question to you is where, when, and/or why you think the system looks like reddit voting, as you claim? I would be rather sad if it has broken down of late.
note 1: Ironically, the No Original Research [3] rule came about to counter claims of unreliability. I doagree with you that -while it does set a lower limit on reliability- it also sets an upper limit. Besides your objections, it also removes a channel for establishing priority [1], and makes it harder for experts to contribute from memory. [2].
note 2: Wikipedia actually incorporates the 11th edition of Encyclopedia Britannica as the basis for a number of articles [4]
[1] https://en.wikipedia.org/wiki/Wikipedia:Lost_functionalities...
[2] https://en.wikipedia.org/wiki/Wikipedia:Lost_functionalities...
[3] https://en.wikipedia.org/wiki/Wikipedia:No_original_research
[4] https://en.wikipedia.org/wiki/Encyclop%C3%A6dia_Britannica_E...
Of course edit may not be possible in contentious subjects but most of us know that.
This seems like a misunderstanding of what the AI is doing. It does not "know about the world" or assert knowledge of any kind of truth, other than to check a citation given in a wikipedia article, and attempt to verify the reference actually contains said information.
For example, if I edited a wiki article to state "over 12 million people visit Hacker News every hour" and linked a random HN article claiming nothing of the sort, this AI would attempt to parse my citation and if it was successful determine the reference didn't support the claim.
But the title "factcheck wikipedia" is not what the model does. At best the model would make bad actors game references the same way people game page rank on google. Sure, deterrent, meh.
But no way I’d trust Facebook on this.
Any particular examples?
If you look at articles on IQ, the entire page claims any link between IQ and race, which has been repeatedly proven in countless independent studies over decades, is wrong and that all IQ differences are environment, a lie that seems blatantly political.
These pages are locked to anyone except the usual few Wikipedia elites with an agenda.
Conservapedia, while not my choice as a source of truth, has a long list of what they find wrong with Wikipedia https://www.conservapedia.com/Examples_of_Bias_in_Wikipedia
RationalWiki also has a section dedicated to this phenomenon, including multiple points of view: https://rationalwiki.org/wiki/Wikipedia#Bias_in_Wikipedia
There are two freely accessible online philosophy encyclopaedias written by actual philosophy academics - Stanford Encyclopedia of Philosophy and Internet Encyclopedia of Philosophy (IEP) - both are miles better than Wikipedia. (IEP tends to be more accessible for beginners - some of Stanford’s articles can get quite esoteric and technical.) Another very good philosophy encyclopedia is Routledge’s, albeit it is not freely available online (or at least, not legally). Why rely on Wikipedia when there are much more reliable and higher quality alternatives?
>Automated tools can help identify gibberish or statements that lack citations, but helping human editors determine whether a source actually backs up a claim is a much more complex task — one that requires an AI system’s depth of understanding and analysis.
>we’ve developed the first model capable of automatically scanning hundreds of thousands of citations at once to check whether they truly support the corresponding claims.
The road to hell is paved with good intentions.
Fact-check is already a sullied overloaded term. It would be better replaced and served by something like “citation-check”.
There are a panoply of cited sources (“facts”) that needs to be properly vetted (“aligned”) to contribute toward its premise (“fact”).
This is why Wikipedia can often lay claim to being more scientific (through sheer column of citations containing of “facts”) assembled by its editors (citation scientists). This is also why many educators teach their students not to cite “Wikipedia” which is (a poor attempt?) to indoctrinate the students into learning how to root out the misleading source (often mistaken as “fact”.)
Mmmm, but it’s SCIENCE! Doesn’t necessarily means it’s a fact.
Meta (“Facebook”) would be venturesome to claim the science of citiogenesis as there are money, prestige, and power to be gained through shaping “science” of these citations. That is, by using artificial intelligence (AI).
Arguably, today’s Fact-checkers would try to use a science process that rarely achieves its “factual” (but really called a premise) claim … in a clean and unarbitrary manner while free of bias: always with unnecessary fillers with a goal to sway the readers with cemented anchors to keep it away from their basic but unwanted “fact”. We call those “fact-checkers” an opinionated citation checkers; they save the readers from doing the work by its artificial power of singular analysis through the curation of its premise (“fact”).
Fact-checkers are basically wannabe- citiogenesis scientists that are just merely interpreting their point of views. And readers (students) who failed their educators’ lesson would claim it as “fact”.
AI may or may not help and into both directions toward and away from their desired premise.
How many different algorithms would be used to dislodge this badly-abused citing of these singular analysis efforts by seemingly “fact-checkers”?
Who would be in control of this Machine Learning of AI? Meta (Facebook)!
Who oversees these AI algorithms? Meta!
And who would be the one that watches the watchers? Meta?
And will today’s educators teach these future generation of discerning readers on the much needed distinguishing between citation checkersd vs. today’s “fact-checkers”? (*cricket*)
Put another way, not all topics’ depth need be created equal, and arguably, Wikipedia is upside down.
The only reason for doing something like this is to ultimately subvert the Wikimedia editors, setting up Factbook/Meta as the sole arbiter of what's correct and true on Wikipedia.
https://tvtropes.org/pmwiki/pmwiki.php/Film/TheSocialNetwork
Reading through it, I strongly disagree with FB's example for "Better citations in action". I don't see an improvement in the wording and IMO they would be making it worse by switching from an official first party source to a third party one.
While their tool might be useful to find semantic (mis)matches, a much more important part of verifying citations is to verify that the source has any business to make claims about the matter in the first place. https://xkcd.com/978/
QUOTE Better citations in action Usually, to develop models like this, the input might be just a sentence or two. We trained our models with complicated statements from Wikipedia, accompanied by full websites that may or may not support the claims. As a result, our models have achieved a leap in performance in terms of detecting the accuracy of citations. For example, our system found a better source for a citation in the Wikipedia article “2017 in Classical Music.” The claim reads:
“The Los Angeles Philharmonic announces the appointment of Simon Woods as its next president and chief executive officer, effective 22 January 2018.”
The current Wikipedia footnote for this statement links to a press release from the Dallas Symphony Association announcing the appointment of its new president and CEO, also effective January 22, 2018. Despite their similarities, our evidence-ranking model deduced that the press release was not relevant to the claim. Our AI indices suggested another possible source, a blog post on the website Violinist.com, which notes,
“On Thursday Los Angeles Philharmonic announced the appointment of Simon Woods as its new Chief Executive Director, effective Jan. 22, 2018.”
The evidence-ranking model then correctly concluded that this was more relevant than Wikipedia’s existing citation for the claim. /QUOTE
I think you misread something. The article that was originally cited is not relevant. It was about someone else becoming CEO for a different organization.
It's more accurate to say that this AI is fact-checking citations. There are lots of ways you can skew or fabircate information on Wikipedia. One well known way is to create some source for a claim and then cite it on Wikipedia. What will an AI do in this case?
3-4 years ago I stood in Menlo Park when Mark Zuckerberg got up and announced in response to the misinformation issues of the 2016 election that an effort would be made to fact check articles. My immediate thought, which hasn't changed, is "that's never going to work". You will always find edge cases where reasonable people would disagree but that's not even the big problem.
The big problem is that there are lots of people who aren't the slightest bit interested in the "truth". I've heard it say that whatever ridiculous claim you want to fabricate, you can find 30% of Americans who will believe it. As soon as you start trying to label content as truthful or not, you won't change the minds of most people. For many you will be contradicting their world view and you'll simply be dismissed as "biases" or "fake news".
I honestly don't know what the solution to this is. I do think sharing links on Facebook itself was probably a mistake for many reasons.
So this effort to fact check citations just seems more of the same doomed policy.
Disclaimer: Ex-Facebooker.
[1]: https://www.theonion.com/wikipedia-celebrates-750-years-of-a...
“Read not to contradict and confute; nor to believe and take for granted; nor to find talk and discourse; but to weigh and consider. Some books are to be tasted, others to be swallowed, and some few to be chewed and digested: that is, some books are to be read only in parts, others to be read, but not curiously, and some few to be read wholly, and with diligence and attention.”
"The humanity's biggest problem with quotes you find on the Internet is that people tend to immediately believe in their authenticity."
Who fact-checks the fact-checkers?
Which authoritative AI determines the authoritativeness of the other AI's?
?
Or, perhaps phrased in "American Dad" terms... "Who's manning the Internet?" <g>
(And, related questions, like "Who trolls the trolls?", "Who thought polices the thought police?", "Who propagandizes the propagandists?", "Who spams the spammers?", etc., etc.! <g>)
What Meta is gonna try here is to use AI to overwhelm reality based on Human Resources to meet whatever Mr Zuck believes reality is.
IOW they’re gonna DDOS Wikipedia to show Mark’s reality.
Anyone who thinks that’s not gonna happen (whether it happens intentionally or not) is just fooling themselves.
https://en.wikipedia.org/wiki/Wikipedia:Bot_policy
For edits that require judgement, a common mode is to use semi-automated tools where there is still a Human-In-The-Loop.
(edit) And if you look at their demo, the latter is exactly how it works. https://verifier.sideeditor.com/ . With a human in the loop like this, this could be a powerful tool.
I don't think I will trust wiki ever again if they do this. It is already like an encylopedia made in bar bathroom graphiti. Now they want to turn it into 4chan.
Are you familiar with Encyclopedia Dramatica? It's basically Wiki for Internet Culture, if all of Wiki's editors where autistic drama merchants from 4chan. It's glorious: https://encyclopediadramatica.online/Category:Weeaboos
That site was immortalized by Cracked: https://www.youtube.com/watch?v=VgQMTLKmwrA
"That's a blumpkin yo."
...including this comment! <g>
It’s taking subtle human judgement out of decision making.
Some systems don’t need subtle social decision making. Using AI for those systems is great. But a lot do. And in those that do, if you’re not careful and automate too much, errors compound. Quickly.
AI recommendation engines have already ripped society apart and increased division, because the subtle bridge building that human curation enables is removed.
Hospital systems have been being increasingly systematized and are becoming increasingly hellish and expensive. They’re a non AI example of the same phenomenon. The risks are in trying to systematizes things with a social aspect.
I love automation and think we should strive to free up as much time for exploration, creativity, and human connection as we can. But AI is just plain bad at understanding us, and always will be, in my view. It will always be a mimic. It simply cannot do the subtle social things humans do because it does not have the motivations that enable that subtle behavior.
The comment above is nonsense; dangerous at worst, idiotic at best. There's no comparison between AI and writing.
The problem is so common, I really think there should be a term.
Edit the article? I don't expect it to be accurate enough to avoid getting quickly banned or auto-reverted.
Raise some issue for volunteers to manually review? Probably not accurate or important enough to be a priority given the likely volume.
Honestly a good portion of the sources I check out in a wikipedia article just 404 or have become paywalled, and that would be pretty trivial for a bot to detect, so there's obviously not a huge desire to have bots checking sources in the first place.
There is however, an open-source AI, and a github link to it. If you're familiar with Python you can literally read the code yourself and look for anything nefarious.
> Thou hypocrite, first cast out the beam out of thine own eye; and then shalt thou see clearly to cast out the mote out of thy brother's eye. - Matthew 7:5
(Sorry, I grew up on KJV, it what I remember and what sounds right in my head)