How I trained fake news detection AI with 95% accuracy, and almost went crazy
towardsdatascience.com
towardsdatascience.com
Where’s the dataset? How did you verify the ground truth? Where are the annotation/labeling guidelines?
What’s the definition of factual/real articles? The dataset appears to be created by the author - which isn’t necessarily wrong but to paraphrase Karl Popper (in the context of human knowledge and scientific endeavors):
There are no ‘pure’ facts available; all observations are functions of subjective factors such as interests, expectations, wishes etc.
Constructively criticism to the OP: I'd suggest they read the nuance and discussions on the Fake News Challenge [0] and then look into their datasets + evaluation code [1] instead of hand-coding their own "biases" into a {"Fake news","Not-Fake-News"} binary classifier. Feel free to replace "Fake News Challenge" with any other similar effort so that OP isn't tasking themselves with the massive task of "Solving Fake News" all alone.
Disclaimer: I don't have any stake in FNC-1
References:
A more accurate way of detecting "fake news" would be interesting, but I fail to see how such a thing could be designed, past simple detection of wishy-washy and avoidant word patterns.
https://en.m.wikipedia.org/wiki/Evaluation_of_binary_classif...
This is one of the reasons why I recommend that Medium thought pieces disclose their data and code instead of just saying "I did AI magic!" to sell a product (and they do charge for their product on their website).
A good question and I'm not surprised he went a bit crazy.
https://plato.stanford.edu/entries/truth/
> The problem of truth is in a way easy to state: what truths are, and what (if anything) makes them true. But this simple statement masks a great deal of controversy. Whether there is a metaphysical problem of truth at all, and if there is, what kind of theory might address it, are all standing issues in the theory of truth. We will see a number of distinct ways of answering these questions.
What is "wrong" in news is "an implausible narrative relative to the reader."
Plato deals with truth which is something entirely different, not in the same realm of the current state of news.
To be more specific, news can be one of four things to a reader: plausible and true, plausible and untrue, implausible and untrue, and implausible and true.
Most readers seem to more concerned with a stories' plausibility and not it's trueness.
Most importantly though is readers no longer value the "truth" component of news. They value whether or not the narrative aligns with their own view of the world.
Because what truly happened doesn't matter to most people, only that they have a way to make sense of it themselves. Even if it's a partially or completely false narrative.
The author describes a "fake news detector AI", that is actually a "typically legitimate source of news" data model, combined with a fake news domain blacklist. It doesn't detect fake news. It detects whether a story possibly came from a source you find to typically be legitimate.
This article is fake news.
A "true fake news detector" would use only the text of the article -- without the URL. And so I agree that this article is kind of like fake news of the "misleading" variety. :)
False fakes--and a wisely designed system would probably say "probably fake" in the first place, not outright "fake"--are less damaging than false truths overall.
That depends on what your goal is: to discover groundbreaking truths (which are almost always considered taboo/false at first) in order to improve society, or merely attempt to cover up apparent falsehoods as damage control.
In the latter case, such a system would be highly prone to enforcing the prexisting biases and prejudices of its developers onto the users, making it much more difficult for them to discover truths that the developers are ignorant of.
From the article, it does seem that the training and test sets were both generating from a dubious data collection process: scraping URLs they knew were fake/satire/real and then applying a label based upon the domain.
But, they had more features than just "source" in their model. so while it's possible their data collection method means that their model is way over-fit and basically just tests of proxies of "source", it's not prime facie obvious that this is the case...? Or am I missing something?
> Domain name - Some domains are known for hosting certain types of content, Fakebox knows about the most popular sites
Fake news sites typically register bespoke domains, though.
If the test data was collected in the same way that the training data was collected, this was a rather ridiculously round-about way of writing what should've been a 5 line perl script, and I'm really curious where the 5% error could've possibly even come from :-\
Then, in your account page, you scroll down and there are instructions for installing/running fakebox (and others).
The "demo" is actually a docker app that spins up a web application of some sort. I'm firing it up as I type this, and I'll checking it out soon-ish.