Fake News Detectors Are Biased Against Texts Generated by Large Language Models
arxiv.org
arxiv.org
I think this is more interesting:
> [LLMs] are more prone to flagging LLM-generated content as fake news while often misclassifying human-written fake news as genuine.
> To address this, we introduce a mitigation strategy that leverages adversarial training with LLM-paraphrased genuine news. The resulting model yielded marked improvements in detection accuracy for both human and LLM-generated news.
I didn't see any concrete identification in the paper itself, but following one of the references[1] I found this:
> Within our list of news sites, we differentiate between “unreliable news websites” and “reliable news websites.” Our list of unreliable news websites includes 1,142 domains labeled as conspiracy/pseudoscience” by mediabiasfactcheck.com as well as those labeled as “unreliable news”, misinformation, or disinformation by prior work (Hanley, Kumar, and Durumeric 2023; Barret Golding 2022; Szpakowski 2020).
> Our set of “unreliable” or misinformation news websites includes websites like realjewnews.com, davidduke.com, thegatewaypundit.com, and breitbart.com. We note that despite being labeled unreliable every article from each of these websites is not necessarily misinformation.
> Our set of “reliable” news websites consists of the news websites that were labeled as belonging to the “center”, “center-left”, or “center-right” by Media Bias Fact Check as well as websites labeled as “reliable” or “mainstream by other works (Hanley, Kumar, and Durumeric 2023; Barret Golding 2022; Szpakowski 2020). This set of “reliable news websites” includes websites like washingtonpost.com, reuters.com, apnews.com, cnn.com, and foxnews.com
The methodological issues here, and the trajectory following from ignoring them wholesale, are considerably more interesting than the study itself, to me.
1: "From January 1, 2022, to April 1, 2023, there was a dramatic surge in synthetic articles, especially on misinformation news websites (Hanley and Du- rumeric, 2023)." which links to https://arxiv.org/pdf/2305.09820.pdf
I read on HN thal LLMs know literaly everything (electronics, computer science, physics). /s
Humans can't fly, but can build machines that do.
And if humans could build some other type of truth machine, how would they know it was working if they themselves are not truth machines?
All those tests have to do is mimic system 2 thinking all the time, while we mere humans have to switch to system 1 often because 2 is slower.
As it should be.