> How would an AI know the truth?
There's been some research on this. I think I even saw at least one article about it here on HN, though I can't find a link any more.
What I remember seeing is a kind of world (?) map of the Internet with a set of different-colored interconnecting webs of sites.
There is one set that deals mostly in verifiable facts, usually based on genuine, usually academic research. Universities are major nodes in this net, but you get some overlap with "mainstream" news sources, Wikipedia and so on.
There's another major set that includes InfoWars, Collective Evolution, Natural News, Deepak Chopra, Mercola and other Anti-Vaxxers, various AGW deniers, and other well-known sources of bullshit.
Both of those networks have many more connections within themselves than to the other clouds, and have their own frequent-viewer audiences who will frequent the one set more than the other.
There were, as I remember, other sets that were less significant.
Purely on probability, a story will be dramatically more likely to be true if originated and/or propagated in the first of those clouds than the second.
Mapping these webs is a pretty mundane mechanical task, yet it can provide an excellent filter for some classes of information.
Other stuff is "squishier" because it's more opinion than fact. Politics and policies will be found here, and you're likely to see (I don't remember if the article I have in mind showed this) "conservative" vs. "liberal" networks rather than "fact" vs. "bullshit". On the other hand, there are experts on many aspects of politics, and if you're looking for ideas on policy that are informed by history and research, you're more likely to get them from professors of history, economics and law than from Fox News or televangelist channels.
I suspect there are a lot more inferences that can be drawn, via probability or Bayes, from the Internet sources of information, its propagation behavior, and the shape of the networks it travels.
Similarly, there's a wealth of information to be gained from analyzing the language of an individual page. Emotional words vs. dispassionate ones; weasel words like "probably" or "it was said" as opposed to "according to" or links to other sources. Bullshit is often simplified for consumption by simple people, while academics often carelessly lapse into "high-bred" or technical language.
Some of those features could be gamed, simulated, faked, and certainly will be. Many bullshit articles in the health, field, for instance, now link to "outlier" or discredited papers in PubMed to enhance their credibility. But like our grandparent poster, I think AI, working patiently and dispassionately, will be able to continue to find hints to distinguish the wheat from the chaff.