755 karma · joined May 6, 2013
Yup. Pretty sure I've built that system before. Can relate.
Fair and accurate. In the best cases the person running the model didn't write this stuff and word salad doesn't communicate whatever they meant to say. In many cases though, content is simply pumped out for SEO with no intention of being valuable to anyone.
Lol
In case you want other mental health tips from the study: "Individuals who abstained from alcohol showed lower cognitive scores than those who consumed alcohol"
At indexing time:
- run LLM over every data point multiple times ("gleanings") for entity extraction and constructing a graph index
- run an LLM over the graph multiple times to create clusters ("communities")
At query time:
- Run the LLM across all clusters, creating an answer from each and score them
- Run the LLM across all but the lowest scoring answers to produce a "global answer"
...aren't the compute requirements here untenable for any decent sized dataset?
"Since 2014, it has raised its menu prices by 100%"
... during the same time period, US inflation has been 32%.
- Using a RAG architecture on top of a database of factual information. Wikipedia is probably your best bet. It is not 100% factual or correct either, but maybe as good as it gets. Scaling RAG to wikipedia size is not trivial, but I think it can be done.
- Prompting the LLM to cite its sources so people can fact-check the fact-checker
- Prompting the LLM to say it is unsure when something does not have a clear answer. I don't expect this to be reliable, but maybe somewhat better
- most current LLMs are trained on large amounts of web data that itself contains facts, opinions, and misinformation. These things are treated equally, so I would expect the LLM to get common facts right, but also to represent opinions or misinformation as facts when they are pervasive.
- LLMs "hallucinate" and tend not to know when to say "I don't know" or to not try to fact-check something that is not factual in nature.
...in short, I would expect LLMs to be an unreliable fact checker, which has the potential to do as much harm as good.
Besides making ridiculous and unsupported causal claims (as other commenters have noted), the results are really poor (R score of 0.2). They show that the model achieves human-level performance, which is also really poor. Then they run it on natural images of politicians and the results are significantly worse. What the paper actually shows is that facial appearance can not consistently predict political alignment, but the authors want to try really hard to do it anyway.