Loki: An open-source tool for fact verification
github.com
github.com
* The idea behind using Serper is great, however it would be cool if other search engines/data sources can be used instead, ie. Kagi or some private search engine/data. Reason for the latter: there are tons of people who are sourcing all sorts of information which will not immediately show up on google and some might never do. For context: I have roughly 60GB (and growing) of cleaned news article with where I got them from and with a good amount of pre-processing done on the fly(I collect those all the time).
* Relying heavily on OpenAI. Yes, OpenAI is great but there's always the thing at the back of our minds that is "where are all those queries going and do we trust that shit won't hit the fan some day". It would be nice to have the ability to use a local LLM, given how many and how good there are around.
* The installation can be improved massively: setuptools + entry_points + console_scripts to avoid all the hassle behind having to manage dependencies, where your scripts are located and all that. The cp factcheck/config/secret_dict.template factcheck/config/secret_dict.py is a bit.... Uuuugh... pydantic[dotenv] + .env? That would also make the containerizing the application so much easier.
Regarding the first version, we are currently working on enabling customized evidence retrieval, including local files. Our plan is to integrate existing tools like LlamaIndex. Any suggestion is greatly appreciated!
Regarding the second point, we have found OpenAI's JSON mode to be greatly helpful, and have optimized our prompts to fully utilize these advances. However, we agree that it would be beneficial to enable the use of other models. As promised, we will add this feature soon.
Lastly, we appreciate your suggestion and will work on improving the installation process for the next version.
I'm a native English speaker, and traditionally when it comes to formal/professional written English (emails etc.) my instincts take me to sounding quite GPTish - luckily I've got a good grasp of the language and have found it fairly easy to alter my formal writing style to be a bit less traditional and a bit less formal too, but if it wasn't my first language and I wasn't a fair bit above average in writing ability even for native speakers, I suspect it wouldn't be nearly as easy to go against how I was taught at school to write in formal situations.
It's really not enough to see that somebody writes roughly in that style to assume they're using LLMs, because the reason LLMs so often sound like that is because they've learned from humans very often sounding like that.
In an example such as this particular case, it maybe set off your LLM suspicions because culturally you wouldn't expect somebody to sound so formal in comments on a site like HN, and choosing the wrong tone of voice for the context is something an LLM is likely to do - but actually, if a) English isn't your first language nor part of your primary culture, and b) you're wanting to make a good impression as the subject of the thread is something you've created and are therefore essentially acting as a spokesperson for in the comments, then all of a sudden writing formally rather than as if writing throwaway forum comments makes sense rather than looking like an indication that AI wrote it.
https://arxiv.org/abs/2305.14902 https://arxiv.org/abs/2402.11175
Who could better know the patterns of liars than the god of lying.
[1] https://github.com/Libr-AI/OpenFactVerification/blob/main/fa...
[2] https://github.com/yuxiaw/Factcheck-GPT/blob/main/src/utils/...
https://arxiv.org/abs/2403.18802
https://github.com/google-deepmind/long-form-factuality/tree...
Compared with our initial version, we have mainly focused on its efficiency, with a 10X faster checking process without decreasing accuracy.
A 2020 Meta paper [1] mentions FEVER [2], which was published in 2018.
[1] "Language models as fact checkers?" (2020) https://scholar.google.com/scholar?cites=3466959631133385664
[2] https://paperswithcode.com/dataset/fever
I've collected various ideas for publishing premises as linked data; "#StructuredPremises" "#nbmeta" https://www.google.com/search?q=%22structuredpremises%22
From "GenAI and erroneous medical references" https://news.ycombinator.com/item?id=39497333 :
>> Additional layers of these 'LLMs' could read the responses and determine whether their premises are valid and their logic is sound as necessary to support the presented conclusion(s), and then just suggest a different citation URL for the preceding text
> [...] "Find tests for this code"
> "Find citations for this bias"
From https://news.ycombinator.com/item?id=38353285 :
> "LLMs cannot find reasoning errors, but can correct them" https://news.ycombinator.com/item?id=38353285
> "Misalignment and [...]"
Sorry, I think an individual who is not only aware of reliable sources to verify information, and who is not familiar enough with LLMs to come up with appropriate prompts and judge output should be the last person presenting themselves as the judger of factual information.
We present the results at each step to help users understand the decision process, which can be seen from our screenshot at https://raw.githubusercontent.com/Libr-AI/OpenFactVerificati...
We will try our best to ensure this tool makes a positive difference
I imagine somebody feeding a live presidential debate into this. Could be a great tool for fact checking
Does it...work?
However, errors can always occur. We try to help users in an interpretable and transparent way by showing all retrieved evidence and the rationale behind each assessment. We hope this could at least help people when dealing with such problems.
While it answered a general "yes" when the more precise answer was "no", the motivation in the answer was perfectly on point and exactly the same things.
As a general LLM for regular user fastGPT (their llm service) is in my opinion "meh" (lacks conversations for instance). But it's really impressive that it contains VERY recent data (like news and articles from last few days) and always provides great references.
Intuitively, just because you put your LLM into a workflow/pipeline this doesn’t really address how to eliminate hallucinations.
For those of us that don’t follow the research closely, can you explain how your findings and approach allows you to utilise LLMS and work around this hard limitation. Said another way, how are you getting around the fact that LLMs themselves regularly output lies/false answers?
The idea that "specificity," such as what scientific research aims for, can be better evaluated for truthfulness or approach what "truly matters," as this project purports, is dubious. E.g., why would a notion that is more limited in scope matter more than something more vast (to use the word that it cites as an example)? In addition to its dystopian idea of a "source of truth," it completely dismisses "vague" language in the name of "science" or "factuality," which is utterly the opposite of science, which I thought was to understand ourselves and nature with as few presuppositions as possible.
I am not commenting on the actual software and I know names are hard and often overlap, but with something as popular as Loki already used for logging I think it might get confusing.
When we named our project, we were unaware of the overlap with Grafana Loki. We appreciate you bringing this to our attention! I will discuss this issue with my team in the next meeting, and figure out if there is a better way of solving this. If you have any suggestions or thoughts on how we can better differentiate our project, we would love to hear them.
Thank you again for your valuable input!
The best case scenario would seem to be that results are derived from certain biases built into the model, unless it weighs “factuality” by the number of occurrences of certain statements on the internet which is as far from a qualification for truthfulness as the biased model.
[0] - https://open.substack.com/pub/thetechenabler/p/trust-in-brea...
We have it everywhere. The problem is however well-known: Human bias, political engagement from the fact checkers, etc.. AI (without any kind of lock, political bias built-in etc) could be the real deal, but because it may be not political correct, it will never happen.
I think the only proper way to verify facts is to derive them from "fundamental facts". E.g., that the earth is round (and even for that there are ppl believing the opposite).
This is some giant BS that is for sure. Some stupid, literally brain-dead AI searching things created by humans to determine what is a "fact". This is beyond dystopian crap.
We all know all the fact-checker orgs. used by big tech like Facebook and others are filled with hyper biased woke people who do not actually fact-check things but get off on having the power to enforce their beliefs, feelings and biases.
I can already tell this is total BS without even looking into it, what kinds of sources will it use? What ranking will they give them? Snopes? ROFL. Probably just uses some woke infested, censored and curated language model to determine a fact based on what has the most matches or THE MOST LIKELY because that how AI works. Has absolutely nothing to do with facts.
And it's even worse, we are literally in a time when AI hallucinates things that do not exist. I won't use a stupid AI to find me "facts".