SafeGPT: New tool to detect LLMs' hallucinations, biases and privacy issues
giskard.ai
giskard.ai
Along this line of thought: was it a massive oversight for them to not train the model to say "math detected, let me pass that to a solver" instead of trying to guess what token should come next in a math problem?
GPT is quite useful, but not because it solves the problem of "I don't know where the question I have is answerable by a calculator"
So personally I think the problem is that people see "this product does X" and interpret that to mean that it does X well. I don't think it's necessarily bad that we're seeing an explosion of AI tools that are a bit underwhelming if people understood it as such -- we're on, after all, a site with a heavy startup focus and saying "your product doesn't do everything that I want" is a bit antithetical to that.
But yeah specifically for this one there are arguments that "X is not even possible, especially not with this approach" so it's a bit more egregious.
(And no one cares that you used to work at Microsoft or whatever).
And for a large swats of things, how can it possibly work? It’s not possible to say if or if not it is hallucinating code for almost all code and apis, for instance. And I see similar issues with many fields outside pure facts. With privacy issues as well.
It would appear that this is not automated monitoring but more like a second stage of human reinforcement learning or perhaps a classifier. It seems that you create input/output examples and the LLM responses are examined by a secondary system (which I’m guessing is probably NOT an LLM, otherwise it would be vulnerable to attacks) and perhaps force regenerates the LLM response if it doesn’t meet the classification threshold.
At least, that sounds more believable to me than someone claiming they’ve fixed the inherent flaws in LLMs.
Currently, fact checking works on straight facts. It does a Google Search and uses LLMs to shorten it. Once it has the short version, it will compare the short results with the answer provided by ChatGPT itself. Premium tiers would get better fact checking sources than just google. We're investigating various data sources and comparison methods.
Note that fact checking / hallucinations is just one of the types of satety issues we'd like to tackle. Many of these are still open questions in the research community, so we're looking to build and develop the right methods for the rights problems. We also think it's super important to have independent third-party evaluations to make sure these models are safe.
This is a new tool we're building in the open, and we're interested in your feedback to prioritize!
Wow, you guys have a database of all the facts?
> It does a Google Search and uses LLMs to shorten it.
Oh...
...actually, this is an empirical fact checker. I wouldn't call it "fact-based", as it's epistemologically an absurd statement, but "empirical fact checking" sounds good and presents an idea that is very close to how humans verify information in the first place - by checking multiple sources and searching for correlation.
For what it's worth, I think your approach makes sense. Good luck.
So your fact-checking LLM is also vulnerable to injection and unethical prompting then when it ingests website text. And a Google search is far, far away from fact checking, particularly for the subtle errors that GPT-4 is prone to making.
Our society is actively declaring that falsehoods are truth, and should be celebrated. We're hallucinating ourselves. All this software does is make sure LLMs hallucinate with us.
Combined with ever-improving ways to fake video and voices, things could get even uglier than they've been over the last few years.
And yet there's something sinister about twisting an evolutionary model into a reeducation camp.
LLMs are not human, but strikingly similar. If we have no qualms about how we treat it, then what will we do to real people?
But truth isn't political. As long as we think it is, we will continue to follow the descent into madness. Truth is just that: truth.
The only reason we think truth is political, is because our chosen leaders so heavily depend on lies that the truth would destroy their reign.
Curious what you consider to be true though? I'm coming at it from the perspective that even in physics where we can isolate so nicely we still aren't divining any truths, just making models with increasing explanatory powers.
Personally, I've been reaching more towards 'shared values' than 'truth', this is likely the pedant in me but truth doesn't feel tractable whereas shared values feels like it has less baggage?
*pretty sure lies exist though
Are those true things? Good candidates, I like 'leptons exist'. Do you mind if we just gently ignore the math one? Feels like inviting the whole 'is math invented or discovered' thing.
1) carbon atoms in a mol - a mol is a counting number so it seems tautological to declare this one a truth
2) pass :)
3) this seems like a good candidate but it also seems to reduce truth to just the things we measure and only to the extent that we can be accurate (I'm also assuming you meant neutrons, protons are static by specie). Purely hypothetically there could be a whole heap of unusually heavy or light carbon out there that would disprove one or another of our theories. To put it another way; is the average number of apples that a trees grows in a year 'true'? It'll change year after year after all. I'm fine with a definition of truth that implies error bars and best efforts but I feel it falls short of the colloquial definition.
4) I think the pure observation that a thing somewhere exists is probably the closest to true, the rebuttals against that would all be self consuming anyway. The specific claim that leptons exist seems a little more fraught though - we could conceivably come to another conclusion if that better fit the facts.
So, can we call these things true if our concept is potentially incomplete or incorrect?
EDIT: I have avoided using "truth" here because it's a more general term than "fact" which has the connotation of being in reference to something concrete.
I don't claim to know the answers to all of the questions (and I certainly don't know where COVID came from), but clearly there are plenty of cases where dubious statements were strongly enshrined as "True" in a way that required major online players to suppress alternative beliefs as "False".
It is directly embedded in the ChatGPT website when you get the extension. As you ask questions, it will be added on the sidebar of each answer
TLDR: It lies on fact based information which is mentioned in very very few places on the internet and not repeated too much. Short of having a human with the context, how do you even detect it.
Example: Ask it to describe a "Will and Grace" episode with some guest appearance. It will always make up everything including the episode number and the plot, and the plot seems very believable. If you have not watched and can't find a summary online, it is hard to say that it is a lie.
The metamorphic testing approach is interesting and might work.
Take for example a credit scoring model: you can reasonably expect that if you increase the liquidity, the credit score should not decrease. In general it is relatively easy to come up with a set of assumptions on the effect of perturbation, which allows evaluating the robustness of a model without knowing the exact ground truth.
The utility of it is there and has nothing to do with making up episodes instead of quoting the current one. Like you can ask it to write new episodes with specific settings and specific constraints. Hallucination is not the value add. Nobody is excited because it hallucinates. People are excited despite it since the other value add is too much.