When and with what subjects it makes shit up is also heavily dependent on training data, and the result is straight up a black box. What good is a fact generator that can't be trusted?
When and with what subjects it makes shit up is also heavily dependent on training data, and the result is straight up a black box. What good is a fact generator that can't be trusted?
If I'm openAI or Google or whatever, I'm definitely going to run extra classifiers on top of the output of the LLM to determine & improve accuracy of results.
You can layer on all kinds of interesting models to make a thing that's generally useful & also truthful.
The nice thing about the easy to bamboozle GPT4 is that it can’t hurt anything, so its flaws are safe. Giving it these arms and legs is where the risks increase, even as the reward increases.
If you ask Wolfram Alpha - something which I think is actually meant to be a fact generator - "Which is the heaviest Pokemon?" it will happily tell you that it is Celesteela, and it weighs 2204.4lbs.
Is that a 'fact'?
It certainly 'true', for some definition of the word true. The game Pokemon exists, and in it Pokemon have a weight. Of all the official Pokemon, that one is the heaviest. Wolfram Alpha has given you an accurate answer to your question.
But it's also completely made up. There's no such thing as a Pokemon, and they do not actually have weights.
So sure, transformer models can't be relied upon to generate facts. But so what? There's a lot more to the world than mere facts.