I’m blown away by the competence of the language model but its willingness to make up facts makes me leary.
I’m blown away by the competence of the language model but its willingness to make up facts makes me leary.
“everything you read in the newspapers is absolutely true, except for the rare story of which you happen to have firsthand knowledge”.
Humans are pretty good bullshitters too!
I've asked BingGPT about myself and it gave me three answers. One was more or less on-point (it found my linkedin profile), and the other two were hallucinations. What happened was Bing found two unrelated pages and GPT has tried and failed to make sense of them.
Either that, or I am a prince whose name means "goose" in Polish.
I really hope they find a way to have it apply context from future conversations such that when it learns the error of its ways it emails you a retraction, but that's probably a ways out because humans can't be trusted to not weaponize such a feature into sending spam.
The weight of phrases like "you are wrong" is in fact so strong, that it fools the chatGPT to apologize for its 'mistakes' even in the scenarios where its text was obviously correct - like telling it 2+2 doesn't equal 4
Sure, grep has never flat out lied to me the way chatGPT does, but it's a statistical model, not a co-worker, so I don't feel betrayed, I just feel... cautioned. It keeps you on your toes, which isn't such a bad state to be in.
I see you haven't met humans
Bing, IIRC, has a way to provide feedback, not sure how useful it is for today's users and if it will be able to solve hallucinations one day.
But I'm generally not interested in asking an average human. I'm interested in asking someone who knows their butt from a hole in the ground in whichever topic I'm asking them about.
When ChatGPT tells me something, I have no idea if it's paraphrasing information gathered from Encyclopedia Britannica, or from a hollow-earther forum.
phind.com has been incredibly good for me.
Is the work of judging the accuracy of a summary not just the work of comprehending the non-summarized field?
For example, a summary could be completely correct and cite its facts exhaustively. Say you're asking about available operating systems: it tells you a bunch of true info about Windows and OSX, but doesn't mention the existence of Linux. Without familiarity with the territory, wouldn't verifying the factuality of each reference still leave you with an incomplete picture?
At a slightly more practical level, do you actually save any time if you've gotta fully verify the sources? I assume you're doing more than just making sure the link doesn't 404, as citing a link that doesn't say what it is made out to be isn't exactly a new problem, but at that point we're mighty close to the traditional experience of running through a SERP.
Finally, even if you're reading all the links in detail, isn't that still a situation prone to automation bias? There's a lot of examples of cases where humans are supposed to check machine output, but if it's usually good enough the checkers fall into a pattern of trusting it too much and skipping work. Maybe I'm just lazy, but I think I'd eventually get less gung-ho about verifying sources and eventually do myself a mischief.
I'm asking because I've been underwhelmed by my own attempts at using LMs for search tasks, so maybe I'm doing it wrong.
Or it's something it just hallucinated out of thin air.
Especially when I talk to it about fiction and ask questions about - for example - a specific story and you see it invent whole quotes and characters and so on...it is a masterful bullshitter.
Checking the links is a good practice.
I feel like we just created an interesting novel problem in the world. Looking forward to seeing how this plays out.
Leary is a rare variant spelling of leery.
I mention this, because you seem to care about correctness.
Cold War strategy. Trust but Verify.
What could be more 2020s than a post-truth search engine?