GPT-4o is consistently hallucinating
platform.openai.com
platform.openai.com
1) it's insanely chatty, to a point where it ignores instructions about not doing certain things. I think this behavior is heavily favoured by benchmarks but as somehow who expects concise answers, this model annoys me. Custom instructions don't fully fix this for me.
2) It likes repetitive answers a lot more than the previous version. Meaning that it will try its hardest to generate the followup answer in the same format as the first one. I think this is also the problem in your example.
To my understanding, this is a measure against laziness, where the model would exclude information from the first answer that haven't changed in the followup. I always liked this behavior but maybe you remember the time from a few months ago where many people complained about the laziness of (I believe) 0125.
Btw, while I type this, I notice that this is probably the highest level of first world problems I've ever complained about. There is this amazing almost free tool that answers all my questions and does most of my coding and I dislike it because it provides me with thorough context.
Based on a tweet in X[1], I had to add "I REPEAT" to the instruction to get the model to not to ignore instructions.
No thanks.
Not like an LLM hallucinating is particularly surprising or newsworthy anyway.
Author, do some more robust tests, write a blogpost about it, and submit that instead.
In the middle of a conversation, try going "SHUT UP, STOP TALKING ALREADY". For me, it just keeps repeating the last output. Very cool.
I routinely fix this by toggling from GPT-4o to GPT-4t.
Having to constantly correct incorrect answers by LLMs only for them to apologize and give another incorrect answer is what made me lose complete interest in using them.
I figured if I'm knowledgable enough to correct LLMs it's more efficient to not use them at all. What's the point really? Am I teaching them? Because I felt like a teacher who is quizzing a student who keeps on guessing but failing.
For example, to inform the University of California about the content of my courses, I have to go through a course articulation which is several pages long, is written in a formal academic voice, and is pretty time consuming to create. GPT-4t can take my informal course outline and an example of a past articulation that I've written and do the job to a point where I just need to ask it to make small changes for 10 minutes and then make a last couple edits myself. I turn a couple of hours to 10 minutes and 25 cents of API calls.
(Also, sometimes when it's explaining example assignments, it thinks of nice things to include that I hadn't planned on, and I end up shamelessly using them; other times it thinks of garbage and I have to coax it to articulate what I actually meant).
I'd say GPT-4o is slightly better at the task... except it commits so strongly to its answers in the context buffer that it doesn't do effective rewrites/corrections. So I've settled into a workflow of using GPT-4o to do initial work and then use GPT-4t for the final cleanup.
The models are non-deterministic and there is no way to predict hallucinations.
And there is a way to predict the presence of hallucinations, but it’s expensive (SelfCheckGPT)
It is at best 90% accurate.