Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?"
Either way, disappointing.
Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?"
Either way, disappointing.
That is indeed an area where LLMs don't shine.
That is, not only are they trained to always respond with an answer, they have no ability to accurately tell how confident they are in that answer. So you can't just filter out low confidence answers.
I’m presuming that one class of junk/low quality output is when the model doesn’t have high probability next tokens and works with whatever poor options it has.
Maybe low probability tokens that cross some threshold could have a visual treatment to give feedback the same way word processors give feedback in a spelling or grammatical error.
But maybe I’m making a mistake thinking that token probability is related to the accuracy of output?
Isn't that what logprobs is?
> Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?"
That's literally what ChatGPT did for me[0], which is consistent from what they shared at the last keynote (quick-low reasoning answer per default first, with reasoning/search only if explicitly prompted or as a follow-up). It did miss one match tough, as it somehow didn't parse the `<search>` element from the MDN docs.
[0]: https://chatgpt.com/share/68cffb5c-fd14-8005-b175-ab77d1bf58...