The 80085 case is only interesting insofar as it reveals weaknesses in the tool, but it's so far from tool-use that it doesn't seem very relevant.
Agreed on the meta-point that deliberate tool mis-use, while amusing and sometimes concerning, isn't determinative of the fate of the technology.
But the failure rate without tool mis-use seems quite high anecdotally, which also comports with our understanding of LLMs: hallucinations are quite common once you stray even slightly outside of things that are heavily present in the training data. Height of the Eiffel Tower? High accuracy in recall. Is this arbitrary restaurant in Barcelona any good? Very low accuracy.
The question is how much of the useful search traffic is like the latter vs. the former. My suspicion is "a lot".
The problem with your judgement is you click on every “haw haw, ChatGPT dumb” and you don’t read any of the articles that show how an LLM works, what is is quantitatively good at and bad at and how to improve performance on tasks using other methods such as PAL, Toolformer or other analytic augmentation methods.
Go read some objective studies and you won’t be yet another servomechanism blindly spreading incorrect assumptions based on anecdotes from attention starved bloggers.
Wanna try again? Alternatively you can keep riding the hype train from techfluencers who keep promising the moon but failing to deliver, just like they did for crypto.
More specifically, without language, can you know that someone else knows anything?
But speaking the truth is just minor and rare application of the language.
> More specifically, without language, can you know that someone else knows anything?
Honestly, just ask them to show you math. If they don't have any math they probably don't have any true knowledge. The only other form of knowledge is a citation.
Language and truth are orthogonal.
So as a product, that’s the game it’s playing and failing at. It’s unhelpfully pedantic to try and steer into technicalities.
If that is the measure you are using that's cool, but
>So as a product, that’s the game it’s playing and failing at.
It is failing that measure by such a wide margin that if "everyone" (certainly anyone at MS) was using that measure then the product wouldn't exist. The measure MS seems to be using is it entertaining and does it get people to visit the site. Heck this is probably the most I have heard about bing in at least 5 years.