OpenAI just announced a new search tool. Its demo already got something wrong
theatlantic.com
theatlantic.com
I don't get it. I mean, how hard can it be? These are billion dollar ventures. For god's sake, at least have an intern fact check your press releases for 15 minutes before publishing them.
On the other hand, these repeated basic mistakes certainly help in keeping expectations in check. But I doubt that's the goal of the demo...
(Granted, your point about the desirability of basics, still holds ...)
It might as well be dogfooding, but to me that seems a bit too, hm, say, intellectually honest for what is basically a PR move in a hype market.
On one hand:
Marketing copy should be word-perfect, with no grammar errors.
On the other:
Video of baby's first steps should be raw + unedited.
The more you move towards a "real moment", the less sense it makes to polish it to a mirror sheen
It certainly can, but not always. In any case, this effect is why when I see a product that has a high degree of polish, I take it as a weak signal that the product is not great.
TBH, seeing those press releases lately, they seem to do it intentionally.
Treating your users as idiots has become the norm.
Yesterday my coworker was talking about using Gemini and wanted to show me how neat it was. So he typed a questing into Google. The search results came back instantly with the correct information as the #1 link. 5 seconds later, the Gemini box displayed the correct information. What the hell is so impressive about that?
This is one of the points missing in most conversations. Each has strengths and weaknesses, LLMs are not going to full-on replace search, but can often be a better starting point. The main advantage to LLMs is that you can either (1) provide vague questions or (2) provide a whole bunch of context
https://topicalsource.dev/chat/023e7e54-947b-490d-bcd8-89cc2...
This example shows, for a python error printed at the terminal, how you can get a really nice response by also sharing a code snippet and directory layout (extra context)
A demo showing an LLM helping to solve some contrived problem of this sort would resonate with people, and show a valid use case for LLMs in the search process. That's what would be impressive.
But hey look at this cool technical thing we can do!
It’s so cringey.
It's an understandable mistake given the dates included on that festival's page, but it's still a mistake. And the mistake is presented as a confident answer, which is the issue with these kinds of tools including Google's search result summary answers.
People need to learn to use LLMs for what they are - useful but fallible and prone to hallucinations.
Until we have a technical solution for that it’s the people that will need to adapt
If you'd like a LLM to provide a stream of characters that look like text, they're very good for that. They're not good for providing information, full stop.
This is like your local newspaper giving the wrong movie times or incorrect scores from the big game. People make mistakes but people’s whose job it is to provide information about something and who are unable to reliably provide that information tend to lose their jobs. Why should I “hire” OpenAI’s solution if it fails to be even as reliable as a human?