CEO of Google Says It Has No Solution for Its AI Providing Incorrect Info
futurism.com
futurism.com
or is that it didn't adjust on Friday and this isn't significant enough to adjust the next time the market is open (Tuesday)?
This is also why MS and other companies are leveraging OpenAI, as they can offload liability onto them.
Emily Bender is correct about how LLMs are just "extruded text machines".
At a Google scale company (and any public company for that matter), CEO's job is not to drive product decisions. That comes down to product leadership.
The CEO's job is longer term strategy (as in these are the types of categories we need to build features for, these are the the kinds of customers we need to target, these are the core financial metrics we need to hit).
> Google should have released a disableable feature leveraging Llms for whoever wanted to use it not shove it in their search results
The PM that owned this feature definitely has UX A/B tested it and found that the majority of users just don't care.
For example, I don't see auto-generated LLM results when searching on mobile which is the majority of Google Search impressions.
However, today, the markets, which can remain irrational longer than anyone can remain solvent, would punish Google for this; it has to go along pretending this is a real thing for a bit longer.
The question is, do the results justify the cost?
LLM is a good information transformation engine. But if I supply it with messy or wrong information in the context, then I’ll end up with hallucinations.
This is why most of the AI systems tend to produce junk. LLM part is good, but the underlying system just provides it with noisy and meaningless data.
I actually believe that Google will not be around as we know it in 10-15 years. And I do think OpenAI and Microsoft is playing a major part of that.
ChatGPT is already replacing search in many professions starting with SWE. IMO they should embrace it 100% and make Gemini a full 50% of the search page or a new button. Just making it a sad widget gives the signal its not the place to be to get AI ar all, not to mentio the subpar results.
Maybe sergey and larry are not taking the reins now. Sundar was a good peacetime CEO but maybe Google need some deep reorg of their product which only an original foubder can have the credential to lead
I know the benchmarks say they’re vaguely similar, but benchmarks said the same of Llama II, and it was almost useless compared to ChatGPT. (so I don’t quite trust benchmarks, more interested in how these tools perform in real use cases)
If we are talking about the ability of models to follow instructions and carry out concrete tasks (as in products or inside RAG systems), then Gemini Pro 1.5 is currently on the eighth place in our benchmark.
Academic benchmarks, HF Leaderboards or LMSYS Chat arena will have different numbers.
That's why I have my own set of simple benchmarks that I'm not going to publish. Everybody can easily prepare such a set - in my case they are various programming tasks that should generate determined output. It is not easy to automatically qualify the quality of code, but at the very least you can filter the results by invalid outputs. With reasonably high number of tasks and their complexity, this can be a fair estimator - provided that it's never published publicly.
My approach is similar - closed source benchmarks with prompts and tests from real LLM-driven products (mostly around boring business automation and enterprise workflows).
Although it would be neat to upgrade the setup to work on the synthetic data. This will at least make benchmarks shareable publicly (not just the results)
Ive used gemini only for vision tasks on document and it was absolute CATASTROPHY (hallucinating the content instead of actually reading it)
The woke part is also extra unhinged on gemini (hating white people is fine but hating other is racist etc...)
Thats not the point of my message tho, the problems could be corrected later. Its more about the product vision
The way for Google to beat OpenAI is to recognize AI chatbot usage is going to make a dent in web search usage and further develop Gemini to make it jump from the current GPT-3.5 to 4.0. When combined with a huge context window, they could quickly make people switch and own both web search and AI search. With the current strategy, they just make people migrate away from their web search.
There are still 60 70% people who havent tried it like my parents. They still have a window to stay relevant one way or the other
I mean, even then you’d surely prefer to be pointed to an article by an actual human than getting whatever an LLM spits out?
Judging by how much crap there already is on the web these days, there would seem to be a market for it.