ChatGPT Opened a New Era in Search. Microsoft Could Ruin It
wired.com
wired.com
The only 'killer' use-case for LLMs is 'summarization'. That is it.
For non current events I'll take a locally run LLM over privacyless google even if the results aren't quite as good. Google knows this and they are scared.
https://weaviate.io/blog/cross-encoders-as-reranker
but these are significant because they noticeably improve search ranking results and a development like that happens every 5-10 years or so. Getting a bunch of results from Bing and then reranking them with a cross-encoder is a low-hanging fruit in terms of making a better than status quo search engine.
If the cross-encoder is well-calibrated then you really can merge search results from different sources the way IBM Watson did back in the day and not have to calibrate the individual sources.
You're certainly not going to run them for every document in the collection but can you run them for the top N? I think the answer to that is yes.
It takes > 1 second(perhaps 1.5-3s if you embed paragraphs not sentences) for re-ranking k >= 8 on a decent say 32GB CPU instance..(requires encoding each document passage with the search query top k times)... if you have GPUs available for your search inference sure it's just a fraction of a second, but right now I would say it's significantly more onerous than bicoders that get you the first jump in semantic search performance (10-40x slower in practise on CPUs) so the use case differs.
1) Re-ranking is only good as the results you originally get. If you get garbage in, you will likely get garbage out, no matter how you re-rank
2) Re-renking using semantic search does little to help with the main problem in search (SEO spam, low quality content). These 'bad' pages typically rank well in semantic search and re-ranking would probably just help boost them.
A good example is something like Brave Search, Independent index with independent AI and its own AD network.
But this change really shows that Microsoft believes Bing can take on google now(and it does! BARD is absolute garbage compared to gpt-4), and it doesn't want anything or anyone getting in its way.
If Google stopped developing Chromium or decided to move its Chromium development to a closed-source project, other chromium-based browsers would have hard time keeping up with the necessary development to keep a web rendering engine up to date.
We need true open source alternatives asap.
As for alternatives, phind is pretty good and uses its own independent model, but it might be killed by this change, as AFAIK, they use bing for results.