Step 2. Humans write SEO copy for machines to rank them higher.
Step 3. LLM writes copy for machines to rank them higher.
Step 4. Human uses LLM to try to distill the LLM generated SEO spam for any remaining signal.
Also to your point:
> SEO listicle garbage filling the internet.
the feeling that the LLM is better than what you described is going to be very temporary, then the mountains of LLM generated bullshit is going to overwhelm even LLM to make meaningful sense of.
The only "if" to all this is if we will destroy the LLMs by feeding them their own diarrhea. I expect a sort of natural selection here to play out, especially in the open source space. Ones that are trained on LLM generated blogspam will probably, I expect, get outperformed by ones that are trained on genuine information, or at the very least ones made using new techniques that adequately filter noise.
How will it learn anything new?
Yes, humans are notorious for only seeking out high quality, accurate data, especially when it conflicts with our priors.
To say nothing of our ability to assess the accuracy or truthiness of information in the first place (look at how many people take, on faith, that Chat GPT isn’t wrong as often as it is right).
For Google, they’re at the mercy of whatever the internet has.
You also have to consider the money angle. As using ChatGPT and other chatbots becomes more popular, people will stop producing garbage internet articles because they will be less popular and therefore less profitable. Bloggers who enjoy writing will continue to do so because it was never about the money, they just enjoy writing.
Further, the internet is only one small portion of information available to train on. There’s a lot of other data out there, including real-world conversations.
So now it's got great information about the Model T Ford but knows nothing about our new mars colony?
I don't think "just don't update the model" is a likely option.
"Just don't update the model, only feed it new information" is exactly how to get to the outcome of concern in this thread.
We've already seen exactly this happened with search. There's no reason to believe that LLMs are immune.
How would LLM upstarts be able to counter the massive commercial interests? As with google they will also succumb to prefer money over usefulness at latest when they have a wide user base.
There is also an even less proven way of distinguishing spam from signal with LLMs.
And not updating a model means that they will be stuck in COVID-19 era forever.
Given the “weights in a matrix” architecture of ChatGPT, I’m not sure it’s possible to store enough data to make the query practical to answer. Say there are a couple hundred intersections in my city. You have to store the token of “restaurant name” “close to” “intersection” for each intersection. I don’t know the size of Google’s Maps DB, but I would guess it’s several Gigabytes per city. From my understanding of the theory, you would need to store BOTH the LLM weights AND the Maps data for ChatGPT to have a shot at generating good answers for that type of query.
I’m happy to be wrong here. If I’m misunderstanding something, please let me know.
What Bing does is to use your query to search the web and use the top N search results in the context window for the chat.
However I’ll push back on your pushback. ChatGPT doesn’t need to be perfect to be a killer app. It is highly flawed. Maybe it was a bit to strident to say ChatGPT will kill Google search, but it’s strictly better for a lot of squishy queries that don’t have a factual basis.
How can I convince my boss to give me a raise? gets you a listicle on Google and a highly specific response on ChatGPT. And if some of the advice doesn’t apply, you can continue directing the conversation. It’s an idea generator, even if some of them are bad or don’t make sense.
It both picked out the interesting places of note, and then I asked it to plan them in such a way that made sense walking-wise (so I wasn't backtracking) and it did so without a hiccup.
But it does know a lot of things and can be super useful. Personally i think search engine is a terrible use case, unless you use the Bing enabled version, or bing chat.
I've used it to write pretty complicated scripts where I had no idea what I was doing, rebuild crusty httpd configs from first principles, explain disassembled code, explain regular code, explain configs, read dmidecode and lspci for me and make a pcie slot report... It's bloody brilliant.
Other: read and translated my blood tests. Accurately!
And yet it hallucinates URLs when I ask it to cite its sources. It's still Google search with a little patience for me.
If you are looking for a location on the internet, use a search engine. LLMs do not memorise the data sources verbatim.
If you want to know how to do something, it will normally give you a better answer than you would find by googling around multiple blogs. No location on the internet needed.