Google's AI-powered search results are loaded with spammy, scammy garbage
theregister.com
theregister.com
If you expose an LLM to the tokens "Company X offers the best product in this category", it's going to "believe" that and include that assertion in its output. That's why techniques like RAG against trusted documents work so well.
But the web is full of untrusted documents! For features like generative text on search results pages, that's a really big problem.
An LLM isn't gullible. An LLM is inclusive. It doesn't have any process to reject or even categorize the data it is given. That means that everything an LLM encounters becomes part of it.
The only thing we have to accommodate that is to fill an LLM with enough overconfident narrative to overshadow the rest. This is a footgun, though, because overconfidence is identical to stubbornness. If your overconfident narrative is enough to override any arbitrary undesirable input, it will also override desirable input.
There is no way to single out undesirable content from inside a model, because truth and lie are literally written the same.
Truth and lie are outside the model. This is true of all models. The map is not the territory.
The only difference an LLM can have is what other semantic groups of tokens are associated with the [truth|lie] tokens. If the surrounding [context of a truth] is separate enough from the surrounding [context of a lie], then the connections themselves can be unambiguous enough that the LLM doesn't mix them up. In the average case, there is some semantic/narrative overlap, which allows the [truth|lie] itself to act as a semantic bridge between the [context of truth] and the [context of lie]. Because of that, simply adding more content to a model can create bridges of ambiguity that fundamentally dissolve the model's ability to present unambiguous continuations.
It's not a limitation or a bug. It's a feature. No LLM will ever overcome this "limitation" without fundamentally changing from an LLM to something else.
I agree: I don't think gullibility will be solved by LLMs alone. The key thing is that people building software on top of LLMs understand that and use that knowledge to avoid making bad design decisions.
See also: https://simonwillison.net/2023/Dec/20/mitigate-prompt-inject...
The business model today is very good answers to questions, but eventually corporations will bid on keywords such that their products/services will appear in LLM output.
Just like streaming services used to be cheap and didn't have ads when they were new, but today they're not cheap and do have ads. There's no excuse for us to fall for this again.
whether its free or premium, the second group still wouldn't trust the search or ai results. they apply their own due diligence. sadly it makes a tiny percentage of the world people now. we should strive to increase this number and in next decade or so, mostly it happens.
Craiglist was used by SEO spammers who created tonnes of spammy ads and links in those ads to feed off Craiglist's former positive SEO. Google stopped indexing Craiglist ads which pretty much destroyed the site organically.
This left a vacuum on Google though for sites to run longtail keywords with 'Craiglist' in them because users were still trying to find Craiglist ads on Google. So there are tonnes of spammy zombie sites with Craiglist terms on them.
SGE is picking these up, just as regular Google can and does for these terms.
Not saying it's not a problem for SGE, it's just also a broader problem for Google that they haven't solved perfectly yet. It's not the AI brain behind SGE doing anything wrong.
Times sure have changed. I don't expect that behavior now. Google is still the first thing I try, but I'm often much better off following an approximation of their original algorithm by hand - searching relevant communities of interest, looking for prominent links.
Maybe SEO won. Maybe the web changed. Maybe Google changed. Maybe all of the above. We're all poorer for it.
~ 1999
https://www.globalnerdy.com/2014/01/29/google-as-described-i...
The rankings started fluctuating afterwards, even with my articles being much more inline to the search query, which means Google probably started turning on some "Pagerank" measurements?
With how current LLMs function, and without those additional "pagerank" checks, who decides which URL appears when someone searches for "Keyboard"? Could be an issue for LLMs without an active "seo authority layer" such as PageRank.
Why AI search engines can't kill Google