Quoted words in the search should return only pages with that actually on the page. That used to work. Now it often shows pages that don't contain it all, not in the search results page itself, nor on the page when you go there. So you'd think it just ignored the quoted text and gave me results without it instead, right? Wrong. Removing the quoted text to either unquoted text or removing it entirely both result in different results. So it DOES process the quotes SOMEHOW. But it's not a clear RULE because probably it's also just processed as an input to the semantics engine. Just tell me you don't have any pages with the text instead, please. Which it actually also some sometimes...
It's an unfixable mess. And I don't think this can be turned back. Building a competitor costs hundreds of billions. A competitor will likely end up taking the same approach anyway.
I just wanted a list of google's big search engine updates but even searching this I get SEO-d blogspam ABOUT THE SEO IMPACT OF THE UPDATES.
And please spare me the excuse that now Google can answer questions. It can't, it just answers with snippets extracted from websites it deems relevant, and often the answer is flat out wrong or irrelevant.
I just hope the ML/AI mania that has taken over these big tech companies proves to be a fad that just goes nowhere like it did in the 70s and we return to plain old algorithms and good software engineering.
Of course their actions have not been perfect but it is a mistake to say that their search would be better as a "plain old algorithm" no matter how well-engineered it is.
I'm certain that search results would be worse than they are now if the algorithm was just "grep but for the whole internet". Or that, in that case, the careful complexity necessary in each search to exclude all the SEO garbage would be unbearable.
Now there's still spam, and search is useless.
Also, I don't think people remember how much interesting stuff is on the internet. There used to be tons of results from small sites of blogs which are still there, but not listed on Google anymore. Modern Google has made the internet incredibly smaller. Everything is still out there, hidden from our searches. It's like their algo has been tuned to favour silos, spam or content silos that is, to the detriment of the long tail of small, hobbyist websites with high signal-noise ratio.
Worse search results -> more Google searches -> more ad views
Google has completely lost the plot if DDG is the engine with the better results.
I wonder what the consequences are for discoverability on the web. It can’t be good. Maybe I’m waxing nostalgic, but it seemed like there was a time I discovered new and cool web sites. Now all I discover is ad spam clogging up the tubes.
I feel the exact same way too. I can't even discover blogs with solutions any more (even ones I know exist because I've seen them before), it's all just spam.
I think it is a second level machine learning system that has gone haywire and nobody knows how to fix it, or they lost the keys, see https://news.ycombinator.com/item?id=5397797 for a real life (or not : ) story.
I recently searched the name of a YouTuber I was watching, the top "People also ask" suggestion was "Is [name of youtuber] married?"
I didn't particularly care but clicked out of curiosity. It expanded to show a snippet from an old (1980s) NY Times article about a completely different person getting married long before this particular person was even born.
Google AI using other companies' content to provide wrong answers to questions I didn't even ask... That says it all.
It's now being used by law enforcement. Sponsored by an errant swat raid near you. Don't worry, they'll prosecute you anyway to cover their ass. Can't risk losing their pensions, ya know. /scared
https://www.policechiefmagazine.org/product-feature-artifici...
I don't know why Google would even suggest this -- Miep was one of Anne's helpers. Imagine all the other people out there having their names unfairly smeared by Google's algorithm.
Be suprised the answer isn't drive/read/shoot^^.
Translates to: “The users can’t find what they’re looking for, so they’re clicking around a lot.”
I have stack exchange, github, and all the canonical documentation sites for my projects in mine.
Google clearly isn't even trying.
So many of these sites are polluting search results for months. It isn't a case of sites that pop up for a few hours until Google notices and blacklists them.
Google Search has gone so far downhill. I'm not sure what they're optimising for. Long-term irrelevance, it seems.
There are a lot of very smart people from all over the world putting everything they have into getting to the top of that site. It's not a trivial task for them.
A stop-gap solution would be to simply penalize anything with ads. Legitimate websites will still have enough "SEO juice" (for the lack of a better term) to offset the penalty, but brand new copycats with otherwise no inbound links to them from other legitimate websites (and no other business model beyond scammy ads) will never be able to rank high, essentially removing the incentive from setting up these sites in the first place.
Not to mention, these problems can be identified, prioritized and then tackled manually one-by-one. Stackoverflow copycats can be dealt with by downloading the SO data dump, parsing it and then severely penalizing any website where the bulk of the content matches the dump. Pinterest can be penalized by simply excluding it from image searches until they actually display the searched image without asking for login. So on and so forth.
You can't win them all, and you can't train an algorithm to win every time either, but you can manually observe what's happening and deal with the biggest problems.
The problem however is that the current status-quo is good for Google. They've got the monopoly on search (every other one is typically even worse when it comes to these issues), and spammy copycat websites happen to have Google ads or analytics so Google actually benefits from them too.
They have a site that’s been polluting image search results for years, isn’t even trying to hide (unlike the SO copycats which could technically churn out an infinite amount of domains to work around bans) and they can’t even deal with that.
And when you get to that state, you've pretty much turned the entire web into content marketing for your ad sales business.
So - of course search quality suffers. Search quality is not the point.
At the end of the day Google doesn't exist for the good of the public, they exist to return the highest possible investment for shareholders. It's literally the opposite of their job to improve search quality when competition doesn't force them to (because investors see that as a waste of money that could have been spent on further eroding people's privacy to enable higher and higher ad revenues.)
Related is if you have a lot of content from git repositories that are mirrored from different locations (GitHub, GitLab, etc.), all of which are showing the same content. Or if different sites are hosting versions of public domain texts. You don't want to derank those results, even if they are similar to the "copy a popular site" websites.
So they could know if a code snippet was already X years old? And from where, originally. (Unless it got edited a lot but that's not the case here?)
The core function of that is actually pretty simple:
1. Strip all X/HTML tags 2. Run `diff`
Sure, it's not perfect, but an organization that pursues academic quantum computing research can sure as hell afford to run the results of the above against an AI to check for similarities.
In the case of things like other sites copying from stack overflow, github, etc. you can figure out which is the source of the information.
Lets day you have decided to make github the source of that information, and derank any other sites that also have that information. As a result, you will derank gitlab for having git mirrors of projects, as the source will match that on github. You will also derank sites like lkml as that contains the commit message descriptions, and the patches will partially match content from the kernel source hosted on github.
Lets also say you've decided to make wikipedia and other wikimedia sites the source of information. Congratulations, you've now deranked project gutenberg for hosting the same public domain texts as wikimedia, along with other sites like the official sites for authors like Jack London. Plus any blog that includes portions of these to discuss or analyze them.
Do you want a blackbox AI making these decisions?
I interviewed with a 3 or so person company that did just that I think.
That was over a decade ago.
I assume they're optimizing for the most common denominator: standard, non-power users. Early search didn't return great results for human-like questions such as What are the wavelengths of the colors on the visible spectrum? The result might be within the first two pages, but that query had too many irrelevant search terms. A better query would have been wavelengths color visible spectrum. That query only has the necessary key terms. Sometimes queries required the user to know search operators (e.g. exact match, date range, synonyms) just to get relevant results.
The average person probably didn't know that early searches gave better results when constructed in the second way. Google changed search to adapt to how normal people search. Now the human-like query will return good results. Combine that with locales, search history, and personal interests, even the most basic user can get worthwhile results from asking Google question. The cost is that power users who understand operators and the power of key terms get less relevant results but likely still correct.
Then again, I have been using Google Maps for 16 years and it still cannot show me distances in km rather than miles. A boolean setting ffs that I've had to manually switch probably a thousand times.
If billions of dollars of AI R&D can't even figure that out...
What it looks suspiciously like to me is a lack of an effective feedback loop for user frustration — if it takes a number of queries to get correct results or someone doesn't stop using Google search entirely, this would be easy to confuse with improved engagement and I'd especially believe that managers whose job it is to get a number to go up are not in a hurry to question whether that growth is meaningful.
So the nitwits just harvested too early, before the HN discussion was fully grown?
I want it to absolutely destroy Google/Microsoft.
Discussion: https://news.ycombinator.com/item?id=29392702.