New ways we're tackling spammy, low-quality content on Search
blog.google
blog.google
> We’ve long had a policy against using automation to generate low-quality or unoriginal content at scale with the goal of manipulating search rankings. This policy was originally designed to address instances of content being generated at scale where it was clear that automation was involved.
> Today, scaled content creation methods are more sophisticated, and whether content is created purely through automation isn't always as clear. To better address these techniques, we’re strengthening our policy to focus on this abusive behavior — producing content at scale to boost search ranking — whether automation, humans or a combination are involved. This will allow us to take action on more types of content with little to no value created at scale, like pages that pretend to have answers to popular searches but fail to deliver helpful content.
It's that we can't tell the difference between the AI generated SEO spam and the human generated SEO spam, not that the AI has started to generate actually useful websites. And there was already so much human generated SEO spam that it couldn't be effectively moderated, but now the problem is a hundred times worse.
https://thehill.com/policy/technology/4492468-google-paying-...
That blog post was about the WA founder joining them, I don't remember reading something about integration.
Hi!
Personally, I think they deserve the money, they offered a better paid product and never sold advertisements, which I think is a scummy model.
I guess we should be glad that Google make bad product decisions though, to hurt the company.
I’m not paying Google any money though. And $10 a month is over my budget
"Use this pattern:
example.com##a:has(.b)
to create separate rules for these domains:
example.net
example.org
..."If I had trillions like google does, I'd invest in hiring a whole bunch of humans to curate a search index that is actually useful. It's something a new player will never be able to scale up to.
Algorithmic search has been useless for 5+ years. LLMs will just generate those spam pages faster, but they're already there.
They obviously have the training data. Even if they needed to manually label a bunch of it.
The problem must be that they make too much money from the web pages that are just "500 words rephrasing the headline plus constantly-refreshing ads between every paragraph"
This honestly seems unrelated to their poor search quality—hell, I'm even open to AI generated content for some queries. I blame catering to their clients and attempting to manipulate the content on the internet in the name of SEO for why I find useful results buried beneath products and ads.
Where should the ones you ask get information about the best options?
It seems much more likely that someone bought something for the wrong reason (easiest to find, being fooled by fake reviews etc) and then are recommending it onward because they are reasonably happy with it.
You ask multiple acquaintances who've purchased different things, and focus on the specifics of their experiences with those products or services, and figure out whether those things would work for you.
> and then are recommending it onward because they are reasonably happy with it.
If they're reasonably happy with it, why wouldn't I be reasonably happy with it?
This is going to be a honey pot for the antitrust authorities. Google's main defense is that its ranking is done in an algorithmic way and human rating can affect the result only in an indirect manner following their rating guideline, which is carefully reviewed by lawyers. When you add some subjectivity from random human, they occasionally make some unfair mistake then it makes all the way through the media outlets, you know what's going to happen.
Curated and sorted link lists were a big thing back in the early days of the WWW. So ... back to the roots, yes?
I'd love a way to search for things, and exclude anything created after 2020.
Sure, it won't help me find reviews of new restaurants, or understand new products or programming languages.
On the other hand, if I'm interested in reviews of a novel written in 1988, or how to solve some particular problem using calculus, that information isn't likely to go stale.
Google definitely isn't needed for that.
The temptation for a for-profit company to corrupt that index would be too great to resist forever.
What modern capitalism really seems to do is produce cycles of improvement to gain marketshare followed by enshittification to cash in. The average of the cycle leans more towards the enshittified state than the improved state.
But this isn't the problem.
People are just generating boatloads of "AI" generated pages with a bunch of ads. They don't need to be #1 in the search list, they just need some of their pages to appear in the search list frequently enough that the ads google places on those pages turn a profit.
Which google also knows, but again they're selling the ads, so they're getting a cut of it as well.
As long as google doesn't penalize pages on the basis of how many ad services and similar they use (and penalize google ads at the same rate) this will continue, and as long as it continues google profits from it. I get that a number of the engineers might not like that, but enough of them do that they long ago threw out basic privacy, accuracy of search results, and even the relevance of ads.
1. Tons and tons of content about the history of the recipe, how humans first domesticated wheat, when the ancient Egyptians first discovered yeast as a leavening agent, etc. etc.
2. Then, after thousands and thousands of words, there is a short recipe for bread at the bottom of the page.
In fairness, this is a hard problem to solve. Some sites (rarely) have good information on techniques or things like substitutions in all that copy. But the vast majority of the time all that copy is just useless drivel to appeal to search engines. So my question is whether a great but pithy recipe for, say, focaccia or whatever would make it to the top of the search rankings, or will Google always favor the verbal diarrhea that sullies all recipe sites these days?
I sometimes wonder how much of the costs associated with running a website are to pay for all the bandwidth for ads, SEO experts, etc… when a simple site with HTML and a dash of CSS could likely host recipes very cheaply.
Unfortunately for Google, G+ was probably their best chance at something like that. Now they just have stuff like time on page to try and figure that out.
Perhaps this is a case where Google could rip the recipe off the page and display it in Search.
Now we’re supposed to believe they definitely haven’t been abusing their monopoly, have just realised all this and are taking this grave threat to quality search Very Seriously Indeed, to ensure that you, dear customer, can find what you need. Couldn’t possibly have anything to do with encroaching threats to their core business.
Google: it’s been obvious for a good many years that you think us all a bunch of fucking peons. I, for one, will enjoy watching you and your clusterfuck of a product portfolio go the way of the Yahoo while I funnel my $10/month to Kagi.
I don't know of anyone who does this well, unfortunately.
As much as I appreciate what Kagi is trying to do, I don't think I necessarily want a replacement for Google. Or if I did, it would have been Google from the 2000s.
I want something different and better.
I DO NOT GIVE A FUCK about "popular" or "trending" matches. I want to see results that EXACTLY match what I am looking for.
Why is this such an impossible ask for search engines? 15 years ago it was the norm, not the exception. If SEO is the problem, then BLOCK THE SEO TRASH AT THE SOURCE.
-word is still a thing that works a bit.
...nope, I still can't.
Maybe it's a regional thing (I am in the US)? Or your profile is messed up? Try searching in incognito window.
I also just went to the site. When I went to tap the jump link, a modal popped up and blocked me. I dismissed that and tapped the link and it jumped me down the page to yet another request to sign up for a newsletter I don’t want. Now the top 3rd of the page is a video that started to auto-play; I had to tap to dismiss that. When I scrolled into the recipe section I got yet another modal trying to get me to sign up for an account so I can save the recipe… as if bookmarks don’t exist. I dismissed that, and now I can see the recipe. There are 3 ads in the middle of the recipe that break it up, some with more auto-playing video, and the scroll keeps readjusting itself based on ads and BS popping in and out. The bottom of the page also has an ad banner to dismiss, a bookmark icon that bounces occasional (which is certainly just another attempt to get me to create an account), links to other recipes are sliding in and out depending on the direction I scroll, to try to get me to go other places on the site… and to top it off, the Pinterest flag blocks content, depending on where I scroll, while sometimes being blocked itself by the aforementioned recipe slide-in.
How can anyone defend all this? It’s horrible.
Returned Love and Lemons first too. Never heard of them, but their SEO must be amazing.
Interestingly the first result for "brownie recipe" for me is BBC Food which is about as close to no-BS as you'll find.
I can see why GP was not impressed.
That would be a major problem for hand curated indexes.
I can’t even fathom how anybody thinks that this is fine on any level (from the mentioned site, the first result about brownies for me, and one of the first few paragraphs). For a while, I rather pay for cooking and recipe books, because wasted time on these texts would worth way more than that money. Especially that even the recipes themselves are terrible most of the time on the internet.
But after a thread a few months ago here on HN, in which people praised w3schools how it’s really a good site, I’m not surprised on anything. If the people who really should know how these things work, and that there are better free alternatives for all content on that site, even encourage this bullshit, then this won’t improve at all.
- you can take notes in the margins
- you can find your favorite recipes quickly because the pages are all stained and wrinkled
- you can reward the author for all the time they spent testing and honing those recipes
The real depressing thing here is the amount of whining about "wasted time" from folks looking for free cooking instructions on the 'Net. It's almost as bad as the "gimme teh codez" jerks on forums.
FWIW, I like this brownie recipe (and have purchased the book it comes from, twice): https://www.seriouseats.com/bravetart-glossy-fudge-brownies WARNING: contains even more words than the "love and lemons" page.
Cue a flood of aspiring L6 promo packets with detailed numbers on improved brownie recipe ranking.
https://www.google.com/search?hl=en&q=site%3Asmittenkitchen....
I realize adding `site:...` is probably fucking around. But trust me, her brownie recipes are worth it. :-)
Google lists dozens of recipe sites when I search for "brownie recipes" and honestly, how many different ways can you change a basic brownie recipe. So I'm sure almost any of them will do. Any seems like reasonable result. I'm not sure what algorithm Google should use to pick the best or most reliable recipe site.
I just happen to know, having cooked many of SKs recipes, that they are good recipes that almost anyone can follow.
http://everyspec.com/MIL-SPECS/MIL-SPECS-MIL-C/MIL-C-44072C_...
They can't do this as 95% of the internet would disappear overnight including a lot of Google ad revenue.
The proposed project has a paradoxical use-case defined, and as such will not function as expected.
Where do we send the invoice for stating the obvious? ;-)
lol
i.e. people get a local version of the results consistent with their local cultural values. You are correct in that the cons/spammers on this forum wouldn't notice anything change from their perspective.
Maybe we should deploy it on the ISP DNS servers for fun =)
What they should do is let people set preferences for how they want domains ranked in their personal results, and then use longer term trends on how those lists evolve.
Statistically, sociopaths are <3% of the population.
I’m not saying this is happening (I’m not saying anyone is a shill). It’s just making me feel a little “hmm” about it.
Another common response is
>Your table is wrong. I tried these queries on Kagi and got Good results for the queries [but phrase much more strongly]
I'm not sure why people feel so strongly about Kagi but, all of these kinds of responses so far have come from Kagi users. No one has gotten good results for the tire, transistor, or snow queries (note, again, that this is not a query looking for a daily forecast, as clearly implied by the "winter 2023" in the query), nor are the results for the other queries very good if you don't have an ad blocker. I suppose it's possible that the next person who tells me this actually has good results, but that seems fairly unlikely given the zero percent correctness rate so far.
For example, one user claimed that the results were all good, but they pinned GitHub results and only ran the queries for which you'd get a good result on GitHub. This is actually worse than you get if you use Google or Bing and write good queries since you'll get noise in your results when GitHub is the wrong place to search. Of course you make a similar claim that Bing is amazing is you write non-naive queries, so it's curious that so many Kagi users are angrily writing me about this and no Google or Bing users. Kagi appears to have tapped into the same vein that Tesla and Apple have managed to tap into, where users become incensed that someone is criticizing something they love and then write nonsensical defenses of their favorite product, which bodes well for Kagi. I've gotten comments like this from not just one Kagi user, but many.
The testing methodology is also odd, looking specifically for yt-dlp when asking for "download youtube video", when the web based downloaders that most of the big engines turn up on top make more sense for the average user, both due to not having to install a command line tool and due to not necessarily being on a device the tool can be installed on. I've often used the top Kagi result, and it has worked fine, yet has been labeled as scammy by the evaluator. IIRC that site also ranks high in Google's results. For adblock they expect specifically ublock origin, and consider ABP to not count despite it also being a very popular blocker.
Them complaining about why it's only Kagi users reaching out to him also seems unnecessarily aggressive, eg I've found their article through you, in the context of Kagi, so if I were to comment, it'd obviously be in that context.
However, how likely is this article to be brought up in the context of Bing, Google, ChatGPT or DDG, where there are many more carefully thought out and less dismissive comparisons and discussions out there?
In a sense it's targeted advertising, since HN users are very likely to be power users who care about a paid search engine for one reason or another. But it's still organic in the sense that the articles posted are ultimately just the regular search engine discourse or otherwise not content specifically crafted for the purpose of advertisement.
However, when I'm trying to find some archaic technical solutions that I know is out there on the web somewhere, Google just can't really do the job any more at all. Exact search doesn't even work! Kagi's filters DO work, and the ability to black list stackoverflow scrapers cuts out a lot of the useless noise. It's at those times when Kagi earns its place.
I have long hated that “free with ads” is the first and only business model anyone thinks can work on the internet, so I want to support companies that are willing to push back against that. Ads in search create a conflict of interest that I don’t like. Ads in most products create a conflict of interest I don’t like.
I pay for YouTube Premium for a similar ideological reason. I want to support the business model that doesn’t rely on showing me ads. In a perfect world, this would mean that the features and algorithms of YouTube would change to prioritize me as a user, rather than pushing constant consumption, but with Google its baby steps.
There's only "free, with ads" and "paid, with ads eventually".
https://www.amazonforum.com/s/question/0D54P000079nP9wSAE/ho...
Edit: but I'm glad Google is noticing their results are getting to be shit.
They don't show your history.
> the ad profile Google builds on you is probably more damaging to your privacy than a photo of the entire contents of your wallet.
I can take measures to prevent Google from linking together a full profile. It sucks that it's necessary, but once you log in it's effectively game over. And your payment info has enough stable identifiers that it's trivial to link you to pretty much anything else you've paid for.
---
The core misconception here seems to be the idea that there is such a thing as a honest for-profit corporation.
> The core misconception here seems to be the idea that there is such a thing as a honest for-profit corporation.
And I would say there's degrees. Even if all companies eventually slouch into a google-like organization, at least you can jump on the bandwagon until that happens and find another when it does. And maybe none of the wagons are headed to a strict privacy, user controlled data, and free-software utopia, but maybe not all of them are headed to Gomorrah.
You are the one claiming that Kagi is somehow special, trustworthy, and immune to the same influences that drags down everyone else.
> I see lots of promises, but very little proof
Yet you trust google and any company that might do business with google? This view is some sort of mind game you're playing alone.
My point is that Kagi's business model requires more trust in Kagi than Google's requires in Google.
From Kagi Settings page, search history toggle:
Save My Search History
Currently this option can not be turned on. Kagi does not save any searches by default. In the future we may add features that will utilize your search history and then we will allow you to enable this.
I encourage readers to think through the privacy practice implications of firms that may see your information "first party", and firms that make their money leveraging your information among third parties and the entire adtech ecosystem.
The easiest way for Google to detect spammy, low-quality content is to ... simply add up the number of Google Ads placements on the page.
Those flooding the internet with garbage content are doing so because the verbiage is tightly intermingled with Google Ad placements, that's the business model.
Google -could- go a step further and deactivate those accounts, but that would impact their revenue, so nah.
It's a long held complaint that Google's ad/analytics/tracking scripts slow down page load/page render times.
It's also long held that Google prioritise speed in rankings.
My guess is that they'd be excluding their own scripts and ad load times as part of that page load calculation, even though that is not a genuine result and would down-rank pages that utilise competitor ad platforms.
In a links post I just published (https://jakeseliger.com/2024/03/05/links-a-moon-landing-rapi...), I wrote:
Still, the Google search monopoly is under more serious threat than it’s been in the company’s history, and that may inspire real change. Amusingly, I used Google search to try and find a video of Sergey Brin saying that he’s un-retired to come back to work on AI, and Google search didn’t easily find it—but it turned up a bunch of spammy YouTube videos.
An eye opener for me was that I didn't question it. Has Google really fallen so low?