Google is apparently struggling to contain an ongoing spam attack
searchenginejournal.com
searchenginejournal.com
Years ago, Google had an explicit policy that sites that showed different content to Googlebot than they showed to regular unauthenticated users were not allowed, and they got heavily penalized. This policy is long gone, but it would help here (assuming the automated tooling to enforce it was any good, and I assume to was).
More recently, Google seems totally okay with sites that show content to Googlebot but go out of their way not to show that content to regular users.
I was working on the system that picked ads to show on our pages (we had our own internal ad system, doing targeting based on our own data). This was the most computationally intensive part of serving our pages and the ads were embedded directly in the HTML of the page. When we realized that 90% of our ad pick infrastructure was dedicated to feeding the crawlers, we immediately thought of turning ads off for them (we never billed advertisers for them anyway). But hiding the ads seemed to go directly against the spirit of Google's policy of showing their crawlers the same content.
Among other things, we ended up disabling almost all targeting and showing crawlers random ads that roughly fit the page. This dropped our ad pick infra costs by nearly 80%, saving 6-figures a month. It also let us take a step back to decide where we could make long term investments in our infra rather than being overwhelmed with quick fixes to keep the crawlers fed.
This kind of thing is what people are missing when they wonder why a company needs more than a few engineers - after all, someone could duplicate the core functionality of the product in 100 lines of code. At sufficient scale, it takes real engineering just to handle the traffic from the crawlers so they can send you more users. There are an untold number of other things like this that have to be handled at scale, but that are hard to imagine if you haven't worked at similar scale.
The best explanation I can come up with is that a failure to notify them of a change makes them look bad when their search results are out of date. Especially if the failures are malicious, fitting in with the general theme of the article.
[1]: https://www.indexnow.org/
[2]: https://www.bing.com/indexnow
[3]: https://blogs.bing.com/webmaster/october-2021/IndexNow-Insta...
Should be easy to crosscheck the reliability of update notifications by doing a little bit of polling too.
"most of everything is shit" comes to mind, but "most of email being spam" and "most of web-traffic being porn" are well known.
Actually everything we discussed here is the result of genuine human activities.
I wonder if there's some sort of "time" threshold for how long an AI can speak/write before it is identifiable as an AI to a human. Some sort of Moore's law, but for AI recognizability
- https://developers.cloudflare.com/cache/advanced-configurati...
And on the other side, that means every customer or ad placer, has to try and filter all the bots so people with actual credit cards and money will see the Google, TEMU, or FB ads (or others).
In some ways, almost feels like Microsoft is griefing online search by burying it under massive robot crawls. Like an ad DDOS.
On top of that, they were ads for doing more of what the user was doing right then, tailored to tastes we'd seen them exhibit over time. Our goal was that the ads should be relevant enough that they served as an exploration mechanism within the site/app. We didn't always do as well as we hoped there, but it was a lot better than what you see on most of the internet. And far less intrusive because they weren't random (i.e., un-targeted). I have run ad blockers plus used whole house DNS ad blocking as long as I've been aware of them, but I was fine working on these ads because it felt to me like ads done right.
If we can't even allow for ads done right, then vast swaths of the internet have to be pay-walled or disappear. One consequence of that... only the rich get to use most of the internet. That's already too true as it is, I don't want to see it go further.
In fact one of my bigger problems have been that Google has served me generic ads that are so misplaced they go far into attempted insult territory (shady dating sites, pay-to-win "strategy games" etc).
This is the case. Advertising is a scourge, psychological warfare waged by corporations against our minds and wallets. Advertisers have no moral qualms, they will exploit any psychololgical weakness to shill products, no matter how harmful. Find a "market" of teenagers with social issues? Show them ads of happy young people frolicking with friends to make them buy your carbonated sugar water; never mind that your product will rot their teeth and make them fat. Advertisers don't care about whether products are actually good for people, all they care about is successful shilling.
Advertising is warfare waged by corporations against people and pretending otherwise makes you vulnerable to it. To fight back effectively we must use adblockers and advocate for advertising bans. If your website cannot exist without targeted advertising, then it is better for it to not exist.
Oh, and they don't get to vote because voting day and locations can't be advertised by the government, especially in targeted mailings that are personalized with your party affiliation and location. The US Postal Service will also collapse, so those mailings can't go out, even if allowed. At least the rich can still search for their polling location on the web [<- sarcasm].
None of that is okay with me. More/better regulation? Yes! But our world doesn't know how to function without ads. Being absolute about banning ads is unrealistic and takes focus away from achieving better regulation, thereby playing into the hands of the worst advertisers.
Not my problem. Those companies, and any other with business models reliant on advertising, don't have a right to exist. If your business can't be profitable without child labor, your business has no right to exist. This is no different.
Seems like that's exactly what they did...
Years ago (early 2000s) Google used to mostly crawl using Google-owned IPs, but they'd occasionally use Comcast or some other ISPs (partners) to crawl. If you were IP cloaking, you'd have to look out for those pesky non-Google IPs. I know, as I used to play that IP cloaking game back in the early 2000s, mostly using scripts from a service called "IP Delivery".
It was, however, one of the most difficult types of spam to detect and penalise, at scale.
And then there’s whatever Pinterest does, which seems awfully like cloaking or bait-and-switch or something: you get a high ranked image search result, you click it, and the page you see is in no way relevant to the search or related to the image thumbnail you clicked.
For context, my team wrote scripts to automate catching spam at scale.
Long story short, there are non spam-related reasons why one would want to have their website show different content to their users and to a bot. Say, adult content in countries where adult content is illegal. Or political views, in a similar context.
For this reason, most automated actions aren't built upon a single potential spam signal. I don't want to give too much detail, but here's a totally fictitious example for you:
* Having a website associated with keywords like "cheap" or "flash sale" isn't bad per say. But that might be seen as a first red flag
* Now having those aforementioned keywords, plus "Cartier" or "Vuitton" would be another red flag
* Add to this the fact that we see that this website changed owners recently, and used to SERP for different keywords, and that's another flag
=> 3 red flags, that's enough for some automation rule to me.
Again, this is a totally fictitious example, and in reality things are much more complex than this (plus I don't even think I understood or was exposed to all the ins and outs of spam detection while working there).
But cloaking on its own is kind of a risky space, as you'd get way too many false positives.
And byw (unless we are talking about different things) it was possible to get to the image on target page, but it was walled off behind a log in.
https://developers.google.com/search/docs/crawling-indexing/...
and the there are AMP pages which is Google Enforced cloaking...
It’s about serving different pages based on User Agent.
https://developers.google.com/search/docs/appearance/structu...
basically cloaking + json-ld markup
That + spam sites spamming as many keywords as they can just mean whatever you search for 95% of the sites are spam after the first page.
Idk why we've let the Internet get like this. There's gotta be a way to sign off on real/trusted content. That's certainly not ssl certs. Could probably crowd source the legitimacy rating of a site or something.
That's another reason why people flock to the big names, reddit, youtube, etc. It's like McDonald's, people know that what they get this time will be exactly what they got before.
See also, pages behind Red Hat and Oracle tech support paywalls.
Still using a lot of other google stuff including gmail and maps. Just not search anymore.
But also, for the past few months, I’ve completely stopped searching the internet. ChatGPT-4 does the job way more effectively and I don’t see why I would go back to searching the internet (assuming the chatgpt experience doesn’t get nerfed in some way).
Much better than bing gpt in my experience.
Though I asked it how to dry a comforter and it wrote me pseudo code lmao.
For quick intros to technical issues GPT4 gives a decent summary if the topic has been around for a while.
For going in-depth, though, I still rely on technical docs…
But for now you must keep your searches to a very limited, sanitized, corporate, non-copyright infringing, non-adult set of knowledge. And it's impossible to know what that will be beforehand, which makes using the tools very frustrating.
For example, try searching for anything medical related. Even if you're clear that you not looking for medical advice, you're just looking for info, it won't give it (sorry, as an AI I can't give medical advice). I imagine this is very frustrating for medical students.
And yes, I'm sure that I could coerce it into responding. Pretend your my grandmother telling me about her old medical recipes or some such. But that's still too annoying to do as anything except for testing the boundaries of the tool.
It needs to directly answer all the questions that are given to it. If it must finish out the post with a disclaimer like "I am not a doctor/lawyer, this is not medical/legal advice" or whatever, that's fine.
This is absolutely not to say that Google can be considered ‘good’ these days.
That's creativity and innovation Microsoft style.
Any time I do end up going to Google, I’m so disappointed by the search results that I just leave. The only thing its good for now is searching site:reddit.com
I also have the same experience where kagi doesn't find something I think it should, so I go try google. Holy hell is google bad. Shockingly bad. I genuinely can't believe how bad it is now compared to 10 years ago.
Kagi is designed to show you what you ask for, and not for showing you the ads you're most likely to fall for. It simply takes your query and returns results matching it. That's really it. It's sad that "does what you ask" is a defining feature, but that's what kagi is.
They have advanced features like "lenses" that bias your results toward a specific topic like programming, research, forums. It also lets you add weight to certain domains. For example, I have Pinterest and Facebook totally blacklisted, and some small sites boosted in my results.
They also support advanced query syntax with double quotes, +/-, and other operators.
Which is to say, kagi is the standard for "search engine that works". Nobody else sells a search engine that just works and does what you tell it to, which is literally all I want out of a search engine. It's mundane and unsexy, but it works, it doesn't advertise at me, and it lets me get my work done faster.
- Especially if you feel the need to use doublequotes.
But yes, it definitely misses some results that Google doesn’t in that regard.
That’s probably different in other regions, but it’s been my experience with pt-br
They almost always follow some sort of structure where the actual answer to your question is all the way at the bottom of the page.
Most of the content appears relevant on the surface, but when you actually read it, it’s completely generic junk. Stuff that a high schooler would use to fluff an essay to hit a word limit.
You have to scroll past the adverts - that's why they exist. These sites are generated from templates.
When I think about my typical web queries across the past year or two, it seems more and more likely that I'd be better off replacing Google with several purpose-built systems, none of which search the "entire web" (whatever that even means anymore). Technical queries? Just search StackOverflow and Github directly. Searching for a local venue of any kind? Search against a dedicated places database where new entries have to pass at least a cursory scrutiny. (Arguably Google Maps or Yelp already serve this purpose today, but I'm not sure if they have enough vetting today). Medical question? Search across a few sites known to be trustworthy.
We have become accustomed to go to Google because it's more convenient to type in a movie title, "chinese restaurant philadelphia", "flights to miami 4/12/24" or "Error code 127 python" into the same single place, but something tells me we'd be better off if that one place made some LLM-assisted guesses of what kind of search it is, and then went to a specialized search that is curated. If we go back toward the DMOZ/Yahoo model of directories that humans curate, I wonder if we could even reverse the trend toward spam and clickbait that has been so lamented in recent years.
Never, ever show me Pinterest when I do an image search.
I imagine my search results would improve quickly in short order.
Better still, aggregate those lists from all users and you can improve search for users that have not yet built up a black list.
Try https://github.com/iorate/ublacklist, otherwise Kagi also has a similar feature.
the yahoo model collapsed for this very reason. back when you went to more than 5 websites to look at screenshots of the other 4, the directories would not necessarly show you the latest thing, because it wasn't on the list of sites manually added to each directory.
i think the current problem with google isn't to do with spam. i think google has become complacent because their ads are on all the sites anyway, so the function of "maximize revenue per search" doesn't actually care if you find what you're looking for, because you will get shown google ads anyway, and will be coming back to google anyway. in fact, they probably get to show more ads by feeding you bad results, because then you're loading more pages. this didn't used to be the case when google search was on top of spam sites, but it doesn't feel like they're doing anymore algo updates to curb the current trend, and spam sites have caught on to what ranks higher in the results.
Hmm... You started to backpedal but then persisted. In the today world, your SO competitor would have that (slim) chance to rank if you started getting links from sites like HN or from people on Twitter who matter and know about tech. This would give you some PageRank and then you'd start possibly ranking in Google (in theory. In reality, no you probably wouldn't rank for anything since you're competing with 1,000,000 spam sites including whole verbatim clones of every page on SO that Google can't even get under control)
If any directory would be worth using, it would be run by humans who would HAVE to look at each submission. They could also look at who's linking to it, and evaluate "Is this a backlink from like, a gibberish page on `prawns-01-blork.info` or from like, Joel Spolsky's Twitter account?" Yes, it would take a lot of work, but like, it would be creating a truly useful product that people might pay for. And we have examples of other professions where "just rubber stamp everyone who pays" is frowned upon, like building inspectors and journalists. It's a hard problem, but it's far from hopeless.
The beauty of search engines (in theory) is that you can find something NEW. Keeping the "open web" out would just entrench and ossify the current players.
Personally I'd rather see a standard which allowed you to add as many directories as you wanted to what your search engine would metasearch across. This also avoids the political problem of "who decides what's trash" -- if you want to add a directory whose main deal is they'll add literally any site, you could. If you want to only add directories which don't allow any <insert hated party> leaning content, you could do that.
I have seen others comment on this as well.
I cant say I know when the trend started.
Could this have been running / going on for a long while without getting the scrutiny, it needed?
Is this spam attack the final act?
Basically it used to be optimized to be a sharp knife but now it's optimized to be a safety knife.
Instead I think search results have been getting worse for everyone because of SEO. Companies want to optimize for number of ads viewed, not quality of content, thus quality goes down in favor of clickbait and keyword stuffing.
I don’t think there’s much Google can do here to resolve this issue. It’ll always be a game of cat and mouse between Google and companies using SEO to push more ads for less money.
Garbage content seems to be making massive gains year on year, while informative or high quality content has stagnated or even declined from data decay
That said I only comment out of casual interest as I stopped using Google more than a decade ago.
TL/DR is that spammers are likely exploiting two loopholes. 1. Longtail keywords are low competition and may trigger different algorithms. 2. Some/many of the search queries the spam ranks for trigger the more permissive Local Search algorithm
Plus there are other reasons why those sites are getting through, which are discussed in detail in the article.
Are you saying you're the author of TFA?
Startpage is a good wrapper for Google if you care about privacy. (Or it was, I haven't checked in many years)
DDG is worse than google, IMO. Bing works, I guess, but I trust Microsoft almost as much as I trust google.
At this point I've given up on Google. If I can't find it in kagi after a bit of effort, I'll either work around it or ask a person I know in the given field.
God, the internet sucks so much now. I miss the early 2000s :(
For queries that fail on kagi, I check ddg and google, but it never helps.
Kagi’s FastGPT works well on queries where the search engines fail, but is worse at search on average.
A search engine that ignores my query and shows me something else is not super valuable to me
Google is a marketplace, and they let most "engaging" results that adhere to a certain content structure win.
By now most paid results offer more value for the users than the organic results. Cause thats what Google wants. Click the paid, ignore the crap.
No. Google is a glorified advertising agency who just wants to make the absolute maximum profit for the absolute minimum amount of work.
> By now most paid results offer more value for the users than the organic results. Cause thats what Google wants. Click the paid, ignore the crap.
Your first paragraph implies that the paid results are crap too?
I don’t generally compare to Google, so I can’t say for sure that the results are ‘as good’, but my experience sure as shit is better.
I search for a thing and I get a page with links. Usually the thing I want is in the first page.
Sometimes it isn’t, or I’m searching for something that I know is recent that Google probably has a later version of, so I just add the !g to the search and there I am at Google.
It’s great. It works. It’s not stressful or horrible or annoying. I recommend it.
Google gets their money. They don't care about anything else.
You are playing up the noise and playing down the signal. Search, Gmail, and YouTube are far from spam. There are obviously many scenarios/URLs that contain spam, but all 3 of those products are overwhelmingly useful.
Here Sivvi, toyou, Namshi, noon, and OUNASS are all brands of shopping websites and you can see their logos in the image.
Clearly this is some sort of keyword spam, though it's hard to tell more than that from your screenshot. It's also not clear why they'd bother to use HTML entities... a bug in the spam code? Or perhaps exploiting some parser differential between different twitter systems? Who can say.
Hypothetically say a website has an internal service to index posts for keywords for search, that just so happens to unescape HTML entities during keyword normalization due to a seemingly harmless bug.
Plus a second internal service to identify keyword spam that _doesn't_ do any HTML entity unescaping (because why would you?)
Then you could end up in a situation where a spammer uses HTML entities to avoid spam detection while still showing up in search results. They hope that the user ignores the nonsense text and just clicks their link based on the image (a list of big shopping brands in the middle east) instead.
eg (November 2022) https://usa-casino.com/casino-news/spammers-take-over-google...
and right now I just tried a query from the article above "Alabama Casinos" on Google maps, and sure enough I see "Bovada casino" (offshore "illegal in US" casino) affiliate links in the third spot in the "locations" list.
yuk.
All I get is the same damn advertisement for the Internet provider I already have. You’d think they’d be able to say “Google don’t send this ad to our customers.”
install with https://www.phind.com/search?q=%s
Google is an ad company and it's entirely reasonable to assume they are simply selling lists of their own email addresses. Same with Yahoo! Mail! Which! I! Ran! The! Same! Experiment! From! With! The! Same! Results!
Pro Tip: if you're going to boast on how you're spamming Google, then expect it to be shut down, especially if it's a hole in their algorithm.
If anything, this will help them sell more ad spots on Google Search.
I haven't tried Kagi's FastGPT yet. How does it compare with ChatGPT 4? Is it updated regularly? What is its knowledge cutoff date?
Google wants to be too relevant , to the point it s unusable
Search is unusable.
It's better than the Rogan and IDW spam I was getting a few weeks ago.
Ye for like two weeks. Searching for Lady Gaga or whatever, get those disgusting video thumb nails in the result.
It is impossible to not look at the thumb nails since they are so disgusting. I wonder if that feeds the ranking somehow?
I’m willing to bet that >90% of Google Search users aren’t even aware that alternatives exist.
They are not suddenly going to stop using Google Search. There might even be a significant short-term increase in usage, for example, if they really need to go to page 5 to find the first relevant result.
They will use it less if it becomes a waste of time to try.
Over two years now I see specifically two people deterministically failing to find stuff through G search, and yet they still start with it in an almost Pavlovian ritual. Sometimes they will desperately scroll down, click on a blatantly irrelevant result and switch to facebook or some other bookmark aggregator of theirs to continue their brutally inefficient search process, thumb flick after thumb flick after thumb flick.
Spam on!
Source?
It's incredible how brain dead Google have become.
Hah. Google's been 98% spam for well over a year now.
Try googling "What is the fifth book in the wheel of time series" and see what you get? All spam.
1. Wikipedia - The Fires of Heaven
2. Amazon - The Fires of Heaven (The Wheel of Time, Book 5)
3. a bunch of videos which I skipped over
4. Wheel of Time Wiki - The Fires of Heaven
5. Macmillan Publishers - The Wheel of Time, Books 5-9
6. Goodreads - The Fires of Heaven (The Wheel of Time, #5)
7. Novelnotions - Book Review: The Fires of Heaven (maybe this is spam?)
8. Esquire.com - Wheel of Time Books in Order (not specific to book 5)
9. FictionDb - Wheel of Time Series in Order (not specific to book 5)
10. Barnes And Noble - The Fires of Heaven
I am wondering how different everyone's results must be. [Also, I don't know how to format a list on HN…]
I wonder how many people on HN are infected by such malware and don't realize it? A lot of the complaints about search results are clearly not this, but when someone complains of outright spam for reasonable queries, I do wonder...
Result quality can be significantly different
Wikipedia, Goodreads, Fandom.com and MacMillan Publishing; these all seem to be reasonable results. I could share the whole page if I could find a place to upload my screenshots (RIP imgur)
Google has been borderline useless for productive work. I always attach Wikipedia or Reddit to my search to get anything useful
Google gave the correct answer. Didn't see any spam.
I find this hard to believe. How do you even measure for this?
I'd love to see a few more examples of searches you are making that show spam, because the example you gave provided me with the appropriate results. I almost suspect you are either being disingenuous or just have some malware on your computer.
They'll complain at the thought of paying for YT premium ("the internet should be free bro! Except my new SaaS calendar app, of course"), pirate Factorio, pay for kagi. A real eclectic bunch.