Google Search Results Plagued with spam “.it” domains
community.cloudflare.com
community.cloudflare.com
I thought that this was a massive nono from Googles side, has something changed?
This changed with the Google "machine learning" days, where you no longer have humans at the helm laying down explicit rules, so no more "change the world" updates, you can only slightly nudge the parameters towards what you want, meaning the same old tricks keep being effective for far too long.
That's just what the scheduled "core update" days are now: https://developers.google.com/search/blog/2022/05/may-2022-c...
A lot of updates are targeted at specific problems such as low quality product reviews but there are still broader updates taking place.
10 years I gave up on a large project where I rehosted and organized dead Usenet forum content because Google's dupe-penalty detector was too good and too aggressive for content that you could barely find beyond a six-year-old cache hit where the origin website was long gone.
Meanwhile these Stack Overflow scrapers are just `<html>{copy-and-paste}</html>` and the same domains are still alive despite years of cloning.
Looks like it's time to boot my project back up.
https://www.searchenginejournal.com/ranking-factors/google-a...
It blocks copycats and hide them from multiple search engines. You may also use the list with uBlacklist.
* the identical text copied from stack overflow should be easily identifiable
* volunteers put together a list of these sites themselves
it should be obvious to Google apoligists that Google is either negligent or intentionally allowing these sites in their search. I'm sick of hearing about how "the world is different" and it's an "arms race" between spam sites and google. Bullshit.
Google starts matching content from SO => Spammers start tweaking the text slightly => google implements some expensive similarity score to down rank copy cat sites => spammers use more complex scrambling=> ...
> volunteers put together a list of these sites themselves
These lists only work because they're used by a tiny minority of people. If Google were to do this the spammers would start switching domains more quickly (or find some other workaround).
I'm no Google apologist but I think you're underestimating how hard search ranking is when spammers are actively trying to game the system.
That's what ML is perfect at detecting, which is Google's forte.
Some of these sites have been returned as top results for a while, so are you suggesting that Google just gave up because spammers would be able to evade them with an update?
You underestimate the resources google has at its disposal.
They simply don’t care because there is no real competition to worry,even with this spam you are still likely to use google, so why would profit motivated company bother ?
At the very least they're being deliberately neglectful because they don't feel the bad experience harms their revenue because there's no other substantial competitor so they can abuse their monopoly status.
I guess they may just not care enough about software developers and figure we're mostly using ad blockers so its wasted effort and we'll develop blocklists ourselves. With no monetary value that they can assign to the ill will that it engenders they figure it must not matter so they don't bother. Pissing off a large chunk of the entire IT community via obvious neglect seems like a poor move to me, but then I've never felt that I'm cut out for management.
It feels like economy-wide that decision makers in corporations and governments have just arrived at the conclusion that there's no money / no point in trying to stop scammers (and there might be an actual cost to revenue of doing so). It won't goose their quarterly numbers and might hurt them so its better to allow it.
But the crazy part is that, for example - Ahrefs says that StackOverflow has "Organic traffic" in the range of 22 million per month. A lot of these copycat sites, at least the ones I saw - have a traffic range anywhere from 10k to 500k per month.
I mean, it's pretty insane just how well such sites can rank in Google, and you bet those copycats are making absolute bank from ads even if the majority of developers immediately close the site.
There's a lot going on with Google Search these days, a lot of people are complaining that sites that scrape content can easily rank really well for long-tail keywords. One case in particular, a site will scrape Google to collect "featured snippets" and "people also ask" - then combined anywhere from 20 to 40 of these answers and publish them as a blog post.
None of the words are changed, all questions/answers worded exactly the same. And Google puts these sites on page 1.
What a joke.
I bet the majority of developers block ads
Would they just move to creating and using new domains with the same content as soon as traffic to the old becomes drops? (What looks like the spammers in the original post are doing)
But something does need to be done to these sites.
This is a decades-old spammer trick. Google used to not rank brand new domains very high for this reason.
It's hard not to think that the only reason Google abandoned most of its old site ranking heuristics was that they were filtering out too many sites with lots of Google ads. The spam sites now infesting Google's first-page results don't look very different from the spam sites I saw back in the early 2000's. (There's more JavaScript, but modern search spiders run every page in a VM before reading the DOM, so that doesn't fool anyone.)
I don’t think this is a cultural issue, I fail to see how this can be considered value add by anyone.
yeah i've been hitting a ton of those lately.
It is simple. Google is making more money from copycat sites then from original content...
Actually forcing a search engine back into the reliable index of valuable sources would be great.
Imagine you to a white list approach to a search engine where a human or AI does an approval first.
Just go to SO and use its search bar. It's actually quite good.
I mean, you know that's where you'll want to find the answer anyway - not some random corporate webpage or ad-infested splog. Why not cut out the middle man?
Only if that fails do I bother with Google.
I think a lot of others formed their opinion (myself muchly included) about this from sites where the search bar was a joke played on people.
Edit: let me upgrade that 'fine' to 'great', now that I think about it it was actually better than a google search which was not my previous experience.
I've found some fantastic articles out there, yes SO is a fantastic resource but there is an entire internet out there :-)
Maybe that means we should be searching in yahoo rather than google.
Surely Google can spare an engineer or two to do a deep dive into the way any one of these spam sites manages to get itself to the first page of Google, work out their scheme, and fix the algorithm? This problem isn't exactly hard to reproduce!
For a while now Google has suggested that the best way to rank well is to have human readable content and focus on user experience. At the same time, natural language generation has come leaps and bounds, to the point where sometimes even I, a human, can't tell if an article has been spun by a bot or not.
So if Google starts ranking human readable content, and robots can now produce human readable content, what is the next ranking signal they can use to differentiate spam from humans? Are we going to end up with "Verified Websites" ala verified Twitter handles?
A huge portion of the web at this point is just bots communicating with eachother, and legitimate business systems having to process bots participation on the internet. I imagine the portion of the web that Google crawls that is legitimate versus that which is bot generated would surely be majority bots, just because of how fast they can generate content. One thing they can't do as easily though is register domains, so it may be one of the better points of defense.
If it can be effectively blacklisted, then Google is dropping the ball. This isn’t difficult algorithm foo failure.
I don’t agree with your sentences, but I do agree with your point.
The problem exists outside (Google-controlled) web: with (not fully Google-controlled) email, too.
Around 2020 I did a per-tld checks on wanted/unwanted messages (ham and spam). With thousands of messages sent from .xyz domains (envelope sender host or PTR record of sending host; I ignored the From header) there wasn't a single legit message. 100% SPAM.
Anyone trying to infect others with Trojans and viruses just need to check user agents or use dynamic redirect URLs, and suddenly this clearly illegal activity becomes black magic that is way beyond the comprehension of the folks at Cloudflare.
Cloudflare is basically making the shittiest parts of the Internet safe for scammers and spammers, and this is just one example.
If that's not bad enough, they're trying like crazy to become a monopoly. If this what they do now, imagine how bad it'll be when they control even more and feel even more immune to making money from scammers.
(1) Hosting is providing services on the Internet without which a site would not function. Providing DNS is hosting. Providing proxy is hosting. Providing email is hosting. Don't fall for Cloudflare's "we don't host" bullshit.
For example, SO copycats are legitimate in that they respect the license and otherwise just serve the content to whoever sends them an HTTP request. As far as I know they don't spam links to their domain anywhere. They are low-quality and of dubious utility for sure, but I'd rather not make the Internet a place where you need to prove quality & utility to someone to be able to host an HTTP server.
The real problem is that a dumbass like Google comes along, sees this and decides that it should rank higher than the source content.
But Google maps better not drive me to someone's yard when I ask to navigate to a nearby mechanic.
If it did, it would be hard to blame anyone but Google.
It's really tricky to enforce open licenses on this scale as it's each contributor that licenses their content rather than the platform host.
Do they? SO contributions are under CC BY-SA. Haven't seen copycats providing attribution let alone specifying that the content is under the same license.
Somebody upthread suggested it was just the use of Google ads, which I suppose is possible, but somehow it seems unlikely. Google sure does love money but they also need to be considered a good search engine, and I'd expect them to be at least a little wary about things like that.
Is there something else I'm missing?
This is how it gets done and Google used to be brutal about crushing it, somewhere along the way they seem to have given up on being so brutal.
However, I don't want Cloudflare to preventatively police what is and isn't a bad website. When these scam sites go live, they can quite easily contain real content (say, a blog, with articles written by AI good enough not to be immediately obvious) and then change into malware on a schedule.
Cloudflare can't see what code customers run on the backend and that's probably a good thing. They're already holding too much power over the internet and requiring the backend to be transparent would only make them more in control of the web.
Any registrar hosts thousands if not millions of spam sites because every single one of the billion registrars have DNS set up in some way.
Despite being almost exclusively used for spam and amateur projects, the .TK TLD barely shows up in Google. Spam sites are a symptom of other services linking to them and making them worth the investment. If Google, Bing, Qwant and Yandex weren't falling for the SEO scams these scammers use, we wouldn't have this problem.
Hosters have some immunity by design, and that's very much a good thing. They have to respond to abuse complaints, but they're not responsible for filtering out all of their customers. Requiring them to do so is exactly what the EU is trying to force upon the internet, which is terrible for online freedom.
This is clearly Google's issue.
Every time a scammer puts up a web site trying to sell counterfeit goods, the company which sells the real goods should file a lawsuit? One for each scammy web site, perhaps? Because Cloudflare shouldn't be expected to do anything at all, until they're compelled to do so by a court?
I don't think you're thinking this through.
> I don't think you're thinking this through.
love the sassines tho
- abcedasdfff.io = €59.29/year
- abcedasdfff.tw = €25.20/year
- abcedasdfff.nz = €25.40/year
- abcedasdfff.mx = €48.28/year
Most of them appear to be €10-20/yr, but it's certainly not uncommon to see them go for €25 or higher. Note: EUR and USD are roughly at parity so I don't think it's really necessary to do a conversion.
Why they decided to ".xyz the TLD", I don't know. ¯\_(ツ)_/¯
Now I get something from McAffee Pratners(sic) every other day warning my computer is about to expire. Back in May I kept winning things from Home Depot and Lowes; and gmail would categorize it as "forums".
No idea if its related, just odd.
I don't know if it's related either.
I had the Slack google drive integration and I needed to mute it because it was couple doc invites every few hours.
The added benefit is I don't get any tech calls for help from my parents who also don't end up clicking random spam and wondering why bad things are happening
I can't help plug kagi.com, which has the amazing feature of grouping SEO'd stuff like recommendation lists together, so a thing that's contextually useful is still available but without polluting the other contexts.
This is not helping to organize the world's knowledge.
And Google is not about organizing world's knowledge but creeping on people for YoY financial results.
> Google's mission is to organize the world's information and make it universally accessible and useful.
Oh they stopped doing that long long ago...
I get your point though about the multiple results for something where there clearly is no authoritative answer.
It's not a firewall.
Seems like you're describing cloaking (https://developers.google.com/search/docs/advanced/guideline...), one of the oldest SEO tricks, and you can imagine that search engines started defeating it on Day 2 of crawling the web.
to register an .it you must prove you are a person or a business working or residing in one of the EU member states and need to provide the ID of a person who's gonna be listed as admin-c of the domain.
Yes, you do!
of course it worked.
you just committed a crime.
you can fake your id everywhere in the World, it is a crime everywhere in the world and if something happens doesn't mean you won't get caught.
you can drive a stolen car, it will work.
> yes there is a field in regstritation where you should enter a "identity card id"
so it is required! you simply ignored it, lied and broke the law.
your criminal behaviour doesn't imply laws do not exist.
if you tried to buy an insurance policy with that fake ID, you would be in troubles now.
https://en.m.wikipedia.org/wiki/1998_Cavalese_cable_car_cras...
>if you tried to buy an insurance policy with that fake ID, you would be in troubles now.
But this is more a "are you 13 or older"-style of "crime".
tl;dr: I managed to find the servers behind it, most likely anybody who are still affected can do the same thing I did pretty easily. We also followed the money, which is a tad more work.
Sounds about as trustworthy to me as a .tk domain.
I think in recent time, .icu and .xyz have been the most problematic, to the point where you to this day probably don't want to host a mail server on those domains.
The same with cloud providers. A fairly significant amount of sketchy websites seem to be hosted on cheap cloud providers with weak rules enforcement. I've taken to blocking all of Alibaba's IP ranges from my search engine crawler, the signal to noise from those sites were so bad it just wasn't worth looking for legit content.
Junk copypasta and news-squatting (posting regularly about the same thing with no additional data) is a decade(s) old problem. Crowd sourcing and verifying junk domains could be a weekend project.
But nothing.
Locked?
(Have no idea how reputable that data is, but it seems about right to me. In 2016 there were 3.6 billion Internet users. Now there are 5.3 billion.)
I meant that google dropping search quality for English has not much to do with growth users in the last few years as that growth has largely been non english
Warning, do not click on those links as you will get your PC infected.
Their abuse form is getting abused too. It sends an email to site operator and the server hosting company in single submit so its getting abused. It not even have a captcha.
The only thing they'll do is forward the complaint to the user. Leaving you with no recourse other than to take legal action before Cloudflare will lift a finger.
Unless there's CSAM, of course.
They store the content of the website on their drive to serve to visitors. Whatever processes lie in the backend of whatever website to fetch up-to-date content from an upstream source is not my concern. They are NOT a neutral ISP, they are providing a service to their customer which includes hosting (doesn't matter if it's temporary hosting because they expire files). From our point of view, it is their IP addresses that are hosting the website. They have all the responsibilities a traditional hoster has, no motter how they try to frame this debate.
A trademark dispute is a civil issue between two parties. We have legal systems to solve these. Cloudflare should ensure that their customers get timely notification of complaints, and that’s pretty much it.
If a website lies to Google itself, I believe the only way to solve it is by reporting the search result as spam or Google contracts people to somehow visit all billions of web pages (again the same problem – from different IP ranges) to verify it as a legit page.
I would like to know how Google currently handles it and probably how it could be improved
I'm not sure on what grounds could someone sue for crawling from random, unaffiliated addresses as long as the crawling isn't causing a denial of service (they can always check robots.txt using the main IP then use that to throttle crawling from random IPs as to remain compliant).
> The reason being that some pay-walled news article websites won't be indexed properly, as the “unofficial IP-ed” Googlebot will not get the paywalled content.
Good riddance? That would be a welcome change.
They can see all the domains that have served a given snippet.
They also have history to identify where each snippet was first seen.
If SO has a lot of traffic and a good reputation, and if the same snippet is found first at SO and then later at bunch of newly created, low volume, low reputation domains, then show the SO result and not the others.
Just blanket block the lot with the following uBlock Origin filter:
google.*##.g:has(a[href*=".it"][href$=".html"])
Google ain't going to fix itself ;)now s/\.it/every TLD/ and you solved domain spam forever.
/s
You might not know that 99.99% of .it domains with urls ending up in .html are completely legit, including some official government one.
I could block all .it sites on my network and I’d likely never even notice.
¯\_(ツ)_/¯
the problem is not .it domains, it's clearly stated in the linked postA large number of spam pages are indexed when searching by our product name. It’s very similar to Japanese Keyword hack, but the difference is that our site is not hacked
so it's definitely an indexing issue, those .it domains are being indexed for the Japanese word hack for some reason, it's not that .it domains are particularly spammy per se.
Your "solution" would filter the vast minority of the abusers at the cost of banning an entire TLD, not much different than turning off the internet connection entirely.
Most of the spam on the internet comes from .com domains though, even more so because registering a .com domain is much easier than getting an .it
Are you willing to ban .com too?
Again, we’re talking about client-side filtering. The original comment about blocking .it domains was talking about a uBlock Origin rule. No one’s talking about blocking .it domains from the web.
Yes, as an American, I could block all .it domains on my end and my web experience likely wouldn’t change at all. I rarely, if ever, need to visit .it domains. So maybe I will.
This is a personal solution to an extremely disruptive and long standing problem, and only affects those who choose to employ it. It's not hurting anyone.
There are 8 thousands towns in Italy, each with their own .it website.
No network connections are blocked...
This and other information is then used to filter out various types of visitor.
In this case, requests claiming to be a Google Search crawler will receive a boring page with lots of text that it can index and use as search results.
Most browsers' devtools let you change your user-agent string, and a listing of the ones used by Google crawlers is publicly available. Not saying that you should, but you could check this out for yourself... entirely at your own risk of course :)
https://en.wikipedia.org/wiki/User_agent
https://developers.google.com/search/docs/advanced/crawling/...
local-zone: "it" always_nxdomain
to NXDOMAIN all requests for the .it TLD and protect non browser devices. I use this method to stay off sanctioned country TLD's and to remove the cheap/free spammy domains and TLD's that often contain more malware than anything useful.