I'm shadow banned by DuckDuckGo and Bing
daverupert.com
daverupert.com
The automated tool says it's in violation of some unnamed rule, but I can't figure out which. There's zero SEO, tracking, or ads, and the content is educational and G-rated.
All the other guides index just fine.
I asked for a review and they came back with the same ambiguous message. Eventually I just gave up.
Recently I split the C guide in two. I'll have to check to see if that made any difference.
But it left a bad taste, and now I don't trust Bing or DDG to provide complete results. Google's overrun with spam, but at least my stuff actually shows up on Startpage.
But I'm pretty sure mine is the greatest that you can't pay money for. ;)
If I search "Beej" on Google without Safesearch enabled, I get 14.3M results, if I turn on the Safesearch filter, it still returns 14.3M results.
If I repeat the same experiment with "blowjob", it's 1.5B results vs 23M.
If I search for "Beej" on Bing with SafeSearch Off, I get 2,840,000 results, while with Safesearch on Strict, I get 2,800,000 results. I couldn't search for "blowjob" at all with Safesearch on Strict.
But MASH is getting a little too far removed these days. :)
EDIT: but maybe they did, ever since voice assistants became a thing?
But they index everything else on my site and don't prude out over that...
(Yes, not adding much insightful conversation. I don’t care if I get downvoted.)
I always wondered about this. What exactly is a backlink, and why should I need one?
https://duckduckgo.com/?q=c+guide+stdalign&t=ffab&ia=web
brings up your guide as the 6th result.
Because if you put the headlines (in quotes) from two of his recent articles into Bring, e.g. either "Megan Smith explaining the General Magic prototyping process" or "Denialists, Alarmists, and Doomists", both point as their first result to a URL starting with "https://www.scien.cx" which seems to be the spam site with a copy of each article. (The URL isn't loading right now, however, when I try to visit.)
How to fix it really depends on what techniques they're using to mirror your site, of which there are many.
Example search and resulting URL:
https://www.bing.com/search?q=%22Megan+Smith+explaining+the+...
https://www.scien.cx/2022/12/25/megan-smith-explaining-the-g...
Compare with Google getting it right:
https://www.google.com/search?q=%22Megan+Smith+explaining+th...
https://daverupert.com/2022/12/megan-smith-general-magic-pro...
If that happens I'd use language about the exclusive "public display" right that you have over your work.
I'm not aware of this claim being used for a dmca, but I'd like to see how such a claim turned out.
In the first case, I sent a DMCA to Google & Bing as well as to Cloudflare. Cloudflare responds by giving the name of the actual host, and I sent another DMCA to that host (they were US based, otherwise YMMV). The content was delisted (not the site, even though it was made up entirely of verbatim scraped content) from search engines and from the site.
Bottom line is you can send a DMCA notice to search engines and it appears to be effective. Actually, in case search engines demote sites like this in some way, I would send the DMCA notice to search engines _first_, because if the content gets removed from the original site they may not be able to verify the duplicate content.
Have we gotten to the point where websites (and their content) need to be verified like Twitter, Instagram, Facebook, and TikTok do for personal accounts?
If so, will search engines be the ones verifying - using this as a new revenue scheme (with the dangers inherent in this... ie; pay to be listed or ranked higher)?
Bing is the problem. It's broken.
> "Thank you for your patience during our investigation. After further review, it appears that your site did not meet the standards set by Bing to remain indexed the last time it was crawled. To ensure that this was not a false flag, I also escalated the issue to our Product Team and they manually reviewed your site and confirmed that it is in violation of our Webmaster Guidelines detailed here:
https://www.bing.com/webmaster/help/webmaster-guidelines-30f....
We are not able to provide specifics for these types of issues but we recommend that you review our Webmaster Guidelines, especially the section Things to Avoid, and thoroughly check your site for any deliberately or accidentally employed SEO techniques that may have adversely affected your standing in Bing and Bing-powered search results."
Before snarking, please check that link and the long lists of things - I did not find to find my website https://linmob.net to be offending their "things to avoid list".
That was a reply to my first ticket requesting re-indexation, later tickets only got what I would call "non-replies".
The more sophisticated and popular the copycat site is (scraping from a distributed network, stripping most HTML tags, etc.), the harder it becomes, and the only thing is to contact the search and hope they can manually mark your domain as the authoritative one. Your success may vary according to your popularity/importance.
That would almost certainly be regarded as a "doorway page"[1], resulting in you getting manually ranked downward or even deindexed entirely.
[1] https://developers.google.com/search/docs/essentials/spam-po...
I'm talking about, if content on legitimatesite.com includes JavaScript that detects if it's being loaded on any other domain, then erase the entire article's HTML from the DOM.
Obviously this is easily defeated by stripping out JavaScript, so it's useful only for very primitive mirroring.
It feels like the Internet is a more hostile place than ever for small-time websites. You get squeezed from below by wily criminals, and crushed from above by careless megacorps who want to filter out anything that doesn't make them money.
The spammers would use that offensively to destroy the original content creators with that system
I'm still surprised no one else seems to be offering this feature.
https://chrome.google.com/webstore/detail/ublacklist/pncfbmi...
The problem is that the two work hand-in-hand, thanks to the advertising driven search model, and the search engines owning the main advertising platforms.
It should be easy for search engines to identify an original site from the SEO spammer rip-offs - the original site is going to have no adverts (or certainly fewer) while the SEO spammer copies are going to be covered in adverts. The problem is that the search engines have no incentive to do so, in fact if anything they have the incentive to send people to the sites with more adverts.
And of course the whole problem has been created by the search engines in the first place - there would be no point in SEO spammers making advert-laden ripoff sites if it wasn't to rake in advertising revenue.
Edit for clarity.
If you use Bing webmaster tools (a logged-in account for the use of the domain owner/content creator) and you can see indications that Bing indexed the content and no indications that errors preventing it from being eligible to show then it is certainly at least closer to a shadow ban than I originally thought.
Still, any reference to a "feed" entirely misses the point unless Bing is the one also serving that feed. I can't see any evidence that Bing displays a feed to the poster.
Regardless of whether you were notified, if you can see that you are banned, it is by definition not a shadow ban.
Shadow ban is explicity "pretend to user they are not banned, but don't show it to everyone else"
Like say getting shadow banned on reddit or HN, you will see your stuff when you're logged in, but anonymous or other people wont.
Search equivalent would be you getting your own site when you're searching but nobody else does
What shall we call that instead, "bubble ban", since it is like a shadow ban for specific bubbles and for everyone else it is as if that bubble never existed on the site?
Bubble ban is also poor verbiage because not appearing in anyone's feed or search results is the default condition not inherently a punishment or redaction.
A search is inherently a selection process and it's perfectly valid to say some content isn't fit to appear anywhere in a listing.
I like reddits choice of "quarantined"
It isn't the same at all, since Reddit tells you about it and it still shows up in search etc. Shadow quarantined, yeah that works. The word "shadow" comes from not telling the user that anything is different, and Twitter doesn't tell you when this happens, it just delists your posts from everywhere except your followers.
Anyway, it is extremely disingenuous to say that Twitter doesn't shadow ban.
Consider this scenario: Person A is shadow moderated by Twitter, has a friend B. A replies to B's tweet, but his friend doesn't follow him, they just talk. And now B will never see A's tweet, and A has no idea that B can't see it, and neither do B, both thinks that they can see each other. What is this if not "shadow banned"? Twitter makes these people post things thinking it will be seen, but it wont, wasting their time and potentially hurting their mental health since nobody responds.
For all intents and purposes this is "shadow banning", when you hurt people like this but say that you absolutely don't shadow ban you are so dishonest that I'd still call it a lie.
Deindexed/delisted. Same as we called it before "shadow ban" was widely known as a term.
Twitter does that, it doesn't tell you that what you post wont be reached by the people you are posting to. Most people don't have many followers, they just reply to tweets and those replies will show up for the original tweeters. Twitter shadow banning you means that the items you post no longer shows up as responses. Sure the small subset that follows you can still see them, but 99.99999% of twitter wont see it, so it is a 99.99999% of a shadow ban.
If they told you that any of this happened anywhere it wouldn't be a shadow ban.
It does not. Delisting = removed from list. Deindexed: removed from index.
Banning implies denying access. There's an important distinction there with Twitter (user authenticates and publishes through Twitter) vs Google Search (indexes public sites). If Google silently stops showing your Google Ads and doesn't inform you, that would be more appropriately described as "shadow banning".
There's no "user" or "account" in a web search engine so the term doesn't really fit.
Imo a logical interpretation of "shadow ban" would be when you are banned but they didn't tell you they banned you, and regular "ban" is when they tell you you were banned. It makes enough sense that people don't think they need to look it up to confirm.
edit: funny enough, I did double-check the wikipedia page to make sure my understanding was correct, but upon reading further it does acknowledge the expanding of the definition: https://en.wikipedia.org/wiki/Shadow_banning
When a user is shadow banned they are normally not told they are banned and are still able to access and perform functions. I think this secrecy is where the word "shadow" comes in. The user is in the dark about the ban...
Banning is just stopping the service for you, wether they tell you actively or not depends ob the service. No search service is actively informing you about the usage of your data, neigther are they telling you they stopped servicing you.
That doesn't sound as cool and victimey though.
Only because this is a search engine do I say this solidly isn't a shadow ban.
Ghost Banned
or
Ghostdexed
Scrolling further, I don’t seem to find my own site either… https://donatstudios.com
I’ve added my site into Bing webmaster tools, we’ll see if it helps I guess.
But I keep thinking that this is a limitation of MIT. It's written with the purest of good intentions but without any way to prevent bad actors from exploiting that.
I wonder if there's a better kind of license that provides a close level of freedom but would prevent some of the most obvious exploitations (e.g. packaging and selling the code, or reposting on a site with advertising)?
There's copyleft licensing like LGPL but perhaps that goes too far in the other direction and besides, it seems to be very unpopular amongst web developers so I'm afraid if I release any LGPL code it won't get used much or attract contributors.
Is there a happy medium between these?
MIT will allow that, Apache/BSD/Mozilla licenses will allow that, GPL and LGPL will allow that, Creative commons CC-BY and CC-BY-SA will allow that - the only difference is extra conditions e.g. GPL will require the seller/redistributor of the code to keep the same license, CC-BY will require leaving attribution to the author, etc, but all of them will allow someone else to redistribute the code for commercial purposes.
It's probably not the reason, but it's worth noting that the author is using Amazon affiliate links in violation of Amazon and FTC rules because they're not disclosing the fact that they profit from purchases through their links.
Per Amazon:
>Anytime you share an affiliate link, it's important to disclose that to your audience... you must (1) include a legally compliant disclosure with your links and (2) identify yourself on your Site as an Amazon Associate with the language required by the Operating Agreement.
https://affiliate-program.amazon.com/help/node/topic/GHQNZAU...
Per FTC:
>As for where to place a disclosure, the guiding principle is that it has to be clear and conspicuous... Consumers should be able to notice the disclosure easily. They shouldn’t have to hunt for it.
https://www.ftc.gov/business-guidance/resources/ftcs-endorse...
But it still doesn't index the first volume...?
Thanks for the info.
[0] https://www.bing.com/webmasters/help/url-submission-62f2860b
[EDIT] I just published a new blog post "Bing and DuckDuckGo removed my business web site AGAIN" https://lapcatsoftware.com/articles/bing2.html
Sigh.
https://daverupert.com/atom.xml
First, he sends it with a "content-type: application/xml" header. In contrast to most sites that send it with "content-type: application/atom+xml". Which seems to have the nice effect that it renders in Firefox instead of opening the usual "What should Firefox do with this file?" popup.
Secondly, he provides this nice header text "Yahaha, you found me! This is my RSS feed.". It seems to be fetched via this part of the code:
<?xml-stylesheet href="/pretty-feed-v3.xsl" type="text/xsl"?>
Pretty nice. Are those best practices? Or will "content-type: application/xml" mess with users who have a native feed reader installed and expect the reader to kick in when they click on a feed url?
"An elegant weapon for a more... civilized age."
One of the ideas people had back then was having a product catalog in XML that your tool would give you, and that you could upload to your website and display as a nice website via XSLT.
It was such a fascinating technology when I learned about it, but I’m not sure if it ever saw any serious use? Maybe in enterprise?
But XSLT was not fun work with.
It’s been a few months since and my website is indeed back in the search results so I advice whoever is having this problem to reach out to Bing.
It says your IP doesn't direct to your site. I wonder if that's the problem.
While it is more likely the poster will get a few people to check on the situation and naively drive up page rank... a personal site is just a rounding error for traffic in a long-tail distribution known as the modern web.
Most search engines will correlate user-side telemetry traffic against crawler and web stats. i.e. if the bots tend to prefer your site for abnormal reasons, the ranking algorithm may blacklist a signature, domain, and IP sets for several weeks as punishment.
Note too, it is still common for a human employee to manually check a suddenly popular site that pops up out of obscurity. i.e. this catches the more sophisticated cheats, and may have legal repercussions in severe cases.
In summary, if you mess with modern search engines, than expect the ban hammer to fall eventually. ;)
Why, Bing does not totally ignore bad links.
The good news, all sites I've seen affected by this recover after a few weeks.
I'll say it again: It comes down to how you treat people. Treat others the way you want to be treated. No one wants to be shadowbanned and we can all agree it is a decidedly cowardly and cruel thing to do.
And you can't use "quality" or anything short of being coerced as an excuse. Techniques and technologies to moderate people without shadowmodding at scale are mot just there but very well established. A site for technologists has no excuse to shadowmod other than elitism amorality.
I have showdead turned on. I see fresh accounts that are automatically dead on each comment which are legit contributions, probably because they're using Tor or a widely abused VPN; that's the only common miscarriage of HN moderation I regularly see, and I vouch for house comments. Those accounts should be in the clear after a week or something like that. I have a couple comments I feel shouldn't be dead, but I can see how others would feel differently, and I have I believe 3 dead comments out of >3000 (many of which expressed views others vocally disagreed with, and I generally feel my views are not particularly popular on HN). But most of the dead comments I see are obviously harmful to discussion. The last time I saw hate speech from a banned HN account - was earlier today. What is it I'm missing here?
It's all well and good to say, treat others as you'd like to be treated. But I don't want to be harassed either. So I forgo harassing people sure. But what's to be done about the people harassing me?
Are you perhaps unaware of a phenomenon called the paradox of tolerance where, if you extend universal tolerance to everyone, including those who use their speech to silence others (through threats, harassment, shouting over people, poisoning the well, etc), you still end up with a forum in which not everyone can share their ideas?
Maybe change this? Simply add:
User-agent: * Disallow:
To allow all crawlers to the site.
After a bit of poking around, it would appear Bing may need an allow block to crawl. I don't know what DDG does, but the author's site effectively has nothing in the robots.txt file, other than a commented out Disallow block. From doing this before in the past, I suggest include the following:
User-agent: *
Allow: /The point of a shadow ban is for the banned user to not notice.
This sounds like a regular old ban/blocklist.
font-family: system-ui, 'Noto Emoji', sans-serif;I recently discovered that one of my services is on stage MS email naughty list, and found out it's because MS uses SpamCop.
I contacted SpamCop and it was very responsive. Unfortunately, the solution SpamCop suggested was to move the entire project to a different provider.
Perhaps some server-side filtering going on, blocking the bingbot user-agent?