Filters to block and remove copycat-websites from DuckDuckGo, Google and other
github.com
github.com
Now that I see the HUGE number of copycat sites in the stackoverflow_copycats.txt file, I am beginning to understand what's going on.
Thanks!
They however have strong incentives to invest all available resources into their own survival, prolong their own life and/or earn more money.
Depending on how the managment is in turn incentivised they will typically also prefer short term success over decisions that would make more sense long term — e.g. short term a search engine saves money by not spending a ton of money on quality when they are in a monopolist position, while in the long term it fould bite them.
People accept help when you give it. This why quite embarrassingly one third of GoFundMe campaigns are to fund medical expenses that wouldn't even have happened with a decent healthcare system. Don't come to me with this argument when literally millions of people are resorting to begging on GoFundMe.
If a child goes home with a coat and comes back without it, they probably have bills and a financial situation at home so dire that they sold the coat.
How is a family of four making $25k/yr supposed to deal with a $120k medical bill when they lack health insurance because this gig economy and the 3 shift jobs they work don't provide it?
Does the rate of ad-clicking on the results page increase if most of the "natural" results are crap? :-(
For an Alphabet company, they sure don't know where to put the apostrophe
Maybe the problem is just that there are more of those now.
I wonder how many players on the page generation side? The economics of it must be marginal, I guess
Copycat sites don't seem to care anymore.
I don't believe there are an overwhelming number for Google et al to deal with as it's often the same names topping search results that such filters can remove through semi-manual user action.
While leads to the conclusion - Google don't care about duplicate content any more.
For a while (maybe around 2012 - 2017) or something it felt like it was almost the rule that if you found a really useful question on Stack Overflow it would always be marked as llw quality.
Eventually I guess they were pruned and that might explain a bit of why they rose.
They annoyed mee too though as they often mixed together unrelated questions on the same page and get hits for very specific queries that are unrelated.
But seriously, Google doesn't need to make anything besides bringing back the option to hide certain domains from the results forever. Even if they don't analyze what domains people are hiding, it would dramatically improve the usability.
https://raw.githubusercontent.com/quenhus/uBlock-Origin-dev-...
What's the worse that could happen? It seems like ublock is already to treat filter lists as semi-untrusted. There's not much it can do other than block stuff.
Then maybe optionally “follow” other people’s favorites as part of your own results. Imagine someone crawling Twitter or SoF and you just “follow” their index and it gets merged with your results and disable/enable them for specific searches.
I’ve been thinking about it a lot lately because I often am trying to remember something that helped me or I wanted to remember. But I can’t find it in my FF history. And Bookmarks feel too clunky.
Where as the “moderator of a filter” is followed on the basis of their output, the filter itself.
I don’t disagree with your statement, but I don’t find it a compelling counterargument to the suggested solution.
That said I'm amazed they are still showing up at the top of google search. My understanding was that that kind of behavior (which I think at least some other people do too) combined with the fact that they are just copying another much higher page ranked website would mean that they are highly unlikely to rank above the relevant stack overflow article that they are duping. So what is happening here?
They did fix this at one point in time by figuring out which site posted the content first and penalizing the copycats, but it appears the fix is once again broken.
I just don’t know how they are managing to get indexed before the big name established sites. Perhaps they are succeeding on some small percentage and that is what we are seeing?
Perhaps they have an additional trick to make it look like they posted the content first, perhaps internal links or something.
There's even a cottage industry around gaming these signals. See SerpClix and the like.
Edit: I have a friend who works there, but not as an engineer haha I'm pretty sure I've told him my woes with pinterest. My wife loves pinterest though. It allows her to come up with amazing design ideas and art ideas.
{google:baseURL}search?q=%s+-site:pinterest.*It is a browser extension and I haven't looked too deeply into it so if that's important to you perhaps have a browse over their repo etc before installing.
Quick edit: I know the domain is ABP but ublock origin picks it up.
For older versions of uBO, you can already use the old way:
abp:subscribe?location=[...]In my case, the issue is that GitHub doesn't allow the apb|ubo protocol in links. However, no problem, I can use the method with "subscribe.adblockplus.org". Yet, subscribing from contextual menu is a great feature.
https://raw.githubusercontent.com/quenhus/uBlock-Origin-dev-...
Adguard picks up 588 rules.
> Specific to dev websites like StackOverflow or GitHub.
Before I noticed that, I had searched for pinterest and found nothing. Even marking the HN title with "dev" would be good.
If this were my list I'd add w3schools because to me, it's low quality, especially compared to mozilla.
So it's working as intended and blocking low effort spam sites
1. What is the criteria for a github copycat?
2. What is the process for having a website removed?
I ask #1 because many businesses use their own git-hosting solutions to host their code but also use github as a mirror. It would be very easy for a competitor to get the website of a rival business listed if the criteria is not strict and specific enough. I recommend never blocking the entire domain unless the website is a repeat offender (to avoid mistakenly harming innocent businesses and the liability that it may cause).
I ask #2 because many of these domains will likely be registered by honest people once the scammers are finished with them. There needs to be a way for the new owners to get their new domains delisted.
For #1 it seems to mostly be a list of SEO gaming sites that I've personally found to be supremely irritating and deserve to be in the list. They basically just mirror stack overflow and GitHub issues and provide obscured links back to the original source to make sure people stay on their site. You can peruse the list in the code yourself. It's just a text file
You still need a browser extension adblocker (uBlock, AdGuard) that can modify the contents of any webpage for this to be effective.
Over the past year, I've noticed that quite a few repos that I used to track have disappeared. I keep a local bookmarks list now because if a "starred" project is removed or DMCA'd, Github does not tell you about it and they remove any mention of the repo from the "starred" list.
Internet Archive or Archive Today.