I never had reason to complain about Google's algorithm until I became a content creator. Now, I wonder what other great web pages I'm missing out on.
I never had reason to complain about Google's algorithm until I became a content creator. Now, I wonder what other great web pages I'm missing out on.
I encounter this several times a week it seems (google, ddg, doesn't seem to matter). Usually when trying to find the actual source of something, like a scientific study, the top hits are all news sites, often whith articles which look like copies of each other. I'm assuming they're ranked by number of visits and popularity in general of the site, but it makes it really hard to drill down to the source of something. To the point I usually give up. I am going to try the '-news' thing mentioned by another commenter though, hope that helps. Or are there any other tips?
All sorts of maximum ad-cram crappy sites (including ones that try to pretend to be a forum, with mailing list posts as "users") that maximize SEO gaming techniques dominate the results for 20+ pages.
http://marc.info, which is the best archive avalible atm, plain, simple no bullshit layout and is linked everywhere anytime somebody wants to reference a mailing list archive, nowhere to be found.
Speaking of which, does anyone know what's going on with Gmane? There hasn't been any updates since the ownership change.
A non-existant website won't get a search placing.
What's really annoying is that if you use Google trends, it will let you know the context of what you are looking for, and you can clarify. I really hate it when I'm trying to look for something on Google, and I know it's just using the wrong interpretation of a word.
Easy example: I do a fair amount of scripting in VBA. I often want to look up documentation for things in VBA. If I use VBA in my search term, I will get a bunch of 15 year old pages on Excel tips and tricks. If I search VBA MSDN, I will get a bunch of pages on VB.net or VB6. I specifically have to use VBA MSDN OFFICE ACCESS in order to get the correct results. DuckDuckGo does not have this problem, but I use date filtering a lot, where DDG falls apart.
Search engines aside, it's ridiculous that Microsoft allowed this situation to happen in the first place.
*I'm not sure if either of those terms is remotely accurate, but they seem to articulate what I'm trying to say.
Remember when we had all those spammy answer sites that dominated Google searches for almost any technical question?
Furthermore, VBA <> VB. I wish I could just tell Google to stop showing me answers for VB6. It would be fantastic if I could use VB. My problems would basically solve themselves, but I'm not using VB, I'm using VBA.
https://bynd.com/news-ideas/google-advanced-search-comprehen...
> VBA AROUND(20) VB6 | vba -vb6
would work?
vba sort array site:msdn.comIf you're looking for an actual paper, why not just use Google Scholar?
"We actually came up with a classifier to say, okay, IRS or Wikipedia or New York Times over this side, and the low quality sites over this side." -- Matt Cutts
----
> Singhal: And based on that, we basically formed some definition of what could be considered low quality. In addition, we launched the Chrome Site Blocker [allowing users to specify sites they wanted blocked from their search results] earlier , and we didn’t use that data in this change. However, we compared and it was 84 percent overlap [between sites downloaded by the Chrome blocker and downgraded by the update]. So that said that we were in the right direction.
> Wired.com: But how do you implement that algorithmically?
> Cutts: I think you look for signals that recreate that same intuition, that same experience that you have as an engineer and that users have. Whenever we look at the most blocked sites, it did match our intuition and experience, but the key is, you also have your experience of the sorts of sites that are going to be adding value for users versus not adding value for users. And we actually came up with a classifier to say, okay, IRS or Wikipedia or New York Times is over on this side, and the low-quality sites are over on this side. And you can really see mathematical reasons …
----
Cutts was using IRS/Wikipedia/NYT as stand-ins for "sources people trust", not "sources of high quality by some magical objective standard that Google decides contra their users". Ultimately, their classifier is validated against what users blocked as low quality sites, since there's no grounding for quality that doesn't come from users.
Hell, the quote you're pulling doesn't even make sense in the context of the exact example we're talking about! A lot of people would find a big newspaper like the NYT at the pinnacle of trustworthiness for factual content, more so than something like NBER or PLOS (where you can find the actual papers that NYT occasionally reports on, often not very well).
[1] https://www.wired.com/2011/03/the-panda-that-hates-farms/
It ends up being the same thing. Brand value being an overwhelming factor is precisely what the OP and everybody in this thread was complaining about in their SERPs.
EDIT: I think I misunderstood your question. I will probably pay that much for quality ad-free service. But of course it's a niche product.
https://chrome.google.com/webstore/detail/personal-blocklist...
3rd party addon for Firefox:
https://addons.mozilla.org/En-us/firefox/addon/hide-unwanted...
Another one for DDG in Firefox:
https://addons.mozilla.org/en-US/firefox/addon/ddg-hide-unwa...
I wish this was a feature for Google when logged in (although this is a better way I guess).
Not a perfect solution by any means, but it should get rid of news articles.
Google Web, Books, and Scholar, desktop versions, are better in this regard, as specific intervals can be specified. The problem here is that for the Web, there's a great deal of newer content masquerading as older. The problem appears in print as well, but not quite to the same extent.
?tbs=li:1
http://jwebnet.net/advancedgooglesearch.htmlhttps://stenevang.wordpress.com/2013/02/22/google-advanced-p...
Sci-Hub will turn up most of the results, LibGen and BookZZ many of the print sources.
At least for scientific articles, it usually helps if you're willing to spend enough time to familiarize yourself with the technical jargon. As a somewhat dated example, "Higgs particle evidence" will turn up lots of crappy new articles, but "CMS Higgs abundances" will immediately return high quality relevant results.
Of course, you have to figure out the right lingo int the first place, but usually you can mine arXiv/PubMed/etc for this kind of thing. Wikipædia usually has links to some kind of primary source.
I used to write for an amateur online magazine with product reviews and analysis of a certain industry. It wasn't big, but the content was good - not just in my opinion, I was told that many times by various readers. In ~2005 we routinely were #1 to #5 in searches related to the products we're covered. Today I often can't find new articles from that website even when I use the exact title as the query. The website is minimalist, fast and has no ads. The articles are extensive and well-conceived. And yet they are buried several pages deep under piles of single-paragraph stub "reviews" on more popular resources.
I assume the same thing happened/happens to a lot, maybe most small-scale content creators these days. Clickbait rules and content doesn't matter nearly as much as the number of incoming links. You can no longer rely on search engines to get you traffic and need to constantly self-promote on other websites (e.g. social networks).
http://firstmonday.org/article/view/833/742 (from 2001)
Google seems quite determined not to risk being seen as biased in that way.
* lack of https * lack of http/2.0 * lack of CDN utilization (aka lack of speed) * lack of Google backed advertising * lack of continuous updates * lack of inbound links * content not minimized * content not zipped
All contribute to a decreasing Google Page Rank (TM) over time.
But depending on the topic, "stale" may not mean much about quality at all. I've had a FAQ site about Tolkien's books online for at least 15 years now, and as it turns out he hasn't been publishing much lately (being dead and all). So apart from minor updates every few years when his son or other researchers publish new information from his drafts and notes, the site has been largely static for a decade or more. It's very good at its intended purpose, and it's hard to imagine what I could do to provide a steady stream of new content to indicate "freshness" without fundamentally changing the nature of the site.
I imagine that the same is true of a lot of sites out there, on a lot of topics that are far removed from the "breaking news" world.
It makes content creators miserable. Some of us just want to write good stuff and get traffic through Google, not build backlinks and be a SEO
Take this book:
https://www.amazon.com/Never-Split-Difference-Negotiating-De...
If I search for '"Never Split The Difference" review', I want to find, well, people's reviews. Note that the book has several ratings on Amazon - it is a popular book.
Yet I found only perhaps 1 "honest" review in the first 2 pages of Google's results. Everything else I find reads like a promotion for the book.
Looking at Fakespot, there is some evidence of light tampering with Amazon's reviews on the book.
The reason I Googled it? I've read a few chapters and am appalled at the book. It essentially is trying to boost its popularity by trashing what is taught in well respected negotiation programs at top universities. But while repeatedly trashing that education throughout the book, he continually advocates strategies that are also taught by the same programs he is trashing.
Given that he continually bashes the most famous book on the topic (Getting To Yes), I wanted to see if anyone has done an honest comparison between the two - pointing out the author's somewhat dishonest stance. And I can't find it in the early Google hits. I see it only in the 1 or 2 star reviews on Amazon.
To be honest, that's similar to my motivation for finishing it. I think this is the first book I decided not to stop reading just for the sole purpose of writing a lengthy rebuttal review.
I actually will not say that his techniques are bad/wrong/poor. But it's really shitty to keep trashing Ivy league MBA programs and then advocating the techniques taught by those programs. To date I've read 3 books from The Harvard Negotiation Project and another from, I think, Duke. For each chapter in this book, I want to highlight his techniques and then specify exactly where the same advice appears in the books he criticizes.
I think it actually would have been a great book if he wrote with more integrity. His book is easy to read and practice. If his techniques work, then he has a legitimate advantage over the other books, which are much more complex. It's a pity he made a potentially great book into a lousy one.
The downside is in the assumption that every difference is correct, but it's as dangerous as the assumption that only what is repeated is correct. Would be an interesting alternative at least.
Basically, only more mainstream sites have the right to present different views which leads you back to square one: Repeating information.
If you want to search for the history of pineapples,
you don't need 100 sites with the same info.
"Pineapples grow from a plant in the tropics.""Pineapples are grown in equatorial countries."
---
"Pineapples, when ripe, have a very sweet taste."
"Ripe pineapples have a high concentration of sugar."
---
"The etymology of the word for pineapple is 'excellent fruit' in the language of Tupi, although the english English translation is an aberration, with a stem of 'pine' due to it's resemblance to a pine cone."
"Most European languages name the fruit using the root word 'nanas', from the original native language where it was grown. Many languages have adapted the root word, although a few have a large degree of uncertainty, such as the origin of the reference to 'apple' in the English word."
---
It would need a fairly decent AI to recognise the commonality of information in these examples.
AI of course can get you much farther still, but you definitely don't need something that powerful for a convincing upgrade from baseline.
The problem is that PageRank is no longer delivering on these.
Please buy ads." - Google
For Google it's a minus sign. (sometimes it's the word NOT)
E.g. "onions -news" (without the quotes) will search for the word onion and not return any sites that have the word news in them. This can have false positives so youd need to play around with the words.
False positives because non news sites may have the word news and news sites might not have that word.
On the other hand, back in the 90s it was a practical necessity to use extended search syntax with services like AltaVista to find anything because search engines were basically terrible. Google on the other hand has and continues to expend considerable effort to ensure such syntax is unnecessary 99.9%+ of the time.
Consequently very few people know it any more. Combine this with the fact that in order to use the extended search syntax effectively you often need at least some understanding of what you're searching for (e.g., at least the ability to discern bad vs. good information), and you can quickly see why the average Joe is going to end up seeing and quite possibly believing the false information.
(I, for example, not being any kind of foodie, only know that it takes much longer than 10 minutes to caramelise onions specifically because I saw a friend who is a foodie do it a few years ago. He specifically made the point that it takes ages, otherwise you lose the sweet flavour. If that hadn't happened I'd believe it takes 10 minutes because Google said so.)
https://bynd.com/news-ideas/google-advanced-search-comprehen...
How about distance between words. (Smith married to Ellen - search for those words close to each other )
How about sounds similar. Like burger.
What about saving complicated searches.
The minus operator either no longer works at all, or it just penalizes results without eliminating them.
I'll have to send links to my Facebook foodie friends and see how long they take to get it.
Of course, I've never been mentioned in the news. Maybe the trick to being in a niche is keeping the niche quiet.
If there are other sites, I can't find them! :)
Being new in town is suspicious to google. Raising honestly to the top takes a lot of time and effort, unless you have incredible promotion habilities and/or deep SEO knowledge.
But now we are actually quite famous in the french Python community. For a blog on such a niche topic in a small country we get around 6000 v/d. The hard part was starting really: we (my co-author and I) have no promotion skill whatsoever.
We actually starting to gain traction after 3 major french bloggers linked to us completely out of the blue (sebsauvage, korben and lehollandaisvolant). It seems linking from trusted source is still the most important part of the game for google.
Eventually I want to create a new site, in english this time, to translate all the stuff and make it SFW so it can benefit more people. Unfortunately this won't be enough to test your theory since as soon as you create content in english, you multiply you traffic by 10000.
Anyway first I'll need to save for 5/6 months of budget before doing so cause as a freelancer I can't really take holidays. Maybe I should do a crowd founding.
It's helped me find unique and interesting sites in the past.
It might be time to consider the need for a way to search all of this content that isn't bound to people trying to monetize the experience. DDG seems like an option in that regard.
I miss the old unhomogenized web.