Google now defaults to not indexing your content
vincentschmalbach.com
vincentschmalbach.com
> Google has transformed from a comprehensive search engine into something more akin to an exclusive catalog.
That alone goes a long way to explain why Google search had become worthless to me. I had thought that it was mostly that their attempts at interpreting "what I really want" are terrible, but perhaps the reason is actually that they don't index what I really want in the first place.
I almost never want brand/big name sites and the like, but that is mostly what I get.
What a joke. :-/
I find much higher quality content on Yandex for those niche topics.
But the main problem now is that the content you're looking for just might not be in Google's index at all.
An additional data point: 10 years ago, Google would provide up to 1000 results per query. Now, it's usually no more than 100. They've even removed the total result count from search pages.
Now, those visionaries are gone. Google is run by managers focused on efficiency. Why maintain a massive, costly index when a smaller, cheaper one generates just as much revenue?
Their visions are gone but they are still there - they just have planes to fly around in, yachts to sail on and interns to sleep with instead of "doing no evil"
Google is doing the same thing with youtube. Youtube was a place to get away from traditional media, cable tv, etc. That was the point of the "You" in youtube after all. But now, no matter what I search, it's mostly corporate/traditional media in the list. Youtube was a major part of the 'cut the cord' movement. Now youtube is rebranding itself as 'cable tv re-imagined'.
It's amazing how google and youtube did a complete 180 in just a few years. Google search and youtube are nothing like what it was 10 or 15 years ago.
I wish - but now, no matter what I search, it's all idiot influencers youtube shock face thumbnails and contant that can only be targetted at retards.
Not being logged in, deleting cookies and rotating IP addresses may be wonderful for privacy, but the reccomendation engine shows you what the masses are watching and it is incredibly sad.
If you watch videos on space flights and search related terms all the time, I expect you'll get related recommendations and links on less mainstream channels.
It always makes me wonder: Who the heck uses these "search engines"?
Their existence suggests there is some completely captive audience to target, e.g., where the audience only knows one source for searching.
An audience that is not going to "reject" the results let alone evaluate them. Even if the results are poorly disguised advertisements. Because this audience is apparently incapable of locating another "search engine".
Perhaps this is also what happens when competition is stifled through anticompetitive practices. People never come to know what they are missing.
There isn't a shred of hard data to support the headline claim that Google now "defaults to not indexing content".
Google never indexed everything, removing duplicates, blogspam, useless pages, etc. Maybe they've changed their thresholds or maybe not. But this post provides zero evidence of anything. It's pure speculation without any facts at all.
Regarding "Google never indexed everything" - I'd say it came close. They did manual de-indexing for heavy spam sites and would even send an email when they did this. Apart from that, nearly everything was in the index, including duplicates. De-duplication happened at the ranking stage, not the indexing stage.
At some point, Google even had a second index, the Google Supplemental Index, for pages of lower importance.
You haven't provided even any comparison screenshots from over the years, much less graphs showing that this is an actual quantitative trend, as opposed to you just noticing different things now and maybe not remembering things entirely accurately.
I'm not saying that you're not telling the truth about your observations. But I am saying that you provide absolutely no basis for making broad sweeping generalizations about Google Search changing its "defaults".
Maybe you work in a category of low-quality content, while people in other categories have seen more of their pages indexed.
The point is, your personal qualitative observations from a miniscule subset of 100 sites is nowhere near sufficient to make a provocative headline claim such as "Google Now Defaults to Not Indexing Your Content".
When you published a new blog post on a fresh WordPress install, WordPress would automatically send out a ping. This feature, called "Ping-o-Matic" is still present in every WordPress installation. You can read more about it here:
https://developer.wordpress.org/advanced-administration/word...
Google likely picked up these pings and quickly indexed the new content. It wasn't an official process, but it worked reliably. The system is still in place, but Google doesn't seem to care anymore.
It'll find anything except what I'm trying to find. Quotes are useless. The content itself is often garbage. It works ok on common queries, but that's not when I need it to work - I need it to work the most when the query is hard. The long tail is the only thing that matters when a user is judging a search engine's quality.
The web itself has gotten worse over time, but that's also partly Google's doing. Google extracted all the value out of the open web and kept it for themselves. Meanwhile online publishers of all varieties are dying, despite being the ones producing much of the value. Google should have identified this as a strategic threat decades ago.
Now I just keep a tab open to ChatGPT all day and use it as a search engine without all the trouble of dealing with webpages.
I can't remember the last time a google search gave a useful result, outside of appending "reddit" to the answer so that the result is (probably) written by a human instead of bs blogspam. And I'm sure that won't work for long either.
A few days ago I was going to send a patch to Alpine Linux, and I remembered they had a document for how they want commit messages to be authored but couldn't remember the URL.
Google search results for "alpine linux commit message style doc" only returned garbage in the first ~50 results - pages from Alpine's wiki unrelated to commit message style, pages from other projects like the Linux kernel about their commit style, and batshit results like https://docs.ansible.com/ansible/latest/collections/communit... (happens to contain the word "commit").
So I asked ChatGPT "What is the URL of Alpine Linux's style guide for commit messages?" and it confidently told me it's https://gitlab.alpinelinux.org/alpine/aports/-/blob/master/C... , which does not exist. I replied "That is a 404" and it responded "I apologize for the inconvenience. Alpine Linux's style guide for commit messages may have been moved or updated since my last training data. To find the most current style guide for Alpine Linux's commit messages, you can visit their repository directly and look for the CONTRIBUTING.md file or [go read their documentation and figure it out yourself, paraphrased]". The first one is a lie - a file named CONTRIBUTING.md has never existed in the master branch of that repo, as git easily reveals. The second one is schizophrenic - it's telling me to read the same file I just told it doesn't exist. The third one is unhelpful.
I had to find it myself. (It's https://gitlab.alpinelinux.org/alpine/aports/-/blob/master/C... )
It got back with "The URL for Alpine Linux's style guide for commit messages can be found in the README.md file of the aports repository on GitLab. The specific URL is: https://gitlab.alpinelinux.org/alpine/aports/-/blob/master/R..."
(+ extra links that include the one from the end of this document, which is what you're after)
It could do better by returning the actual link rather than a page where you can find it, but it's still better then Google. (Kagi pointed at the same readme at least)
https://gitlab.alpinelinux.org/alpine/aports/-/commit/0b7097...
In fact now the first result for me is my very comment above, but the actual COMMITSTYLE.md is still not in the first 50 results. The only result from even the right domain in the first 50 results is https://gitlab.alpinelinux.org/alpine/docs/user-handbook (at #12, and of course irrelevant).
I recently logged into google and asked that they index my domain (djhaskin.com). They asked me to put a TXT record in there proving I owned it and I did so. Then their "website console" thing showed that my website was indexed[1]. They have a console for this stuff now[2]. They recently showed me a page in there where that displayed some URLs that weren't indexed and which were. I requested a re-index of one of the non-indexed URLs, but the others were just broken/junk/RSS feed urls, so it was fine that they weren't indexed. The console gave me a ton of tools for making sure my site was indexed, and told me why if it wasn't.
I had plenty of tools to get my site indexed and felt like I was in control. I don't feel any sense of mystery about what is happening and I receive notifications when indexing fails.
That makes your real content a smaller proportion of the whole web, and therefore less likely to meet the threshold for fitting inside googles finite indexing budget.
They want a single box that can do everything because they don't think an user can be given a minimal amount of training to use a search engine, and they want all results ranked in a single, linear list similarly.
It's becoming increasingly clear that when you do that, all long tail results turn into garbage.
Google has become essentially a phonebook. If you type the name of a thing, Google can quickly find its official website. Anything more complex than that and it quickly derails.
A good example is reviews. Google literally can't find reviews if you type "reviews" in the query, because it can only find webpages that strictly contain the terms you typed. This means that for a review of something to be found, the writer would have to literally write "reviews" in their webpage, or Google would have to interpret the query non-literally and heuristically categorize webpages as reviews, which is error-prone, because then you'll have lots of webpages that aren't reviews appearing when you search for reviews and the word "review" you typed is apparently ignored completely.
The whole authority concept also feels fundamentally flawed.
If the most brilliant researcher in a field has a blog that is a gold mine of information, that blog is never going to appear in any search result because this person would be too busy doing their research to waste their valuable time building backlinks to rank.
There are too many cases where "ranking" will go wrong. It's wrong to assume Google can deliver the best result as the first result, as it has no way to make value judgements on the content of articles, but this assumption apparently drives Google itself to implement policies to hurt its own results. They seem to only care about the first result. They don't want users to browse a whole list of results. I believe "browsing" is the key word here. If it's not possible to browse, it's not possible to find it yourself, and you depend entirely on a machine that is a black box to do the finding for you. When the machine fails, there is no recourse.
with five ads before it pretending to be its official website
Their website actually still says this: https://www.google.com/search/howsearchworks/our-approach/.
It seems from their recent actions their mission is to "organize the world's information and drip it out as slowly as possible, covered in ads".
If you have had it with (c) and paywals and ads, come join the revolution which is the World Wide Scroll: https://wws.scroll.pub/
I remember the olden days when Google was hailed for its brilliance in A/B testing the subtle shades of colors it used on its home page.
Why should we think they aren't producing the product that most people want?
Most people were hosting on AngelFire or Groceries back in the day, but as we see today that's not what people wanted.
While OpenAI etc. is pretty good (so does Google Gemini) what is OpenAI like interfaces prevent me from doing is to segue from a focused topic to related areas to discover knowledge on the periphery, which is the most important aspect of learning in my opinion which chatbots today are not able to do that well.
HN historically has lot of G haters. Which is fine, but I feel a lot of criticism is not really reasonable.
In the past, Google would index nearly everything they could find, then filter content through ranking algorithms. They would downrank what they considered junk, but with enough query refinement, you could still find it. This approach also allowed you to find sites that Google might have incorrectly categorized as low-quality but were actually what you were looking for.
Now, the situation is different. Google isn't just doing its filtering at the ranking stage, but at the indexing stage. This means you can refine your query endlessly, but you'll never find content that isn't in the index to begin with.
I doubt if this is a new thing though and both options are not mutually exclusive. Google probably now has enough signals to detect a vast majority of pages to be not worthy of even crawling. I would be surprised if they figured this out in 2024 and not in 2014.
Chances are these pages are absolute garbage.
PS. Google does not benefit at all promoting Instagram or Tiktok. They benefit from simpler an independent content lot more. Clearly it would be other way around. I have seen Pinterest simply disappear from Google recently to give you an example.
My criticism of Google search is unrelated to my feelings about Google itself. My criticism of Google search is that it gives me truly awful results from my queries.
That's objectively true, as demonstrated by the fact that other search engines (even ones that use Google in the background) are better at that.
1. Register on Google Search Console
2. Go to the page indexing section
3. Look at the rows "Discovered - currently not indexed" and "Crawled - currently not indexed"
For example, my own site has a two-digit number of URLs in both categories. These are blog posts Google simply doesn't want to index for reasons unknown.
I have access to Google Search Console data for over 100 websites, and most/all of them have the same issues. This includes sites (like my own) that rank well for certain keywords and receive traffic.
[0] https://search.marginalia.nu/ [1] https://blog.kagi.com/small-web
Upranking/downranking sources turns out to be more powerful and meaningful than I thought at first. At first, I didn't even bother with it, but it turns out that being able to provide a "correction factor" signal makes a serious difference.
I don't use the lenses myself because I don't need different rankings for different things, but they might address your needs.
Want an answer to a question or to learn more about a topic? Use an answer engine like perplexity.
My need to go to a search engine is dropping quite a lot now because of LLMs.
I’m using Kagi as well which is nice.
It's even more impossible to satisfy them in a way that's also useful for Google's own business model.
The most obvious and potentially naive way of doing that would be to allow end users to upvote or downvote search results. I know that Google already does that in an automated way to some extend, but the problem is that that signal is then used to determine the quality across all the users which, as I mentioned in the beginning, is impossible to get right.
Instead that metric should only affect each user's own search results and not everyone else's. This could improve the quality of the results, bring more people back, and eventually increase the revenue. It would also help prevent "gaming" the system.
What am I missing here? I can't believe that they're not aware of that. I also can't believe that they don't want to fix it or that they're so focused on ai that internal politics don't allow them to do anything else. So what gives?
Should I start to write text next to each image, like:
"Mech approaches another dark, very evil looking mech on a bright day and swings its laser sword to decapitate the evil one."
Gimme a break. For 5K+ images... :D. The topic, title and description is not enough, it needs more text to believe that this is what the images are about. No, your images will not be indexed, will not be included in the image search results, because you are not part of the exclusive club. No monopoly here.
Bing webmaster tools tells me that there are no highly ranked sites that link to my content, and there should be :D. I just started the site, how would there be any linkage to it?! Are they insinuating that I should create fake sites to promote my content or maybe pay for seo? No monopoly here either.
Yandex... I can't really figure out whether their indexing works or not, it sometimes complains, then doesn't do anything for weeks. Then it comes up with another made up problem that is nonexistent. It acts like a drunkard.
I haven't tried Baidu, because they need some local phone number and they clearly can't send activation SMS to Europe.
Next I'll try with a news site and write a blog post about my experiences. Truly interesting times.
You've got a pile of images with no descriptive text. What kind of query do you think would make a search engine include one of those pages in the result pages? They need something to work with.
At most it sounds like it could be people using your site name in the query. But that only happens when you have users, you can't bootstrap your popularity like that.
And if there is no set of reasonable queries where your pages are likely results, what is the point of of indexing them?
Although Google these days definitely runs image recognition and OCR while indexing images for its image search.
That's the only plausible long-term path to keep its search results competitive and relevant.
> For content creators, it presents a significant challenge: how do you gain visibility if Google refuses to index most of your content?
Don't publish junk that it doesn't find interesting under the assumption that it "owes" you a front page search result.
I'm not a Google fanboy. I have a paid Kagi account and I use it nearly exclusively. I want Google to stay competitive though. There's a vast army of SEO spammers who think they know the one magic invocation that will drive traffic into their willing arms, or, at least, are able to convince paying customers of it. If Google could wave a magic wand that could accurately identify all of the junk that exists purely to increase results rankings, and they used that knowledge to permanently remove it from the results they show me, well, I'd probably stop paying Kagi to do that.
In the past, users could refine their search keywords until they found what they wanted. This approach doesn't work as well now. The main reason? The content you're looking for might not be in the index at all. Google's selective indexing aims to reduce spam, but it also limits the diversity and depth of discoverable information.
Countless terabytes of junk web content make it hard to find the needle in the haystack. Frankly, I don't want to see that low-value part of the web. Every time it shows up in search results, that's one less decent result appearing on my screen. I suspect that anyone who says they want to see everything doesn't really appreciate what everything means. There will never be a page on Experts Exchange that I want to see. I'll never be grateful for a result from answers.com. Although those sites are capable of creating useful information, their business model been to dump a ton of junk on search engines with the hope that people accidentally click on it so they can make some ad money. I picked on them because I can name them off the top of my head. There are a million sites just like them. And none of them have anything to offer me whatsoever.
I'd hate for Google's changes to remove the site maintained by that expert kid in Tampa who has precisely the answer I need, along with part numbers and wiring diagrams. If Google has to nuke the Experts Exchanges and answers.com of the world so that the Tampa kid's site does start showing up, yay!
> Many valuable, high-quality pieces of content that people would find useful never make it into Google's index.
I see your point in the light of the article (not indexed = not visible), but it feels like the things that _do_ make it need to follow very particular content and style patterns to rank high.Anecdotally, this observation comes from searching for any term and seeing the results: they are usually similar-looking plausible-looking-but-actually-low-quality results that seem to follow the same or similar structure and have the same content. This does indeed limit the diversity and depth of information, but I'm not so sure it reduces spam, as these low-quality sites seem to be as prevalent as ever before, if not more.
From experience writing articles to a small tech blog, this means that it's quite difficult to get well-researched articles to rank well, even if they're indexed.
For example, I've written an article on how to block hotlinking (I've just checked, and Google says it's indexed). If you search for this, my article on a not-so-well-known blog is nowhere to be found(*)(**), and this is somewhat expected, for a myriad of reasons. The problem isn't that my post doesn't rank, but rather that none of the top-ranking (or even not-so-top-ranking) results are wrong. They are either about how to do this on cPanel or whatever, which is ineffective (but granted, could be what people are looking for), or instructions using the `Referer` header, which is ineffective.
These days, browsers offer headers like `Cross-Origin-Resource-Policy` which can completely solve the particular issue of hotlinking, unlike `Referer` which is easily bypassed using `referrerpolicy="no-referrer"`. However, because most 'authorities' seem to be wrong on this issue, the correct result isn't displayed, because it's a hard problem to solve algorithmically (or even manually).
(*) This doesn't affect just Google, though.
(**) Because it's indexed, adding the right keywords (which you wouldn't do in this case unless you already knew the answer) does bring it up, although from federated high-authority sites instead of the original canonical source.
Give me an embeddable iframe that I can optionally add to my site which allows visitors to give feedback (think a floating Reddit upvote button). Require that button's access to require an authenticated user via Google.
Ranking in the algorithm is weighted such that the organic user votes are the heaviest, and content length, keywords, etc., are the lowest.
Remove all the AI crap (or tweak it so I can chat with a bot to improve my search, but it's not Foie Gras'd down my throat). Make ads a free vs paid experience (want to avoid ad results, pay Google $5/mo for a clean result set).
This would make the only way to "game" SEO authentic, quality content. In essence, it's taking the current hack (appending "Reddit" to the end of a search query) and building it into the core experience.
Wouldn't that mean that tons of people won't be able to "vote"?
Ultimately this needs to be part of the browser chrome, since you can't trust the site.
It could get murky, but you could even have votes weighted by an account score (not quite sure how that would take shape).
And so then the way to rank is to get in a group chat with all your SEO buddies and trade ~~backlinks~~ upvotes with each other. Or make bot accounts that appear real.
Even if this idea worked, the idea of giving Google even more consolidated power/knowledge over the open web is horrifying to me.
Upvote farming could be mitigated by vote frequency. That sort of behavior would be obvious in the data as you'd see significant voting activity over some time range (which wouldn't look like organic search behavior).
You can't outright avoid the consolidation of power with search engines, unless (and I think this would be smart for someone to do), you create a decentralized version of what I described that's part of an alternative search engine.
I'm not saying they're applying the best of ideas and doing it well, but I'd wager none with good experience in this field would bet it would be that easy. I think it's more akin to hackers and defenders, each surpassing the other one in a never ending race.
Yes. Google hasn't really produced anything of substance post-2010. They fell into the typical corporate bureaucracy trap. The rejection of ideas like mine are likely even more common internally in Google. In an org of that size/stature, people jockey for position, they don't create (and that's not an opinion, it's objective reality).
O.G. Goolge would try anything (I remember they offered radio station automation software briefly).
Could I be wrong? Sure. But if public opinion of my core product is waning massively, I'd rather trust the mad scientist than the Wharton dork who's focused on his L-status. Worst case scenario the failure is a rounding error on my balance sheet.
I wonder if Stumbleupon could have achieved something similar to what you're proposing if they had tried to build a search engine.