CNET is deleting old articles to try to improve its Google Search ranking
theverge.com
theverge.com
Can there be any doubt that Google destroyed the old internet by becoming a bad search engine? Could their exclusion of most of the web be considered punishment for being sites being so old and stable that they don't rely on Google for ad revenue?
These are all the sites (like CNET) that Google indexes which are the entire reason to use search. They are having their rankings steadily eroded by an ever-rising tide of SEO spam. If they start dying off en masse and if LLMs emerge as a viable alternative for looking up information, we may see Google Search die along with them.
As for why their revenues are still increasing? It's because all the SEO spam sites out there run Google Ads. This is how we close the loop on the "killing the golden goose" theory. Google uses legitimate sites to make their search engine a viable product and at the same time directs traffic away from those legitimate sites towards SEO spam to generate revenue. It's a transformation from symbiosis/mutualism to parasitism.
Edit: I forgot to mention the last, and darkest, part of the theory. Many of these SEO spam sites engage in large-scale piracy by scraping all their content off legitimate sites. By allowing their ads to run on these sites, Google is essentially acting as an accessory to large-scale, criminal, commercial copyright infringement.
[Disclosure: Google Search SWE; opinions and thoughts are my own and do not represent those of my employer]
Why do you assume malicious intent?
The balance between search ranking (Google) and search optimization (third-party sites) is an adversarial, dynamic game played between two sides with inverse incentives, taking place on an economic field (i.e. limited resources). There is no perfect solution; there’s only an evolutionary act-react cycle.
Do you think content spammers spend more or less resources (people, time, money) than Google’s revenue? So then the problem becomes how do you win a battle with orders of magnitude less people, time, and money? Leverage, i.e., engineering. You try your best and watch the scoreboard.
Some people think Google is doing a great job; some think we couldn’t be any worse. The truth probably lies across a spectrum in the middle. So it goes with a globally consumed product.
Also, note, Ads and Search operate completely independent. There’s no signals going from Ads to Search, or vice versa, to inform rankings; Search can’t even touch a lot of the Ads data, and Ads can’t touch Search data. Which makes your theory misinformed.
That’s the mistake. They should be talking. Sites that engage in unethical SEO to game search rankings should be banned from Google’s ad platform. Why aren’t they? Because Google is profiting from the arrangement.
Not GP, but to me, admittably a complete non-expert on search, there are so many low-hanging fruits if search result quality was anywhere on Google's radar that it is really difficult not to assume malicious intent.
Some examples:
- why pinterest is flooding the image results with absolute nonesense? How difficult it would be to derank a single domain that manages to screw google's algorithm totally?
- why there is no option for me to blacklist domains from the search result? Are there really some challenges that can't be practically solved in a couple of minutes of thinking?
- Does google seriously claim they can't differentiate between stackoverflow and the content copying rip-off SEO spam sites?
You might already be aware of this, but you can use uBlock Origin to filter google search results.
Click on the extension --> Settings --> My Filters. Paste in the bottom
google.*##.g:has(a[href*="pinterest.com"])
Every time I get mislead on clicking onto an AI aggregator site, my filter list grows...There doesn't have to be any malicious intent, just an endless chase for increased profit next quarter. SEO spam has more ads, thus generates more income for Google. Even if Ads and Search operate "completely independently", there must be a person in the corporate hierarchy which has control over both and could push the products to better synergize and make that KPI tick up.
Actually deranking sites which feature more than three Google Ads banners would improve search quality (mainly by making sites get to the point rather than padding a simple answer into an essay like an 8th grader at an exam) - but it would reduce Ads income so you cannot do it, no matter how independent you claim to be.
Unless you have a clear view by leadership of what they desire the web should be and are willing to disclose it in detail, then there's not much to add by saying you work in Search.
At one time both buggy whips and Philco radios had hockey stick growth charts, too.
You must not bet old enough to remember when people thought MySpace would always drive the internet.
You know there's a difference between "people thought" and dollars.
Some people think the earth is flat. Opinions can change very quickly - like 5 years ago, people thought elon musk was the hero of the internet.
Here's the stats on MySpace revenue: It generated $800 million in revenue during the 2008 fiscal year.
Google could be generating $27 billion in revenue in 2038 and be considered a massive failure compared to what it is now.
I fail to see the point you are trying to get at?
At some point a competitor will emerge. The tech crowd will notice it and begin to use it. Then it will go widespread.
I don't think Google dying would be good (lots of things would have to migrate infra suddenly), but the adtech being split off into something else would certainly be a welcome turn of events, IMO. I'm tired of seeing promising ideas killed because they only made 7-figure numbers in a spreadsheet where it'd have been viable on its own somewhere it wasn't a rounding error.
[1] https://twitter.com/searchliaison/status/1689018769782476800
But the vast majority are morons, grifters, and cargo culters.
The Google guidance is generally good and mildly informative but there’s a lot of depth that typically isn’t covered that the SEO industry basically has to black box test to find out.
The more Google helps you get ahead, the more you end up dominating the search results. The more you dominate the results, the more people will start thinking to come straight to you. The more people come straight to you, the more people never use Google. The less people use Google, the less revenue Google generates.
This is a possible outcome but there are people that type in google.com and then the name of their preferred news site, their bank, etc, every day.
The site with the name they search dominates that search but they keep searching it.
Google has a vested interest in creating a good web experience. Consultants have an interest in making their clients money.
Link building is a classic example where good consultants deliver value. (There are bad consultants than good ones though)
That's because websites' goals and Google's goals are not aligned.
Websites want people to engage with their website, view ads, buy products, or do something else (e.g. buy a product, vote for a party). If old content does not or detracts from those goals, they and SEO experts say, it should go because it's dragging the rest down.
Google wants all the information and for people to watch their ads. Google likes the long tail; Google doesn't care if articles from the 90's are outdated because people looking at it (assuming the page runs Google ads) or searching for it (assuming they use Google) means impressions and therefore money for them.
Google favors quantity over quality, websites the other way around. To oversimplify and probably be incorrect.
Perhaps what CNET really means is that they're deleting old low quality content with high bounce rates. After all, the best SEO is actually having the thing users want.
Sure, one page doesn’t matter, but thousands will.
[1] https://twitter.com/searchliaison/status/1689068723657904129...
https://twitter.com/searchliaison/status/1689297947740295168
>Removing it might mean if you have a massive site that we’re better able to crawl other content on the site. But it doesn’t mean we go “oh, now the whole site is so much better” because of what happens with an individual page.
Parsing this carefully, to me it sounds worded to give the impression removing old pages won’t help the ranking of other pages without explicitly saying so. In other words, if it turns out that deleting old pages helps your ranking (indirectly, by making Google crawl your new pages faster), this tweet is truthful on a technicality.
In the context of negative attention where some of the blame for old content being removed is directed toward Google, there is a clear motive for a PR strategy that deflects in this way.
CNet is big enough that I’d expect Google to ensure the crawler has fresh news articles from it, but that isn’t explicitly said anywhere.
Apparently not if this SEO trick is really a thing...
EDIT : sorry my bad it's actually the opposite. One could expect that a site like CNET would include a timestamp and a unique ID in their URL in 2023. This seems to be the "unpermalink" of a recent cnet article.
Maybe the SEO expert could have started there...
https://www.cnet.com/tech/mobile/samsung-galaxy-z-flip-5-rev...
It's not worded in any way intended to be parsed. I mean, I guess people can do that if they want. But there's no hidden meaning I put in there.
Indexing and ranking are two different things.
Indexing is about gathering content. The internet is big, so we don't index all the pages on it. We try, but there's a lot. If you have a huge site, similarly, we might not get all your pages. Potentially, if you remove some, we might get more to index. Or maybe not, because we also try to index pages as they seem to need to be indexed. If you have an old page that doesn't seem to change much, we probably aren't running back ever hour to it in order to index it again.
Ranking is separate from indexing. It's how well a page performs after being indexed, based on a variety of different signals we look at.
People who believe removing "old" content aren't generally thinking that's going to make the "new" pages get indexed faster. They might think that maybe it means more of their pages overall from a site could get indexed, but that can include "old" pages they're successful with, too.
The key thing is if you go to the CNET memo mentioned in Gizmodo article, it says this:
"it sends a signal to Google that says CNET is fresh, relevant and worthy of being placed higher than our competitors in search results."
Maybe CNET thinks getting rid of older content does this, but it's not. It's not a thing. We're not looking at a site, counting up all the older pages and then somehow declaring the site overall as "old" and therefore all content within it can't rank as well as if we thought it was somehow a "fresh" site.
That's also the context of my response. You can see from the memo that it's not about "and maybe we can get more pages indexed." It's about ranking.
If by pruning old content, CNET can get its new articles in the results faster, it seems this would get CNET higher rankings and more traffic. Google doesn’t need to have a ranking system directly measuring the average age of content on the site for the net effect of Google’s systems to produce that effect. “Indexing and ranking are two different things” is an important implementation detail, but CNET cares about the outcome, which is whether they can show up at the top of the results page.
>If you have a huge site, similarly, we might not get all your pages. Potentially, if you remove some, we might get more to index. Or maybe not, because we also try to index pages as they seem to need to be indexed.
The answer is phrased like a denial, but it’s all caveated by the uncertainty communicated here. Which, like in the quote from CNET, could determine whether Google effectively considers the articles they are publishing “fresh, relevant and worthy of being placed higher than our competitors in search results”.
No. Google has their own motivations here, they are a player not a rule maker.
Don’t trust SEOs as no one actually knows what works, but certainly dont think google is telling you the absolute truth.
Unreadable source code is a crime against humanity.
If we wanted browsers to be fed code that for performance reasons isn't human-readable, web servers ought to serve something that's processed way more than just gzipped minification. It could be more like bytecode.
About the byte code: You mean wasm? (Guess that's what you're alluding to.)
Isn’t this essential what WebAssembly is doing? I’ll admit I haven’t looked into it much, as I’m crap with C/++, though I’d like to try Rust. Having “near native” performance in a browser sounds nice, curious to see how far it’s come.
Worth keeping in mind that "performance" here refers to saving bandwidth costs as the host. Every single unnecessary whitespace or character is a byte that didn't need to be uploaded, hence minify and save on that bandwidth and thus $$$$.
The performance difference on the browser end between original and minified source code is negligible.
What really adds to the source code footprint is all of those trackers, adverts and, in a lot of cases, framework overhead.
I'm actually all for binary HTML – not just it's smaller, it can also be easier to parse, and makes more sense overall nowadays.
For me I guess what I was getting at is that I consider source the stuff I'm working on - the minified output I won't touch, it's output. But it is input for someone else, and available as a View Source so that does muddy the waters, just like decompilers produce "source" that no sane human would want to work on.
I think semantically I would consider the original source code the "real" source if that makes sense. The source is wherever it all comes from. The rest is various types of output from further down the toolchain tree. I don't know if the official definition agrees with that though.
Plus I suspect minifying HTML or JS is often cargo cult (for small sites who are frying the wrong fish) or compensating for page bloat
You can always open dev tools in your browser and have an interactive, nicely formatted HTML tree there with a ton of inspection and manipulation features.
Not to mention, there is no need to "minify" HTML, CSS, or JavaShit for a browser to render a page unlike compiled code which is more or less a necessity for such things.
> The program must include source code [...] The source code must be the preferred form in which a programmer would modify the program. Deliberately obfuscated source code is not allowed. Intermediate forms such as the output of a preprocessor [...] are not allowed.
By your logic, there's actually no reason to use compiled code at all, for almost anything above the kernel. We can just use Python to do everything, including run browsers, play video games, etc. Sure, it'll be dog-slow, but you seem to care more about reading the code than performance or any other consideration.
Glad to see the diversity of HN readers apparently includes twelve year olds.
Anyway, you do realise plenty of languages have compilers with JS as a compilation target, right? How readable do you think that is?
The abuse of JavaScript does not deserve the respect of being called by a proper name.
>Anyway, you do realise plenty of languages have compilers with JS as a compilation target, right? How readable do you think that is?
If you're going to run "compiled" code through an interpreter anyway, is that really compiled code?
If you dislike unreadable source code I would assume you would object to minifying JS, in which case you should ask people to include sourcemaps instead of objecting to minification.
And some are the most data-driven people you'll ever meet. As with most people who claim to be experts, the trick is to determine whether the person you're evaluating is a legitimate professional or a cargo-culting wanna-be.
Excellent quote. It's counterintuitive but looking at what is most likely to happen according to the datasets presented can often miss the bigger picture.
or just plain numerology
Right for the wrong reasons is still right.
The yandex source code leak revealed that keyword proximity to root domain is a ranking factor. Of course, there’s nearly a thousand factors and “randomize result” is also a factor, but still.
SEO is unfortunately a zero sum game so it makes otherwise silly activities become positive ROI.
I assume Google is turning all the site's pages, js, inbound/outbound links, traffic patterns, etc...into large numbers of sometimes obscure datapoints like "does it have a favicon", "is it a unique favicon?", "do people scroll past the initial viewport?", "does it have this known uncommon attribute?".
Maybe those aren't the right guesses, but if a page has thousands of derived features and attributes, maybe they are on the list.
So, some SEO's take the idea that they can identify sites that Google clearly showers with traffic, and try to recreate as close a list of those features/attributes as they can for the site they are being paid to boost.
I agree it's an odd approach, but I also can't prove it's wrong.
It reminds me of Apple's "don't run to the press" advice when hitting bugs or app review issues. While we'd assume Apple knows best, going against their advice totally works and is by far the most efficient action for anyone with enough reach.
Nope. It says that Google does not ding you for old content.
"Are you deleting content from your site because you somehow believe Google doesn't like "old" content? That's not a thing!"
I guess that Googler never uses Google.
It's very hard to find anything on Google older than or more relevant than Taylor Swift's latest breakup.
This is not the same as saying that it doesn't prioritize newer pages over older pages in the search results.
The way it's worded does sound like it could imply the latter thing, but that may have just been poor writing.
With archive search, the News section floats links like https://www.nytimes.com/2008/11/09/arts/music/09cara.html
Our ranking system with freshness is explained more here: https://developers.google.com/search/docs/appearance/ranking...
But we do show older content, as well. I find often when people are frustrated they get newer content, it's because of that crossover where there's something fresh happening related to the query.
If you haven't tried, consider our before: and after: commands. I hope we'll finally get these out of beta status soon, but they work now. You can do something like before:2023 and we wouldn't show pages from before 2023 (to the best we can determine dates). They're explained more here: https://twitter.com/searchliaison/status/1115706765088182272
-George Orwell, 1984
My bet is that they don't. My bet is that there is so much old code, weird data edge cases and opaque machine-learning models driving the search results, Google's engineers have lost the ability to predict what the search results would be or should be in the majority of cases.
SEO experts might not have insider knowledge, but they observe in detail how the algorithm behaves, in a wide variety of circumstances, over extended periods of time. And if they say that deleting old content improves search ranking, I'm inclined to believe them over Google.
Maybe the people at Google can tell us what they want their system to do. But does it do what they want it to do anymore? My sense is that they've lost control.
I invite someone from Google to put me in my place and tell me how wrong I am about this.
Once upon a time, Matt Cutts would come on HN give a fairly knowledgeable and authoritative explanation of how Google worked. But those days are gone and I'd say so are days of standing behind any articulated principle.
In addition, if you want an explanation of how Google works, we have an entire web site for that: https://www.google.com/search/howsearchworks/
So there's zero machine learning or statistical modeling based functionality in your search algorithms?
if CNET deletes all their old articles, they're making a situation where most links to CNET from other sites lead to error pages (or at least, pages with no relevant content on them) and even if that isn't currently a signal used by google, it could become one.
It's not that we discourage it. It's not something we recommend at all. Not our guidance. Not something we've had a help page about saying "do this" or "don't do this" because it's just not something we've felt (until now) that people would somehow think they should do -- any more than "I'm going to delete all URLs with the letter Y in them because I think Google doesn't like the letter Y."
People are free to believe what they want, of course. But we really don't care if you have "old" pages on your site, and deleting content because you think it's "old" isn't likely to do anything for you.
Likely, this myth is fueled by people who update content on their site to make it more useful. For example, maybe you have a page about how to solve some common computer problem and a better solution comes along. Updating a page might make it more helpful and, in turn, it might perform better.
That's not the same as "delete because old" and "if you have a lot of old content on the site, the entire site is somehow seen as old and won't rank better."
But lying on the internet isn't a crime. I work for Google on quantum AI solutions in adtech btw.
Your job is not to disseminate accurate information about how the algorithm works but rather to disseminate information that google has decided it wants people to know. Those are two extremely different things in this context.
I work on these kind of vague "algorithm" style products in my job, and I know that unless you are knee deep in it day to day, you have zero understanding of what it ACTUALLY does, what it ACTUALLY rewards, what it ACTUALLY punishes, which can be very different from what you were hoping it would reward and punish when you build and train it. Machine learning still does not have the kind of explanatory power to do any better than that.
He was the guy who replaced Matt Cutts.
But who knows, really? They run things to extract features nobody outside of Google knows that are proxies for "content quality". Then run them through pipelines of lots of different not-really-coordinated ML algorithms.
Maybe some of those features aren't great for older pages? (broken links, out-of-spec html/js, missing images, references to things that don't exist, practices once allowed now discouraged...like <meta keywords>, etc). And I wouldn't be surprised if some part of overall site "reputation" in their eyes is some ratio of bad:good pages, or something along those lines.
I have my doubts that Google knows exactly what their search engines likes and doesn't like. They surely know which ads to put next to those maybe flawed results, though.
It's really hard to not create perverse incentives with KPIs. Targets like "% of tickets closed within 72 hours" can wreck service quality if the team is under enough pressure or unscrupulous.
Generaly speaking, an easy solution is to attach another target to either the nominator or denominator, a target that requires people to move that in value in acertqin direction. That might even be a different team thanthe one having goals on the ratio.
These are good in that they’re directly aligned with business outcomes but you still need sensible judgement in the loop. For example, say there’s an ice storm or heat wave which affects delivery times for a large region – you need someone smart enough to recognize that and not robotically punish people for failing to hit a now-unrealistic goal, or you’re going to see things like people marking orders as canceled or faking deliveries to avoid penalties or losing bonuses.
One example I saw at a large old school vendor was having performance measured directly by units delivered, which might seem reasonable since it’s totally aligned with the company’s interests, except that they were hit by a delay on new CPUs and so most of their customers were waiting for the latest product. Some sales people were penalized and left, and the cagier ones played games having their best clients order the old stuff, never unpack it, and return it on the first day of the next quarter - they got the max internal discount for their troubles so that circus cost way more money than doing nothing would have, but that number was law and none of the senior managers were willing to provide nuance.
My organization tracks how many tickets we have had open for 30 days or more. So my team started to close tickets after 30 days and let them reopen automatically.
Of course the actual behavior in the article is highly disturbing.
https://twitter.com/searchliaison/status/1689068723657904129
My understanding is that if you have a very large site, removing pages can sometimes help because:
- There is an indexing "budget" for your site. Removing pages might make reindexing of the rest of the pages faster.
- Removing pages that are cannibalising on each other might help the main page for the keywords to rank higher.
- Google is not very fond of "thin wide" content. Removing low quality pages can be helpful, especially if you don't have a lot of links to your site.
- Trimming the content of a website could make it easier for people and Google to understand what the site is about and help them find what they are looking for.
There is no way the PR team making that tweet can say for sure that deleting old content doesn't improve rank. Nobody can say that for sure. The neural net is a black box, and it's behaviour is hard to predict without just trying it and seeing.
Words without action are useless.
People have been committing horrifying atrocities in the name of SEO for years. I've seen it firsthand. And it spectacularly backfired each time.
This can very probably be yet another one of such cases.
Also, it will slow down the crawl frequency if you noindex it
So it's a non problem
You know, like a news website that's been on the internet since the 90s
Eventually Google stops crawling noindexed pages.
Wasn't the point of them tracking us so much to customize and cater our results? Why have they normalized everything to some focus group persona of Joe six-pack?
***
Let's try an experiment
Type in "Chicago ticket" which has at least 4 interpretations, a ticket to travel to Chicago, a citation received in Chicago, A ticket to see the musical Chicago and a ticket to see the rock band Chicago.
For me I get the Rock band, citation, baseball and mass transit ticket in that order.
I'm in Los Angeles, have never been to Chicago, don't watch sports, and don't listen to the rock band. Google should know this with my location, search and YouTube history but it apparently doesn't care. What do you get?
Where you'll see integrations like this used to be assistant, but is now bard. Both of which have lawyer boggling Eulas and a brand they can sacrifice if need be.
My ads are:
- Flights - Chicago the Musical - Flights - More Flights
My search results are:
- Citation - Citation payment plan - News report on lawsuit regarding citations in Chicago - Baseball
I also live very far from Chicago, and the only time I was there was for a connecting flight some time in the '90s.
It knows you're in LA and did not look up "Chicago flight", so you probably aren't looking for flights there.
Chicago musical isn't playing in LA so probably not the right kind.
Probably why most people get parking ticket listed higher. It would be interesting to see the results in a city where the band, team or musical has an event soon.
I suspect this is a long-term play for Google to phase out this search and replace with Bard. Think about it all these articles are doing now is writing a verbose version of what Bard gives you directly unless it’s new human content.
Google has in essence stolen all their information by scraping and storing in a database for its LLM and is offering its knowledge of this directly to users, so in a way, this is akin to Amazon selling its own private label products.
Not sure if the kagi folks are willing to share, but I get the impression that pagerank, tf-idf and a few heuristics on top would still get you pretty far. Add some moderation (known-bad sites that e.g. repost stackoverflow content) and you're already ahead of what google gives me.
False. Google Directory (and many other major name, mostly now defunct, web directories) were powered by data from DMOZ which was crowdsourced (and kind of still is, through Curlie [0] and while some parts of the website show updated as recently as today, enough fairly core links are dead or without content that its pretty obviously not a thriving operation.) Also, it was not pre-Wiki: WikiWikiWeb was created in 1995, DMOZ in 1998. It was pre-Wikipedia, but Wikipedia wasn’t the first Wiki.
That said, in NL a lot of people's home pages for a long time was set to startpagina.nl, which was just that, a cool directory of websites that you could submit to the website. It seems to exist still, too.
People discovered that Google measures not only how much time you stay on a webpage but also how much you scroll to define how interesting a website is. So now every crappy "tech tips" website that has an answer that fits in a short paragraph now makes you scroll two pages before you get the thing you actually wanted to read.
I know you came here looking for a recipe for boiled water, but first here's my thesis on the history and cultural significance of warm liquids.
Water boils in pots with different metals because only temperature matters for boiling water. If the water is 100c at sea level, it will boil.
Surely Google is too afraid of our vigorous pro-competition regulatory agencies and would never do such a thing.
I search for "how to do X", and instead of just showing me how to do it, which might take 30 seconds, they put a ton of fluff and filler in to the video to make it last 5 minutes.
Typical video goes something like:
0 - Ads, if you're not using an ad blocker
1 - Intro graphics/animation
2 - "Hi, I'm ___, and in this video I'm going to show you how to do X"
3 - "Before I get in to that, I want to tell you about my channel and all the great things I do."
4 - "Like and subscribe."
5 - "Now let's get in to it..."
6 - "What is X?"
7 - "What's the history of X?"
8 - "Why X is so great."
9 - finally... "How to do X"
Fortunately you can skip around, but it's still a bunch of useless fluff and garbage content to get to the maybe 30 seconds of useful information.
They're probably reaching some YT threshold to have more ads show on it
(The converse is; whenever I turn on a Hololive stream I'd say that I pick up 20% of what they're saying. If they talked slower, I would probably watch more than every 3 months. But, they rightfully don't feel the need to optimize for non-native speakers of Japanese.)
100%, and you can skim it so that your pace subconsciously varies depending on how relevant or complex that section is.
I wish more stuff were available in just text + screenshots..
That said, tinkering before and after youtube has been two different worlds. I really like having video to learn hands-on activities. I just wrapped up some mods to a Rancilio Silvia, and I noticed my workflow was videos, how-to guides and blog posts, broader electrical information documentation, part specific manuals / schematics, and my own past knowledge. I felt very efficient having been through the process before, and knowing when to lean on which resource. But the videos are by far the best resource to orient myself when first jumping in to the project, and thus save me a lot of time.
As someone who learns best by reading, I'm already at a disadvantage with video to begin with. To make it worse, instructional videos tend to omit a great deal of detail in the interest of time. Then when you add nonsense like you're pointing out, it makes the whole thing a frustrating and pointless activity.
Do you have an example of a video that does this? Seems like an interesting problem to solve.
I definitely write super long things when I consciously make the decision to not spend much time on something. Meanwhile, I've been working on a blog post for the better part of 2 years because it's too long, but doesn't cover everything I want to discuss. If you want people to retain the content, you have to pare it down to the essentials! This is hard work.
I think the situation that people run into is something like "how do I install a faucet" and they are getting someone who does it for a living explaining it for the first time. Explaining it for the first time is what makes it tough to make a good video. Then there are other things like "top 10 AskReddit threads that I feel like stealing from this week" and those are too long because they are just trying to get as much ad revenue as possible. The original comment was about howtos specifically, and I think you are likely to run into a lot of one-off channels in those cases.
Making a long video isn't like writing a rambling letter. It takes work to make 10 minutes of talk out of a 1-minute subject. And mega-popular influencers do this, not just newbs who haven't learned how to edit properly yet.
This is something that happens a lot. I'll Google a narrow technical question that can be answered in three lines of text--there's literally nothing more of value to say about it--and all the top hits are 5+ minute videos. That doesn't happen by accident.
You get no random ad content which just cut into the feed at will, which makes for a somewhat better experience. But there's the inevitable "NordVPN will guarantee your privacy", "<some service here which has no ads and was made by content creators so you don't have to look at ads if you subscribe but hey all our content is on YT but with ads and here is an ad>" ad.
There is no escape. I actually pay for YT premium and it's SO much better than being interrupted by ads for probiotic yoghurt or whatever. I know there are a couple of plugins out there which I have not tried (I think nosponsors is one of them) but I really don't think there is any escape from this stuff.
I want the most succinct result possible when I search the web.
We have no guidance telling publishers to get rid of "old" content. That's not something we've said. I shared this week that it is not something we recommend: https://twitter.com/searchliaison/status/1689018769782476800
This also documents the many times over the years we've also pushed back on this myth: https://www.seroundtable.com/google-dont-delete-older-helpfu...
There's no denying Google encourages long rambling nonsense over direct information
You're financially rewarding people for hiding information.
That includes self-assessment questions, including this:
"Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"
That's not telling people to write longer. Our systems are not designed to reward that. And we'll keep working to improve them.
I would assume they didn't answer this because the answer is either "No" because echo chamber or "Yes" but they don't want to say that publicly.
To be more explicit, yes, we're aware that there are complaints about the quality of search results. That's why we've been working in a variety of ways, as I indicated, to improve those.
We have continued to build our spam fighting systems, our core ranking systems, our systems to reward helpful content. We expanded our product reviews system to cover all types of reviews, as this explains: https://status.search.google.com/incidents/5XRfC46rorevFt8yN...
We regularly improve these systems, which we share about on this page: https://status.search.google.com/products/rGHU1u87FJnkP6W2Gw...
The work isn't stopping. You'll continue to see us revise these systems to address some of the concerns people have raised.
So it seems the search quality team exists but gets locked on in the closet by advertising periodically.
I know you can't verify anything directly but maybe we could set a system of code for you to communicate what's really happening...
If you think I am exaggerating, try to prompt him to see if you can get him to acknowledge that Googles current systems incentivize SEO spam. See if he passes the Turing test.
"your actions", give me a break. The parent commenter doesn't own Google, and you aren't forced to use the platform.
Are people on this site really convinced that an L3 Google engineer can flick the "Fix Google" switch on the search engine?
No, it's just when someone speaks on behalf of the company with the terms "we," they are typically addressed with "you." That doesn't mean we think they're the CEO. Are you unfamiliar with this concept? I can send you an SEO guide on it.
And anyways, the sentence "Take responsibility for the results of your actions like an adult" actually does imply he has some personal responsibility here. It's not helpful to the discussion and it's rude.
"a company doing dumb things"
"insult everyone's intelligence"
"lying to them"
That's a little hyperbolic, don't you think? Do you even hear yourself? I fully understand Google hate but directing it at one person who is literally just doing their job (and hasn't lied to anyone despite your allegation) is childish and counterproductive. Save that for Twitter.
Danny has been here since 2008. Your account was created in 2022.
And also, "people" aren't being rude, you are. Own your actions.
How long I've been here is irrelevant.
That wasn't me. Maybe pay closer attention?
>reward longer articles
The age of articles was being discussed, not article length. Maybe pay closer attention?
Actually... you know what, never mind.
You did that, not me. It's you who seem to be having a problem with understanding the thread.
I mean, saying that you should design pages for people rather than the search engine clearly hasn't shut down the SEO industry.
This is such terrible, low-quality, manipulative content. It does not belong on HN.
If this is what online publishing has come to we have seriously screwed up.
And if the content isn't shared somewhere (typically on a non-Google property), then is it even relevant anymore? All the Googlers who defined search relevance outside of freshness have left to other opportunities.
The issue here is very clearly with how Google et al are operating, effectively intentionally favoring blog spam over real content producers.
If the current UI is a reflection of what PMs at Google want users to do, then they don't really care if the user uses, or even knows about, the option to filter by time. So I don't think it flies to point at the user and say, "you're holding it wrong".
Just like it was A/B tested that the best background color to contrast the ad area against the organic results’ #ffffff is #fffffe.
I will not be surprised if in my lifetime it is not possible to Google anything by William Shakespeare.
Otherwise, there is the save page option in your browser if you do want to keep a local copy.
It looks very appealing, but I haven’t had a chance to try it myself just yet.
The approach to just move it to another domain and be done suggests it's not well enough researched for what an organization their size could manage.
Because it's clearly bullshit at this point.
Split that off, and search is all of a sudden producing a whole let less value, since the profiling value cannot be meaningfully extracted by other companies.
How come CNET is getting a free pass here to do something obnoxious because they're losing money, but Google is getting the stick for (maybe) doing something obnoxious to avoid losing money?
I think Google do a lot of bad stuff but let's at least assign blame proportionally.
Do you really not see the power dynamics at play here?
We might as well burn the libraries, they serve no modern purpose.
Storing and indexing and maintaining old content isn't free, in either dollars or environmental footprint.
It's pretty close to free.
Surely Google determines "fresh, relevant" content according to whatever has recently been published, which this doesn't change. If anything, doesn't Google consider sites with a long history of content with tons of inbound links as more authoritative and therefore higher-ranked?
This baffles me. It baffles me why this would be successful SEO -- and assuming that it actually isn't, it baffles me why CNET thinks it would be.
Updated on <?= date(m/d/Y) ?>Google's suggestion isn't to delete pages, but maybe mark some pages with a no index header.
https://developers.google.com/search/docs/crawling-indexing/...
Alternatively, other than ads, what is changing on a CNN article from 10 years ago? Why would that still be getting daily scans?
i am tracking rss feeds of many sites, and on some i get notifications for old articles because something irrelevant in the page changed.
Wikipedia had a long tail of low-value content, but even the low-value content tends to be among the highest value for its given focus. e.g., I don't know how many people search "Danish trade monopoly in Iceland", and the Wikipedia article on it isn't fantastic, but it's a pretty good start[0]. Good enough to serve up as the main snippet on Google.
[0] https://en.wikipedia.org/wiki/Danish_trade_monopoly_in_Icela...
They’re just truly useful pages, and that is reflected in how people interact with them.
That's for stuff like large e-commerce sites with constantly changing product info.
Google is clear that if your content doesn't change often (in the way that news articles don't), then crawl budget is irrelevant.
It doesn't have to fetch every article (statical sampling can give confidence intervals), and it doesn't have to fetch the full article: doing a "HEAD /" instead of a "GET /" will save on bandwidth, and throwing in ETag / If-Modified-Since / whatever headers can get the status of an article (200 versus 304 response) without bother with the full fetch.
There’s some methodology to trying to direct Google crawls to certain sections of the site first - but typically Google already has a lot of your URLs indexed and it’s just refreshing from that list.
It’s easy to change millions of pages once a week with on-load CMS features like content recommendations. Visit an old article and look at the related articles, most read, read this next, etc widgets around the page. They’ll be showing current content, which changes frequently even if the old article text itself does not.
Once a site has been indexed once, should it really be crawled again? Perhaps Google should search for RSS/Atom feeds on sites and poll those regularly for updates: that way they don't waste time doing to a site scrape multiple times.
Old(er) articles, once crawled, don't really have to be babysat. If Google wants to double-check that an already-crawled site hasn't changed too much, they can do a statistical sampling of random links on it using ETag / If-Modified-Since / whatever.
Just a guess though.
No need to invent a new system based on RSS/Atom, there is already an actually existing and in-use system based on SiteMap.
So, what you suggest is already happening -- or at least, the system is already there for it to happen. It's possible Google does not trust the last modified info given by site owners enough, or for other reasons does not use your suggested approach, I can't say.
https://developers.google.com/search/docs/crawling-indexing/...
If the content deleted is garbage, why wouldn't it help? No clue on CNET's overall quality, but I don't have a favorable image of it. Just had a look at their main page and that did not do it any favors.
Second is there could be spam links pointing to old CNET articles that need to be wiped from CNETs site spam score.
I know it doesn’t make sense and that Google says it is not necessary. But it clearly worked for us.
I think a fundamental truth about Google Search is that no one understands how it actually works anymore, including Google. They announce search algorithm updates with specific goals… and then silently roll out tweaks, more updates, etc. when the predicted effect doesn’t show up.
I think the idea that Google is in control and all the SEOs are just guessing, is wrong. I think it’s become a complex enough ML system that now all anyone can do is observe and adjust, including Google.
So it doesn't matter if the photographer or illustrator worked hard to make a beautiful image, Google's robotic decision based on some lifeless mathematical formula crossed out their efforts.
When it comes to Google, sufficiently-advanced malice is indistinguishable from incompetence.
Makes it incredibly difficult to have nice imagery above the fold at least.
Even if updates are mostly minor corrections or batch updates of boilerplate, the capability exists to rewrite any part or all of a story if there's no way to see past versions from third-parties from a single cache from archive.is or possibly archive.org.
Archival navigation and visualization should be deeply integrated into a user-centric, privacy browser.
~~I'm also worried about the deletion of old pages on archive because new owners of a domain update the robots.txt file to disallow it, which I've heard wipes the entire archive.org history of that domain. I hope that gets addressed.~~
Edit: this is no longer the case
https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
Including the part where they blamed anyone but themselves for hiding old snapshots when robots.txt changed?
It was possibly a worthwhile fight for someone to have, but not for the site that hosts the Wayback Machine. Separation of concerns, my friends...
Google's deteriorating performance shouldn't result in deleting valuable historical viewpoints, journalistic trends and research material just to raise your newly AI-generated sh1t to the top of the trash fire.
At this point, we should all realize that every ideal and slogan spouted by tech companies is just marketing. We were just too young and naive to know any better. 'Don't be evil'. In hindsight, that should have set off alarm bells.
In this utopic arrangement, users of search services are more-or-less evenly distributed among different search providers, enjoying a variety of different takes on how to find stuff.
Search providers, continues the sci-fi imagination, keep innovating and differentiating themselves to keep an edge over competition and please their users.
Producers of content, one the other hand, cannot assume much about what the inventive and aggressively competitive group of search providers will accentuate to please their users. So they focus on... improving the quality of their content, which is what they do best anyway.
Its a win-win for users and content producers. Alas, search service providers have to actually work for their money. This is slightly discomforting to a few, but not the end of the world.
One cannot but admire the imagination of such authors. What bizarre universes they keep inventing.
> Are you deleting content from your site because you somehow believe Google doesn't like "old" content? That's not a thing! Our guidance doesn't encourage this. Older content can still be helpful, too.
https://twitter.com/searchliaison/status/1689018769782476800
[1] https://www.businessinsider.com/cnet-slashes-10-staff-says-c...
FWIW, I'm as mad about Google quality going downhill as anyone's. The problem of sites doing stupid things for SEO purposes goes back at least 20 years though.
Source: used to professionally offer SEO services.
Red Ventures is trying to get CNET to be worse. with this and the AI written stories. Google should react by delisting all of CNET.
For all its resources it’s incapable of improving the situation. Since it only understands proxies of quality and truth, and these things can be manufactured and industrialized, they are incapable of winning.
Blogspam and fraud consistently outrank the original, and changes to the rules frustrate legitimate sites more than SEO spam. If anything these frequent changes increase the demand for SEO, not make it harder. You just wind up with more.
Until they figure out some kinda digital pesticide for SEO spam, the situation will continue to get worse.
Information on the internet should be in whatever format best suits the topic. The format that best serves the users looking for that information.
And search engines should learn to interpret that information and the various formats, in order to be able to best connect those searching information with those providing information. Yes, the search engine should adapt to the information and it's formats - not the other way round.
Instead we see "information" (or the AI-generated trite replacing it) adapt it's contents and format for search engines, in a bid to ultimately best serve advertisers. And search engines too adapt and change their algorithms to best serve advertisers.
As a result it becomes ever harder and harder for users to actually find the information they want in a format that works.
It's become so bad, that it's now more practicable to use advanced AI to filter out the actual information and re-format it, rather than go look for for it yourself.
As a human, you no longer want to use the web, and search... you want to have a bot that does that for you... because ultimately the space has become pretty hostile to humans.
https://www.seroundtable.com/google-dont-delete-older-helpfu...
Incentives are incentivizing.
I think older articles will eventually hold more weight in Google Searches. There will be a before OpenAI vs after OpenAI weighting.
> The page itself isn’t likely to rank well. Removing it might mean if you have a massive site that we’re better able to crawl other content on the site. But it doesn’t mean we go “oh, now the whole site is so much better” because of what happens with an individual page.
> “Just because Google says that deleting content in isolation doesn’t provide any SEO benefit, this isn’t always true" Which ... isn't what we said. We said that if people are deleting content just because they think old content is somehow bad that's -- again -- not a thing.
To me that reads like it is possible to improve the ranking of new content by deleting old content. The only thing they are refuting is that the age of the deleted content is the reason for the improvement.
With articles such as "The Best Home Deals from Urban Outsitters' Fall Forward Sale" currently gracing their front page, I'm wondering how long HN commenters expect to need access to this content.
Google presumably used AI to rank pages (it hasn’t been just PageRank for a while).
The AI has noticed people don’t engage with older content, so it deranks older content. It also deranks websites with lots of older content.
So websites pull their older content , which is an important form of historical memory.
Even if the AI isn’t actually doing this, people assume it is.
Because AIs aren’t rules based, we have to guess what it’s doing.
And we guess it’s deranking old sites.
Archive.org deserves all the support it needs. If only the Wayback Machine was actually indexed and searchable too...
https://web.archive.org/web/19991013034959/http://softseek.c...
Also, if you just dump all that content on archive.org, you're kind of just reaching into archive.org's wallet, pulling out dollar bills, and giving them to Google, whose ostensible goal was to index and make available all the world's information. I feel like that's enough irony and internet for today.
The times I have visited CNET pages in the past was to find specific information. If such information happens to be in deleted articles, that would reduce my interactions with CNET in the future.
I think they should archive the old articles or even offload them to Wayback willingly, but it's possible some of the articles they're purging aren't worthwhile. If I write up an article about a cat playing on a scratching post, there's a good chance there's nothing unique or valuable about it and it doesn't need to sit around gathering bit rot :p
'Stories slated to be “deprecated” are archived using the Internet Archive’s Wayback Machine, and authors are alerted at least 10 days in advance, according to the memo.'
Don't worry, Wayback is capable of telling CNET where to go if they were really concerned. If anything they should be thankful for the newly-generated interest that a major company would use them instead of some randoms cherry-picking sites for arguments.
You can also save them some space, it appears, by uploading a file on your website asking them not to retain it and then sending an email to them. See https://webmasters.stackexchange.com/a/128352
For now, I just blocked the bot, but when I have some time maybe I'll ask them to delete the data.
Hopefully, with a little community effort we can reduce the strain on them so they can spend their limited resources reasonably.
I had the same feeling about the Verge.
It feels like going to Britannica online to read about what Encarta encyclopaedia was and why it came in something called a CD-ROM.
---
E: Awesome: https://www.britannica.com/topic/Encarta
I’ve seen people use Britannica Online as ammunition in arguments on Wikipedia. “Britannica says/does X so Wikipedia should too”
Doesn’t really make things better, IMO, but they are at least doing that.
I'm not sure that CNet is relevant. I see their headlines regularly because they're a panel on an aggregate I visit. Most of what they publish is on a level with bot-built and affiliate pages.
We just had a long run of "Best internet providers in [US city]" as if people who live there could choose more than one. Between those will be Get This Deal On HP InkJet Printers and Best Back To School VPN Deals 2023 that compares 10 different Kape offerings.
If there is a point to cnet's existence, I truly don't see it.
I don't understand why they think I or anyone else won't just skip the part where they lard it up with ads and just talk to the robot directly.
My inner historian is screaming.
I’m sure there must be some way to fix this with META tags/etc, but it is often easier just to delete the old stuff, than change the META tags on 100s or 1000s of legacy doc pages
There was a defined process for customers to request access to it – they'd open a support ticket, tell us what they wanted, the support engineer would check the licensing system to confirm they were licensed for it, then transfer it to an externally accessible (S)FTP server for the customer to download, and provide them with the download link. There was a cron job which automatically deleted all files older than 30 days from the externally-accessible server.
So, while there may be some historical interest in how we were talking about, say, cloud computing in 2010, that's the kind of thing I keep in my personal files and generally wouldn't expect a company to keep searchable on its web site.
That 2017 product roadmap slide saying "we'll deliver X version N+1 by 2020" gets a bit embarrassing when 2023 rolls around and it is still nowhere to be seen. Maybe the real answer is "X wasn't making enough money so the new version was cancelled", but you don't want to publicly announce that (what will the press make of it?), you just hope everyone forgets you ever promised it. You can't get rid of copies of the slide deck your sales people emailed to customers, or you presented at a conference, or in the Internet Archive Wayback Machine – but at least you can nuke it off your own website. Reduces the odds of being publicly embarrassed by the whole thing.
There are infrastructure costs and manage costs.
All this SEO garbage may be responsible but I still think google could do a much better job.
Feels like Google is just over monetizing it’s search at this point and the quality of product is in a free fall.
Google definitely has an incentive to push content farms covered in Google ads to the top of its search engine, whereas a premium service's only incentive is to provide the best search engine possible.
Stop giving these crap companies power over you and your data. You don't know what they'll do with your data in the future (think George Orwell).
As we can see these companies can out of the blue force other companies to delete their old articles. There goes freedom of speech just to rank higher on a shi*y search engine that no one smart even uses (I use brave search). Then again, these companies kowtowing to Googl mainly write for the brainwashed sheep. I personally don't use any of their privacy invasive analytics on my blog.
Yes, things are becoming ephemeral. Probably good. Been like that forever except for some naive time window of strange expectations in the 2000s decade. We shouldn't keep rolling ahead of us a growing ball of useless stuff. Shed the the useless and keep the valuable. If nobody remembers otherwise, then probably it's not a thing of value.
So we need some preservation, but indiscriminate blind hoarding of all info isn't necessarily the best.
If you genuinely strive to create a great content that piece of content will drive traffic, leads, sales, help for rankings and etc. for years.
Just by searching a random thing like "how to build an engine" the first article that comes, and the top results is from 1997.
CNET used to, perhaps not be cool, but it certainly was a place you'd go get tech news from every once in a while. Just imagine they were huge enough to buy download.com, back then, even! For a while, they were essentially *the* place you'd go for downloads of various shareware stuff.
Now, it's a ghost town of a site with AI-written junk.
The saving grace here? The existence of the Wayback Machine. A non-profit by the Internet Archive that is severely underfunded. If you ever needed a reason to donate, this is probably it. And even then, the survival of this information depends on a singular platform. Digital historians of the future will have a tough job.
This seems antithetical to the internet approach of 'data wants to be free.'
Data most assuredly doesn't want to die. Why U kill data?
However, many of the links are no longer working. Some websites domain have expired, while others have chosen to remove the old articles. That made me feel so bad.
I did not know that this was a possible motivation as to why my more 'historic' work is disappearing, though.
0. Donate to internet archive now
1. Website operators should aim to not hide content or download artifacts behind JS with embedded absolute URIs or authenticated APIs.
I’ve found deleting the worst quality ads to be good for everyone. Maybe they are deleting the worst.
But then I thought about scenarios where this might be legitimate. If we assume:
1. Google has some ability to assess the value of content that isn't 100% reliable at a per-article level, but in aggregate is accurate.
2. Google therefore has the ability to judge CNET's content quality score in aggregate, and use that in search rankings for individual articles.
3. CNET knows that they have articles that rate well on Google's quality score, and articles that rate poorly.
4. CNET has reason to think they can generally distinguish "good" articles from "bad" articles.
THEN it would (could?) make sense to remove the "bad" articles in order to raise their aggregate quality score. It's kind of like the Laffer Curve of Google Content Quality. Overall traffic could go up if the "bad" articles are "bad" enough.That said, unless CNET was publishing absolutely useless garbage at some point, it would make a lot more sense to just de-index those articles, or even for Google to provide a way for a site to mark lower-quality articles, so Google understands that the overall quality score shouldn't be harmed by that article about the Top 5 Animals As Ranked By Hitler, but if someone actually Googles that, it's fine for CNET to have retained that article, and for Google to send them to it.
In short: Google advising sites not to do this when there are perfectly reasonable alternatives seems silly.
Very few people use it.
And it costs resources to keep up and crawl.
Better to prune content that doesn’t get a lot of traffic to keep your domain authority high.
Also 50k page limit on sitemap.xml.
Also… stop hoarding. Delete everything you can. As often as you can. Clinging to the past isn’t healthy.
Not all content should be immortal.
Best case that new domain starts ranking well.
It's a shame because CNET, surprisingly, has retained some really great editors that are as knowledgeable as they come in their respective domains. David Katzmaier and Brian Cooley come to mind.
GPT4 with WolframAlpha plugin even gave me enough information to implement Taylor polynomial approximation for Gaussian function (don't ask why I needed that), which would have otherwise taken me hours of studying if I could even solve it at all.
PS: GPT4 somehow knows even things that are really hard to find online. I recently needed standard error but not of mean but rather of standard deviation. GPT4 not only understood my vague query but gave me formula that is really hard to find online even if you already know the keywords. I know it's hard to find, because I went ahead to double-check ChatGPT's answer via search.
i find it most helpful when i am not sure how to phrase a query so that direct search would find something. but i also found that in at least half the cases the answer is incomplete or even wrong.
the last one i remember explained in the text what functions or settings i could use in the text, but the code example that it presented did not do what the text suggested. it really drove home the point that these are just haphazardly assembled responses that sometimes get things right by pure chance.
with questions like yours i would be very careful to verify that the solution is actually correct.
Good luck when you need to update it and adjust it - this is the equivalent than copying/pasting a function from Stack Overflow.
But now I'm curious!
And you face the same problem when looking for something in a domain you are not an expert in - no way to tell if a web page is truthful and no way to tell if ChatGPT is right. ChatGPT just lets you make more mistakes more efficiently.
But for those cases where you kind of know the answer, ChatGPT is usually better than search.
With a regular search engine you've already found your answer by that time.