That's for stuff like large e-commerce sites with constantly changing product info.
Google is clear that if your content doesn't change often (in the way that news articles don't), then crawl budget is irrelevant.
That's for stuff like large e-commerce sites with constantly changing product info.
Google is clear that if your content doesn't change often (in the way that news articles don't), then crawl budget is irrelevant.
It’s easy to change millions of pages once a week with on-load CMS features like content recommendations. Visit an old article and look at the related articles, most read, read this next, etc widgets around the page. They’ll be showing current content, which changes frequently even if the old article text itself does not.
There’s some methodology to trying to direct Google crawls to certain sections of the site first - but typically Google already has a lot of your URLs indexed and it’s just refreshing from that list.
It doesn't have to fetch every article (statical sampling can give confidence intervals), and it doesn't have to fetch the full article: doing a "HEAD /" instead of a "GET /" will save on bandwidth, and throwing in ETag / If-Modified-Since / whatever headers can get the status of an article (200 versus 304 response) without bother with the full fetch.