Google Bot Attempts to Crawl Shortest URLs First
russell.ballestrini.net
russell.ballestrini.net
Edit: this would also encourage webmasters to use short URLs, which benefits users by being easier to remember, too.
We're unlikely to ever know the answer unless someone from google explains it to us.
why would you crawl the site in any other way?
I could see crawling pages most likely to have changed first as those pages would most likely lead to fresh content.
Anyway, good point, that deserves more testing to extract some conclussions
I can definitely see the local/relative effects being a natural consequence of prioritizing by pagerank, but the global part sounds more like a separate signal.
Does anyone have insights?
it's not because of sitemap or because of url structure or because of dynamic content.
Mine were blog articles in the same format. This is how it was crawled:
sitename.com/year/mo/day/stub
sitename.com/year/mo/day/stub-one
sitename.com/year/mo/day/stub-one-two
sitename.com/year/mo/day/stub-one-two-three
sitename.com/year/mo/day/stub-one-two-three-four
You can also ask in #seo on irc.freenode.net I know there are knowledgable SEO people in there who might be able to provide you with a decent answer.
I find it is also likely that short URLs (especially in the case of a directory-type site) are seen first by the spider and that this order is respected by the crawler (FIFO).
You can also ask in #seo on irc.freenode.net I know there are knowledgable SEO people in there who might be able to provide you with a decent answer.
Most people posting here are looking for some sort of deep meaning in this when IMO it is more likely just due to a localized side-effect of doing something such as storing the urls in a trie-like structure and then iterating over it breadth-first.