LinkedIn sitemap.xml
linkedin.com
linkedin.com
My Netbook is still trying to open this over dialup, and xPUD has ground to a halt.
Size of file: 8.2mb
Transfer time over ~10Mbit broadband: 18.24s
Firefox onload (from GET): 49.21s
Chrome onload: N/A, still hung at 100% CPU after 15 minutes (Chrome attempts to style raw XML, whereas Firefox just makes a foldable tree like IE)
... I recommend:
curl -s http://www.linkedin.com/sitemap.xml | lessCanary kept on chugging at the expense of the whole OS hanging.
You can use a sitemap index [0] to split a sitemap into multiple ones. This is beneficial because you can change part of it and not have to have google re-crawl everything, since index updates happen right after a whole sitemap has been crawled. As a result, changes you make to a smaller sitemap get indexed way faster.
I've switched to grouping my bulk pages into sitemaps alphabetically, and then putting important ones (front page, about) into different sitemaps, and specific landing pages into yet another different one.
Edit: "If you want to list more than 50,000 URLs, you must create multiple Sitemap files." According to 'icebraining' they have 47785 urls.
xpath -q -e '//*/loc/text()' profiles-sitemap.xml | xargs -n 1 wgethttp://thepiratebay.se/torrent/7031839/35_Million_Google_Pro...
It's frustrating, at least for me, that the legality here is still so gray (at least IMO it is).
Much of their content is likely in the public domain (facts / basic non-creative information) although there is definitely plenty that is not; the lack of 'black-and-white' is what frustrates me...
Clickwrap - http://en.wikipedia.org/wiki/Clickwrap
$ curl http://www.linkedin.com/sitemap.xml | grep -c '<url>'
47785
I wonder why were these selected. Probably the most searched. user@/tmp > grep lastmod sitemap.xml | cut -d - -f 1 | uniq -c
47785 <lastmod>2006
They were probably sitemaping all their users when they started the service, then they have decided to stop shortly after.In case you were looking for me: