A sysadmin's rant about feed readers and crawlers
rachelbythebay.com
rachelbythebay.com
As a somewhat orthogonal observation, I've found it's surprisingly hard to write a crawler that's well behaved with consistent heuristics across a variety of different feed providers.
Usage of published vs pubdate vs updated, which is changed when, (or all three just containing 1970 and/or the current time), reordered feeds, items published out of order _regardless_ of the scheme used, changing urls/ids, etc. Whatever set of heuristics one uses for some sites may not apply to others. Now, this begs the question of "why not make parameterize and tune", and yes, this is largely what I've resolved to, but it's, to the core point, more of a moving target than one would expect from how ostensibly simple RSS is.
At the end of the day, in many cases I just fall back to using a cache of recently-seen-urls and, when possible, short-circuit the enumeration when I cross one. (Similar disclaimer, I do love me some RSS, I've just never had a good opportunity to rant about this.)
Her points make absolute sense for how RSS and Atom are supposed to be used (and probably are used by her and other's who don't monetize content) but are absolutely not how the (monetized content) sites like Reuters use it.
Granted, updating every 10s is really stupid. The only way I could imagine that being sane(ish) is using RSS as a generic protocol for plumbing automated data movement rather than for fetching actual blog posts. Daily is easily fine for 99% of feeds that are not themselves aggregators.