RSS Feed Best Practices
kevincox.ca
kevincox.ca
A pet peeve of mine, is where some enterprising web developer has either built a theme, or hacked an existing one and removed or not added the <link> tag for the RSS feed, when the site engine in question DOES have a working RSS feed.
You see this a lot on customised WordPress sites. I end up having to try all variants of feed URLs I can think of until I find the feed. Surprisingly, most non-bespoke site engines still have working feeds, so the hit rate is quite high.
I also miss the days of browsers auto-extracting the <link> tag and showing an RSS icon when a feed is found.
.../rss , .../rss.xml , .../.rss , .../rss_full.xml , .../feed , .../rss-feed , .../feed/all/ , .../MySection.xml , .../MySection.atom , feedserver.example.com/section/index ...We have a huge internal document of how RSS elements map to UI in every player and all the gotchas, and we still discover new things years after initial development.
So everyone goes ahead and makes their dialect of RSS and pushes that. Apple did this with iTunes and so did YouTube.
What clearly needs to happen is that RSS needs to have a refresh. The common uses cases need to be integrated into the standard.
From another comment:
>The problem is that there are huge swaths of use cases that are not covered by the current rss standard.
Have you considered turning that document into a RFC to establish a standard?
This is considerably more correct, robust and featureful than any others I’ve seen, just because I could (I don’t even use most of the stuff I’ve implemented for it, though some like enclosures I have draft content using). It handles Atom’s text constructs (where content can be provided in XHTML, HTML or text format), and the common attributes xml:lang and xml:base, none of which I’ve seen any other stylesheet handle correctly.
I also produced an RSS variant, since you basically have to use the inferior RSS for podcasts. That’s not in use on my site at this time, so I’ve dropped a copy at https://temp.chrismorgan.info/2022-05-10-rss.xsl for now. But please, use Atom. RSS is a lousy format that causes some genuine problems and should have been completely retired fifteen years ago, and the only thing that actually needs it still is podcasts.
If you wanted an informal list of people making use of this pretty-feed.xsl specifically, it would be possible to solicit that sort of thing through GitHub and see if others come out of the woodwork to identify themselves.
See also https://discovery.thirdplace.no/?q=kevincox.ca for feed discovery. Disclaimer: I built it.
Am I missing something or is this site missing one?
If you just paste the page into your feed reader it should find the feed for you.
<link href=/feed.atom rel=alternate title="Blog Posts" type=application/atom+xml>
instead of <link href=../../../../feed.atom rel=alternate title="Blog Posts" type=application/atom+xml>
?And of course none of this matters much if you have a non-IPNS reference because the feed will never change making it mostly useless.
https://validator.w3.org/feed/check.cgi?url=https%3A%2F%2Fke...
https://feedmail.org/subscriptions/new?url=https%3A%2F%2Fkev...
However I noticed that this site also has trouble: https://discovery.thirdplace.no/?url=&q=https%3A%2F%2Fkevinc...
Which site? For the submitted one, you have to search for 'atom' in the source markup; for HN, for 'rss'.
--
kevincox.ca
<link href=../../../../feed.atom rel=alternate title="Blog Posts" type=application/atom+xml>
HN <link rel="alternate" type="application/rss+xml" title="RSS" href="rss">A nice example on how to do it is The Guardian. Basically any category page you can visit is also a feed.
Agree that RSS is imperative; don’t agree about excluding JSON feeds. One can do both.
In any event, it requires virtually zero extra effort after initial setup, so I see no reason not to offer both. Different strokes...
RSS Feed Best Practices
10 | RSS Feed Best Practices ~~~
37 | RSS Feed Best Practices ~~~~
Many already have the comment count on hand and it's not really respam in your combined feed as you get fresh data at the old position. The ~ just lets you search for hottest things when you are pressed for time as most search boxes require at least 3 chars.A simple proposal would be something like <comments count="23" last="TIMESTAMP"> and <score>72</score>. Then the reader could re-surface these items to you if they pass particular thresholds.
Of course there are some details such as ensuring that entries with recently updated comment counts appear back on the first page (or that old pages are scanned for updated entries).
Also I think WebSub doesn't work here because it excludes entries that have been seen before (which won't be an issue for your original proposal).
The goal of rss is for me that everything I am interested in should come to me. This breaks down the moment a high frequency feed (such as news) is in the mix, which is why I don't use it (outside podcasts).
- Videos. (Basically the same use case as blogs though)
- Releases of various projects. (GitHub for example supports this)
- FeedBurner has a feed for problems detected in your feeds. (Although I don't use FeedBurner anymore)
- Hacker News posts on the front page with >400 points.
- A handful of Reddit searches.
- My Reddit inbox.
- A feed of WebMentions to my blog.
- A feed of packages that I maintain for nixpkgs that are out-of-date.
Really anything that someone may want to be notified about can be a useful feed. For example I can imagine a feed of price changes for a product so people can wait for it to go on sale.
The search feeds can be expensive for the site if they aren't careful. But between "materialized" and caching it can be made pretty cheap. If you are careful you can even make these work with WebSub. (If you run your own hub you can know what the subscriptions are and either run them periodically or actually check the queries against new items in real-time)
As every wp site (that didn't switch it off) has a feed there it is more reliable than favicon.ico you will also find lots of feeds there not linked any place. I think many authors are not even aware they have it.
I like it as you don't have to visit the domain, you can simply point a feed reader at a list of (slightly modified) domain names and discover many precious shiny things.
I have some complex logic to find a hero image that is sometimes in a feed entry outside the HTML, and prepend it to the top of the article -- but only if there is not another reference to that image in the article. It is complicated because sometimes there will be hero-image.jpeg in the article and hero-image-1200x500.jpeg outside the article, and I have to judge whether the images are the same.
I am in favor of embedding the image into the HTML, and not adding it elsewhere.
You can also use enclosures but not all readers will display these. Notably WordPress puts every embedded image into an enclosure so a lot of places just ignore them to avoid duplicates.
In RSS the main media attachment is the <enclosure> element. However like a lot of things in RSS it is poorly specified. For example it isn't even clear if you can have multiple. If you have multiple are they different representations of the same resource or different resources?
https://validator.w3.org/feed/docs/rss2.html#ltenclosuregtSu...
The solution to this is the Media RSS spec. (Which is also commonly used in Atom. XML namespaces are actually kinda nice.) This is a lot better if only because they contain the <group> element which makes it clear what different things mean. (Different representations or different resources.) Unfortunately support is still varied and it isn't always clear if this is just a copy of every image in the post (WordPress...), only semantically relevant images, or images that are the point of the entry but don't appear in the post.
https://www.rssboard.org/media-rss
So unfortunately the best path forward is probably yet another standard that makes these clear from the outset, and hoping that people actually follow the rules. Not an easy path forward.
Again, I'm not an expert on media in feeds, but based on the feeds I have seen trying to write a feed reader I would do the following if I was creating a feed with media in it.
1. Use Media RSS. Don't use <enclosure> at all unless you are a podcast.
2. Only reference interesting images directly. You are welcome to include whatever you want in the HTML for stylistic reasons but don't put reference it outside of the HTML unless it is actually an interesting image.
3. Put every media element in a group. Basically because WordPress puts everything into a media element and it is unclear if these are the same images, duplicates or even interesting at all (I don't need the author's avatar listed in the media elements WordPress...). Even if you only have one representation put it in a group just because it is unambiguous what you mean.
If anyone has more experience here I would love to hear what you think.
Nothing stopping a reader from fetching the linked resource and pulling useful stuff out of the existing Open Graph annotations, of course...
I'd like to get into News aggregation with RSS again now that I think there are some significant population who think Social Media may not be the best way to consume news. So I made a cursory exploration of the status of RSS feed support for major news outlets.
To my surprise, Many do still support RSS incl. those with paywalls. It's surprising because AFAIK Google doesn't mandate RSS to index their news, Apps like Feedly which started out as pure RSS aggregator went onto HTML parsing of major news websites by the time I abandoned my RSS feed aggregation apps.
Are the news outlets still supporting RSS because?
1. Legacy systems which depend upon RSS.
2. There have been a rise in demand for RSS feeds.
3. It's the right thing to do, for journalism.
4. None of the above.