Style your RSS feed
darekkay.com
darekkay.com
But there's a gotcha:
This works fine if you serve the feed with `content-type: text/xml`, because with that content type the browser typically renders the result in-browser. But `text/xml` is technically not the correct content type for RSS feeds, `application/rss+xml` is; and when you serve it with that content type, browsers typically either open a download window instead of rendering the feed, or they render it as plain text without styling.
So you're stuck. Have a styled feed but serve it with the wrong content type, or be technically correct and serve it with the right content type, but no styling for you.
Practically, content type really doesn't matter that much, and most (all?) RSS readers are fine with `text/xml`. But for those of us who like to be technically correct...
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8
Which means it would prefer to get a response from the server in (basically) the order shown, but it will accept any response type (*/*). So ideally the server would be using that information and making the decision to serve the RSS feed as application/xml instead of application/rss+xml. AFAIK if the "subtype" (rss+xml in this case) has a + in it, then it basically means it conforms to the format after the + (e.g. application/xml) and thus is just providing a bit more context, but is still valid as the more generic version.Though I do think the browser should also keep that same thing in mind and attempt to render any application/whatever+xml in its XML renderer.
Edit: another thing you could try to add this response header:
Content-Disposition: inlineWith RSS you have to make the request periodically to stay up to date, so I'd say you trade some of that temporal tracking for extra geotracking.
That isn't to say any claims that feeds are almost a thing of the past holds any water. There are almost certainly more feeds now than at any point before. The second derivative is probably a flat line, though, which is what I imagine is behind the perception of deathliness. Feed consumption, on the other hand, is almost certainly down—but only if you exclude podcasts.
The first few times it didn't pick up a feed, I would inspect the page source, and sure enough, no `link` tag in the head for the feed (nor references to "rss" or "atom" or "feed" anywhere in the page).
I don't know what stack they use, but I find this particularly common for company blogs.
It's important to note that Firefox is all but legally owned by Google at this point. It exists on Google's dollar so Google can point at it during anti-trust lawsuits.
I liked the overall approach of a declarative transformation, but XSLT is absolutely awful and the lack of an alternative that's supported by browser made me not do it again.
It's probably still like that in the backend.
I was always amazed how worked.
What an underappreciated feat!
In a way modern React etc. frontends can be seen as a "fix", in the sense that at least it means most sites effectively have APIs, whether or not they officially expose them to users, but without the pain of XSLT. If they support server side rendering as well, they're getting close to what we were doing.
And that was the original driver for the site design I mentioned - every URL was an API endpoint with a well defined set of expectations, letting you effectively explore the API by browsing the site as normal until you had sliced and diced the data the way you wanted and then just add a parameter to get it in the format you wanted (just as you can with Reddit - e.g. append ".json" or ".rss"). And I try to do that as much as possible still.
Before the "XSLT adventure" I wrote a big webapp in C++, and we actually had a C++ rendering pipeline that had a component model to render not all that different to an initial React server-side rendering, combined with a pipeline where custom tags could be registered to further rewrite the output. I really liked that apart from being C++ and lacking the ability to optionally to client-side rendering.
It's because the browser renders the <xmp> content itself first. I used to have that set to display:none which prevented the flash, but HN complained about not being able to see anything with Javascript turned off :/.
Not to be confused with Symfony framework.
The same way that DTDs are very easy to work with... until you get a complex one created by a lot of people.
XML is so incredibly full of horrible decisions that it's not funny. And they are never on the macro level of "this format is useless", they are always in the details.
To be more charitable I think the problem has always been tooling. There's not a lot of good design tools that have supported XSL stylesheets as output. You've always had to do a lot of manual editing to get make a decent stylesheet and do any complicated functions.
What's annoying is XSL can be used for any XML documents like you said. Your "pages" could just be serialized database entries with a stylesheet on them. Since all the styling was done client side the server is really just an API server.
This was a design conceit of ATOM. It was mostly used for syndication but it was meant to be an XML API interface. It supported posting to a server in the ATOM XML definition as well as pulling. The idea being you'd come across a server with an ATOM API and be able to post comments or whatever to it right from a browser, no HTML form or JavaScript required.
RSS is alive and well and mainstream with podcasts.
Also: how many folks use RSS(-as-protocol) and how many use Atom (RSS-as-feed)?
What's interesting is that this particular post had been posted by someone on HN before I've shared it anywhere. So it got through RSS right onto the HN front page. This is what I meant with "thriving, especially among tech users".
Like you I tried to figure out how many of those are actual users to have a sense of how many people still use it and the numbers still don't make any sense to me.
That said, what do you think is the benefit of styling an RSS page? The point of the rss is to be consumed via an RSS reader so I don't see the benefit of having styles on a page like that. But maybe there's something I'm missing here.
Instead, you can have a readable page with a message like the one in the post: "This is an RSS feed. Subscribe by copying the URL from the address bar into your newsreader. Visit About Feeds to learn more and get started. It’s free."
Also, and not sure why this is happening, I have the Reeder extension installed on safari and it usually automatically prompts me to open an rss link in the app but it doesn't happen in your case.
I can't even use the button in the browser toolbar to open the feed in the app directly. The only way for me to add it is to manually copy paste the url.
Regarding Reeder, I've included the feed in the markup, so I'm not sure what is going on. I will definitely have a look, thanks for the finding!
I'm not sure which you meant here—is your perspective that RSS is thriving (judging from your traffic) or nearly dead?
There are two very similar English idioms that have opposite meanings: "all but dead" means basically dead, "anything but dead" means not even a little bit dead.
But again, appreciate the correction.
Sure, but I would assume that "RSS is dead" isn't so much about consumer awareness of RSS, but rather about the technology being invisibly supported and used in the infrastructure they make use of (through podcast clients et al) such that content producers still care about publishing through RSS.
Like, if I said "TLS1.0 is dead", I wouldn't be referring to greenfield projects not using it any more; but rather to the fact that the browsers and libraries that invisibly use it (and nothing newer) are themselves not being used by anyone any more — so nobody on the content-delivery end has to think about supporting TLS1.0 clients any more, and can drop any code that was supporting that.
But that's definitely not the case with RSS. There are still many important systems that depend on consuming RSS feeds. So there are people who still need to support RSS. So it's alive.
Of course Spotify has incentive to prioritize you listening in their app, but they can still put ads into the audio feed so it doesn’t matter as much to them. And your state-owned radio is absolutely limiting their audience by not using the accepted standards for podcast content.
I would not be surprised if Spotify themselves get the audio from the source via RSS/Atom.
No doubt the podcast creator may have to create an account and tell Spotify manually (and also on Apple, Google, etc), but once the feed is in their system the creators/producers just need to update things in one place and all the 'distribution points' get the new episode.
Spotify (and other big players) are doing their best to end this.
As many as one in four podcast listeners currently use spotify to listen. The growing number of spotify-only podcasts are not available via RSS. Most of the listened-to podcasts may also be available as RSS, but the listeners wouldn't know either way. Spotify's goal is to be the chokepoint for podcasts.
Spotify isn't the only one. The way to make money on podcasts is by being the distribution platform, and by locking people in with non-open standards.
But there are also spotify-only podcasts, which are not distributed to spotify like that, are not available via RSS or any other method but a spotify client.
This is an intentional business move by spotify. They are attempting to change the podcast landscape.
Although googling for recent sources, it looks like there's been some pushback and slight retreat -- looks like their plan wasn't exactly working. Yet. I'm sure they haven't given up yet:
https://www.theverge.com/2023/4/18/23688644/spotify-podcast-...
> After the cancellations and resulting layoffs, members of the Gimlet union blamed Spotify’s exclusivity strategy for disappointing numbers. Although the shows were not behind a paywall (free subscribers to Spotify could access them, as well), they did not enjoy the kind of wide distribution that the shows did before the acquisition. It’s not like they don’t have a big platform — according to a study by Cumulus and Signal Hill, Spotify is tied with YouTube as the most-used podcasting platform. But even then, it only has about one-fifth of the market.
And if you are typical, apparently the gimlet staff were right to be worried that this hurt their numbers!
You have to wonder. What would the world look like if more publishers had gone the route of styling RSS or Atom feeds, and maybe supported and extended the relevant standards in the places they found those standards to be deficient? Could we have ended up with a world where content delivery was all RSS, the relationship was exclusively between you and the publisher, and we didn't need Meta as the middle-man sucking publisher profits dry while convincing our daughters to kill themselves?
...Nahhhhh, I'm sure that going full neanderthal, RSS LOOK SCARY, clubbing it over the head and removing it from a website is the better approach. /snark
Sure, you can use anything you would on a regular HTML page. I was consuming my local news website via their RSS feed in my browser, as it looked like a regular website (but without all the fluff). Unfortunately, they've dropped the custom view completely, and it's now back to raw XML content :(
But we threw out XML for JSON. With JSON we need loads of custom, client side code to turn it into a DOM that the user can look at. With XML we only need XSLT. It won't work for all cases, but the majority of sites wouldn't need a single line of JS to renders sites. Yet here we are: shadow DOM, event listeners, useeffect, JSX, progressive hydration, and so forth and so on. To build web-experiences that we could deliver back in early 2000 but were deemed too complex and too daunting.
To be clear: there are lots of things that XML+XSLT cannot do, but which JavaScript+HTML can do (in an HTTP context).
But for the most typical HTTP usecase: a website presenting information, XML+XSLT is fully up to the task, yet we forego that, and instead pull out the big, complex "guns", entire JS frameworks. A knee-jerk reaction that I blame on the bad rep XML got, and the praise it's incomplete but simpler replacement, JSON, got.
But yeah, JSON is a pain. XML was pretty smart.
AFAIK it's the only major implementation of this technique. Most other big sites that provided an RSS feed didn't bother, and most of those RSS feeds are dead now. The BBC one has hardly changed since those days and it still works really well as a dual-delivery system.
That said, on the RSS as a whole, for most websites (mostly publications I suppose), adding RSS feed is a set-and-forget thing.
Sure, it's great for readers as it allows them to read all the content from one reading app but, on the other hand, the publication can't provide a custom experience, or - what's most important for some publications - display ads. That's likely why RSS is slowly moving into obscurity, at least IMHO.
Sure. But a title-only RSS feed seems like a very reasonable compromise.
And it's so much better than what everyone seems to have moved to, which is email notification.
I'm happy with newsboat[1]; but I'm not surprised that people have integrated scraping into RSS readers.
Fundamentally, that's not a problem with RSS, that's a war between scrapers and content providers. If the email newsletter model persists long enough, I'd expect that people will come out with "newsletter readers" that scrape websites too.
I'm not sure there's a good long-term solution to the problem. Aside from constant vigilance (obfuscation).
---
Earning a living from ads for these types of people.is not sustainable. They will have adblocker installed, so you are also wasting respurces with no income.
RSS or newsletter or whatever scrapper, it's there today already.
Issue is that Google Ads and such don't offer this and you don't get the "typical" ad networks.
If RSS were more popular there wouldn't be a problem to build the required tooling.
Even when feed readers don't send cookies etc. while fetching you can do a permanent redirect to a feed with an unique ID in the URL and most feed readers will store that URL, thus you can do tracking (incl. personalizing URLs in the feed) and all that.
What saves reader privacy currently is the small user base.
On the other hand, I don't necessarily dislike the email model (especially when having dedicated address just for newsletters) but it's much less comfortable than RSS and quite limiting from the publishing standpoint (given how far behind emails are behind modern Web).
Personally I still try to have RSS in all my content websites (with full content) as I'm primarily thinking about reach rather than ads. The ones that I haven't implemented RSS for yet, are just because it was more challenging or required more effort when integrating with e.g. CMS or something.
There is a world of difference between sites maintaining those feeds vs. nothing at all. Having a signal that an article is there with even the slightest amount of context is so much better than the alternative.
It's mildly annoying to visit a site for the full text (with your ad blocker), but signing up for newsletters etc. from each site one-by-one is really painful.
If I'm going to take 5+ minutes to really read something, it's OK to visit the site. That means something is interesting or relevant enough to invest my time in. That signal can usually be gleaned from a title and short paragraph. Compared to the number of new things published every day, it's relatively rare to find things worth those 5+ minutes.
From what I can gather, many people use RSS readers to follow 5-10 feeds, and they slowly look through and read most of the articles. It serves as a convenient way to follow their top few sites and maybe a few aggregators like HN. Other people track 100s of feeds and quickly scan what's happening, only diving into something if it's interesting or important.
I'm building a service for the second type of person (mainly because I'm that type of person, TBH). No idea what the ratio of "completionists" vs. "scanners" is. Having title-only feeds is not ideal for the latter group, but it's usually fine.
I’m definitely interested in a (hopefully FOSS) service for us “scanners”. I average around 300~ articles in my RSS daily, and I’m always hungry for more information. Though I should probably see about re-organizing it all so I’m not as consistently overwhelmed.
I want to pair that with an (opt-in) service for syncing feed subscriptions and handing off a stream of things worth archiving (it's often not urgent that this happens, you just want to make sure it happens soon so the content is not lost in a few years). That service could be a very low cost monthly subscription thing plus a FOSS option that you could run on a cheap VPS, etc.
24/7 services are also essential for generating notifications when something in a filter is spotted, being able to have an email gateway, doing things like POSTing items to other sites automatically, etc.
Making that "full" service FOSS is not in the near term plans, though. This is a distributed system that has run 100s of millions of jobs already, has a very specific security and monitoring setup, uses a number of queues and databases, etc. From my past experience, it's really hard to support people with on-premises distributed systems software like this (FOSS or not). I couldn't do this part alone (bootstrapping and can't afford to hire anyone yet).
I need to read the things in my feed, or maybe skip, or maybe save for later and maybe come back to it.
It's good to know that all I miss is a xml-stylesheet. I'm going to implement that once I'm free.
Good times
Making an RSS file look good for a human is nice, but it isn't for a human to look at.
I would say don't waste your time. Just get something off the shelf and start tweaking that one instead.
Not even sure RSS readers like Thunderbird have features to deal with duplicates like that.
The main time this happens is when people switch website backend or site generator or whatever, and don’t take care to keep the same IDs. In practical terms, IDs are just opaque strings, but they’re supposed to be IRIs (… even though you mustn’t assume it can be dereferenced) and universally unique, so a common approach is to use the page’s URL, which could lead to something like this if you change it:
<link href="https://example.com/current-title/" type="text/html"/>
<id>https://example.com/original-title/</id>
Some systems avoid this by using a different form of ID, e.g. WordPress uses its internal post IDs: <link href="https://example.com/current-title/" type="text/html"/>
<id>https://example.com/?p=12345</id>
It doesn’t need to be an https: URL, either; you’ll sometimes come across UUIDs: <link href="https://example.com/current-title/" type="text/html"/>
<id>urn:uuid:01234567-89ab-cdef-0123-456789abcdef</id>
As a concrete example of changing things and keeping old things working, here’s an excerpt from the template for my Atom feeds: <link href="{{ page.permalink }}" type="text/html"/>
{%- if page.year < 2019 %}
<id>{{ "http://chrismorgan.info/blog/" ~ page.slug ~ ".html" }}</id>
{%- else %}
<id>{{ page.permalink }}</id>
{%- endif %}tag:chrismorgan.info,2019-01-01:blog/slug
But the real problem, of course, is people caring: You'll need to store the ID with the content and continue using them when moving CMSs or domains. People, apart from your notable exception, don’t do that.
(Sorry for minting an example tag URI in your authority! I shouldn’t have done that according to the RFC.)
[RFC 4151] https://www.rfc-editor.org/rfc/rfc4151.html
Here's the code that does the work:
https://github.com/samuelclay/NewsBlur/blob/master/apps/rss_...
The common practice is separating the the Human-readable interface (HTML) from the Machine-readable interface (XML) and build both of them on top of some source layer(e.g. a CMS or Markdown files)
XSLT allows you to use XML as the source layer and build a presentation layer on top of that. Admittedly, the styling workflow can be cumbersome but you get a browser-supported templating language that seems Turing complete (anyone confirms?) and with the power of CSS, you can pretty much get a full featured static site generator that runs on the client!
I've seen styled RSS once in a while and never thought about it until now. I'd love to give this technique a try. It feels like one of those Semantic Web ideas that should have taken off decades ago.
A corollary: if you have a small blog and you'd like people to notice your new posts (without you promoting them on social media), make sure it supports RSS!
I'd love to explore how we could do more with this as well as I see so much value but like it has been pointed out here, RSS is set and forget for many people and therefore isn't in the majority of user's minds. Also the utility of having content in a feed is undervalued IMHO.
Incidentally, given this behaviour your response should include `Vary: Accept`. This is actually messing Firefox up a bit: open the document, it loads the HTML, View Source, it renders the source of the Atom, loading it from the cache according to the dev tools, not sure how it got there, force reload and you get the source of the HTML. Your server is also not handling HEAD requests, but improperly responding 405 to them.
Example: https://www.ghacks.net/wp-content/uploads/2019/04/rss-feed-p...
Edit: Nope, extension: https://code.guido-berhoerster.org/addons/firefox-addons/fee...
I have this installed: https://code.guido-berhoerster.org/addons/firefox-addons/fee...
I know it's not in style—today it's all about exposing data as JSON and then templating it with something like Web Components. Meanwhile this is stuff that XML and XSLT has been doing for over a decade.
I cut my teeth on HTML5 and JSON APIs from 2010 and beyond, so I never really had to work with XML.
I’ve never felt wholly comfortable with the mixing of semantic markup and data with the presentational elements required for styling. While advancements in CSS have improved the situation, I really like the boundaries demonstrated here and the inversion of of priorities: ie. we fetch the data and then we consider presentation rather than downloading the presentational layer which then in turn fetches the data.
I know there is a lot of baggage around XML, but this aspect of it really does seem superior.
...as well as on my Atom feed: https://jordemort.dev/atom.xml
...using the same XSLT stylesheet for both: https://github.com/jordemort/jordemort.github.io/blob/main/p...
https://github.com/org/repo/tags.atom
https://github.com/org/repo/releases.atom
For example: https://github.com/php/php-src/tags.atom
https://github.com/textpattern/textpattern/releases.atom
Chuck a bunch of those into your reader of choice, it works really well.Custom XSLT could be a fun feature for a feed reader.
Update: problem might be that browser sniffing for UTF-8 won't work, and that the doc might have multiple (RSS) title elements where the browser expects one (!) in head content
https://www.askdavetaylor.com/3-blog-pics/safari-4-view-all-...
Would love to get that back!
I don't count how many times I've seen people attempting to create simple reports from some XML taken from some obscure back office and try to reinvent the wheel by doing some JSON conversion and process that from some other program that build up an html report
My XSL stylesheet is rather fancy in what it supports. Sometimes because I actually use that fanciness (e.g. supporting HTML in all text constructs, not just <atom:content>, because I use HTML titles), sometimes just for the sake of correctness or completeness (e.g. passing through xml:lang as a lang attribute, or more obscure elements like atom:rights), and sometimes for things that I might use some day but haven’t yet (e.g. turning audio/video enclosures into <audio>/<video> elements, for podcasts). https://chrismorgan.info/atom.xsl is for Atom, and I’ve also written an RSS variant of it, purely for use in podcasts since podcasts are stupidly stuck with RSS (and it’s almost all Apple’s fault and I hate it): https://temp.chrismorgan.info/2022-05-10-rss.xsl.
Developing using XSLT is an interesting throwback to how web development used to be, because this stuff is basically frozen as it was twenty years ago. Almost all errors are fatal, diagnostics vary between nonexistent and poor (remember when you basically had to bisect to figure out what precisely broke things?), and the dev tools you’ve grown used to can’t help you when anything goes wrong. Also there are probably more bugs than there used to be. And documentation is bad (a lot of load-bearing functionality is completely undocumented). And browsers are much more inconsistent in their behaviour than you’ve grown used to (welcome back to the days of enforced trial and error, and having to check things in multiple browsers). (You get some of the same weirdnesses if you try shipping HTML using XML syntax, but it’s mostly XSLT-specific.)
Some particular issues:
• The stylesheet needs client-side JavaScript to render any type="html" content (HTML serialised as text), because Firefox doesn’t support the (admittedly optional) disable-output-escaping="yes" XSLT feature. (I’m puzzled by this lack even when using <xsl:output method="html"/>—and yes, Firefox and Chromium do both produce HTML-syntax documents in this case, not XML-syntax documents. I check this by searching for an xmlns attribute an all elements’ outerHTML, not sure if there’s a more direct check and I’d love it if someone knew one, ’cos this stuff regularly matters in JavaScript libraries—loads of libraries will break if run in an XML-syntax document due to assumptions made that no longer hold. OK, enough of this long parenthesis.) Therefore I recommend using XML syntax for the HTML instead (type="xhtml" in Atom, no equivalent in RSS because RSS is a disaster), and expect to do so for whatever next refresh I do of my own website.
• The stylesheet needs client-side JavaScript to resolve URLs if you use xml:base, since browsers just don’t support that at all any more (if they ever did?).
• If you make an error in the XSLT, Firefox will often point to exactly where the error is, but just as often give a completely useless generic old-school XML parse error screen. It validates declared XSLT 1.0 stylesheets, and is all round not very forgiving of errors, which is good, except that I wish the error reporting was more reliable.
• Chromium is more forgiving of errors, mostly for the worse, blithely ignoring some XSLT 1.0 errors that you really wish it would point out because you have certainly done something wrong. On errors it won’t overlook, it will leave you with an document empty save for an xml-stylesheet ProcessingInstruction, not put anything in the dev tools console, and print the error to stderr (and with no stack trace in the XML or XSLT, which is painful)—I’m mildly surprised with myself that I even thought to run it in a terminal to see if there was any output there, but I guess I was getting desperate one time when everything was working fine in Firefox but not in Chromium.
• In Firefox, pages that use XSLT will sometimes just hang during loading (seems to happen very often, maybe even always?, when it’s the first navigation in a new tab). Haven’t ever filed a bug about this. My wild guess is that this probably started happening at some point in the switch to a multi-process architecture.
• If you reload the page in Firefox while the dev tools are open, the dev tools are now effectively dead until you close and reopen them. Bear in mind also that the dev tools operate on the post-transform document, not on the source XML. Haven’t ever filed a bug about this.
If you try working on this stuff, you’ll want to keep a copy of the XSLT 1.0 and XPath 1.0 specs open. They’re the best documentation you’re likely to find, and honestly pretty good. Because each is a single document, you can search through for keywords nice and easily.
I’ve been very tempted to try shipping a website where pages are Atom documents (the list, a feed document, and individual pages entry documents), mostly just for fun, but I’d want to inspect browser support more carefully before doing this, and I’d need to nail down that Firefox hang first too.
And try putting xml:base on the atom:content tag.
I’m not sure what you mean. xml:base is supposed to work everywhere, for whatever is the appropriate scope, and my stylesheet copies the xml:base attribute across, and then includes JavaScript to make it more or less work, wherever you use it, since browsers don’t support xml:base at all any more.