I have implemented rss 2.0 parser faster then understanding the atom specification. Atom can do encode stuff like encode html inline the xml instead of as a CDATA string. In theory this sounds great, but is ends up in a big mess of complexity (e.g. a blogpost with handwritten invalid html).
These days there is also JSONFeed which is really easy to parse, simple and flexible, but it is not supported everywhere yet.
https://www.rssboard.org/rss-2-0-1-rv-6#hrelementsOfLtitemgt
> An item may also be complete in itself, if so, the description contains the text (entity-encoded HTML is allowed; see examples)
Note "is allowed", not "is required". This caused SO MANY problems back in the day, because the spec didn't clarify if you should or should not include HTML in that element - and there was no way of telling, when parsing a feed, if the author was in the "entity-encoded HTML" or "YOLO and just stick plain text in there" camp.
IIRC, Atom came about precisely because the RSS specifications didn't provide the level of detail needed for a spec to be truly interoperable.
But I think it’s worth noting that a cultural tradition emerged that papered over the flawed spec. I think that is actually pretty common with specs, even if the rss2 one is extra loose.
Maybe having a correct spec isn’t everything.
Back in 2013 a developer of a feed crawler wrote a selection of things people get wrong with their feeds.
https://inessential.com/2013/03/18/brians_stupid_feed_tricks
(This particular one isn't papering over flaws in the spec, many of thse are advising against doing things that violate either RSS or XML spec, or are subjective opinions additive to the spec (e.g. always have a datetime). But ya this is basically what I mean.)
I don't know what you're doing that RSS 2.0 is somehow faster to parse than Atom. I've written parsers for both over the past twenty years with a negligible difference between the two besides the fact that the RSS feeds often need hacks. I've also wrote a whole bunch of blog and linkblog backends that produce Atom feeds, and have never and issue with any. Let's look at the required elements of an entry: updated, title, id. Nothing remotely onerous there. In fact, it's purposely minimal, more minimal than RSS. And in RSS 2.0, title is a required element (because if something it's explicitly noted as optional in the RSS 2.0 spec, it's assumed to be optional).
In my personal linklog, I use the title of the target page of the link as the title, because it's the sensible option. With tweets, you have half a point. Only half a point, because title is required, but Twitter also post-dates the early 2000s considerably. But here's the thing: 'title' is required in RSS and Atom, but there's nothing saying it can't be empty. I know, I've blown your mind!
And then there's JSONFeed, which, of course, can somehow gracefully cope with people dropping '"' in random parts of the file because people generate JSON files like that by hand, right?
Right?
Just like they write RSS and Atom feeds, right?
Right?
The same thing can also happen in RSS feeds (and JSON Feeds): Entity-encoded HTML strings or CDATA HTML strings do not have any guarantee of well-formed-ness. The direct embedding of XHTML into Atom as namespaced elements just surfaces potential invalid markup higher up.
I wrote a podcast validator, and I don't think that's true — every RSS feed must be "well-formed" XML.
(Note that all "valid" XML documents are "well-formed", but "well-formed" XML documents are not necessarily "valid".)
In a perfect world people would construct their XML documents with an API which guarantees that the generated serialisation is a well-formed XML document. E.g. the API guarantees that the element tree is nested, that namespaces are declared and that the serialiser escapes any text nodes. Then people could add their well-formed XHTML fragments as a child to <atom:content type="xhtml"> and then serialise the whole document, guaranteeing well-formed-ness across namespaces.
In practice people have a tagsoup string from their data store which they concatenate inside their RSS template in <description>. If you’re lucky, they replace "<" and "&" beforehand or do the CDATA thing. But in XML terms that is just a string, not well-formed markup.
As long as you don’t need to consume feeds, just use atom.
The less TL is that RSS1 and RSS2 are basically two different branches of the original:
- Netscape released RSS 0.90 as an RDF application (RSS literally stood for RDF Site Summary) - RSS 1.0 was an update / direct evolution of RSS 0.90 by a dedicated working group using final RDF 1.0 semantics (as RSS 0.90 had been based on an earlier working draft) - RSS 1.1 an evolution of RSS 1.0 by unrelated people
This is called the RDF branch, for obvious reasons.
However a few months after RSS 0.90 Netscape also released RSS 0.91, which dropped RDF entirely, rebranded to “Really Simple Syndication”, and added some elements from Userland’s own syndication format.
This is the start of the “Harvard” (formerly “Userland”) branch, Userland / Dave Winer released his own variant of 0.91 (timeline with netscape has never been super clear to me), then went on to release 0.92 with an <enclosure> element, followed by 0.93 and 0.94. He then released RSS 2.0 to mark a bit of a compatibility break, as RSS 2.0 adds namespace support and removes some elements from his 0.9, and also to fuck with the RSS WG’s 1.0 release.
Because the Harvard branch was the first to support enclosures (embedding audio) and Userland had built support for that, it became the de-facto format for podcast feeds, Atom also supports enclosures but I’m not sure any podcast client (or podcasting source) supports them.
Are there specific things that come to mind as specifically annoying?
Nah, you should just use Atom.
The only vaguely good reason to use RSS is spotty Atom support in podcast apps.
So, better formulation: only bother with RSS if you're a podcast.