RSS mistakes: let's not make them again
blog.superfeedr.com
blog.superfeedr.com
It is not made clear why this is desirable. The author makes no case for why a JSON-based system would be any better than the current XML system. Perhaps the implied "eww XML", but XML is the standard.
(But yes, I agree that "turn this already-existing thing into JSON just because" is a stupid trend.)
The first issue is that you have a field, and that field is rendered in a text box, and is defined to be text; at some point you go "man, I wish I could add a hyperlink to my text", and so you now want to put some HTML in there. However, the field is just text, so what do you do?
Let's say this was JSON, and this text is a string; do you put HTML in the string? This is somewhat equivalent to putting escaped HTML into the text node of the XML document. Alternatively, you could replace the string with an object (vaguely equivalent to putting HTML elements in the RSS).
The result of some people choosing the first option (which seems more reasonable at first glance than the second option, as it provides a better experience for existing readers and better fits the existing protocol) is having to look at a string and guess whether it should be parsed as HTML or not.
The second problem is that as people make changes in the second direction, if they are not carefully organized and centralized in a specification, you end up with haphazard and incompatible changes: someone decides to add a "type" field with a mime type, someone else adds a different element.
Having a ton of people generate the format, and having the thing parsed by some liberal canonical language, leads to too much flexibility in fields like dates: someone puts a weird date format into their date field, and it is parsed correctly, and then tons of other people do it, and you're screwed.
This also isn't helped by going with JSON: the field is just more likely to end up being "anything that a JavaScript Date object is willing to convert from a string to a date", which assuredly supports irritating corner cases that are not supported by the Date parsers from other random languages.
With RSS, it was really "the peoples' protocol", with the specifications only encoding random changes that had become popular over time. What we were left with was a total mess: I can't find it now, but someone once wrote a proof that you couldn't actually parse RSS due to conflicting standards.
That's extremely debatable.
XML is quite loose, unless I use plist theres no data types which is both an advantage and a disadvantage.
Note that I am not, here, endorsing JSON over other alternatives; the same is obviously true of XML, HTML, YAML, or lots of other potential formats.
On sites that provide a feed link, I just click that link, which opens that feed in feedly, and I hit add to my feedly. Process is as simple as twitter. The choice of some sites to hide their feeds and provide no links is the burden of site designers.
I too sometimes have to resort to the tactics that the article describes to get a feed, but it is usually just wordpress sites where the author is unaware that rss even exists or that rss is provided by wordpress.
Here's a Chrome plugin to add it back: https://chrome.google.com/webstore/detail/rss-subscription-e...
Chrome, even when it had the icon, by default linked it to the unformatted XML page. I get the sense that Google's heart was never really in RSS.
You can put the icon back in the toolbar, or you can press on the globe/keylock near the url, press on more informations, and then Firefox will show all the available feeds for the page under the "feeds" tab.
I think also however that this post brings up a lot of good points. Specifically I feel it touches on a very important one subtly... we do have now a "web of data" even if it isn't exactly the W3C spec of such a thing, we still have one. Now what?
Yes, polling has issues, but the author doesn't explain how changing the file format from RSS to some JSON-based format makes those problems go away.
Maybe I am just dense, but neither does the author show real problems, nor does he offer solutions. Well, the latter being somewhat unsurprising considering the former.
Um, less solutions to this exist, but they are much simpler than 'PubSubHubbub or RSSCloud'. And don't require any changes to RSS.
You are requesting this thing over HTTP. You (and the server delivering the feed) simply need to use standard HTTP caching headers. etags, last-modified, etc.
No need to fetch the XML and parse it and diff it.
From an etag in the header, you can tell if the current remote resource is identical to the one you have locally or not, without any fetching/parsing/diffing.
Same with from a Last-Modified header, right?
In fact, you send a conditional GET rather than first retrieving only a HEAD and then deciding whether to continue with a full GET, but either way you aren't fetching/parsing/diffing to tell if there's new content.
Take https://kouio.com for example (a Google Reader replacement I've built), in kouio you just enter a website's address and it'll discover the feed automagically.
Well as jerf mentioned below, HN doesn't correctly provide a link reference to its RSS feed. Now that's fine in our case, as a surprisingly large number of websites fail this test, and simply have a visible link with "rss" or "atom" in either the link's text or URL.
Now we pick these up, but in this case the front page of HN has an invalid link with "rss" in the text, namely this very thread itself. So normally the HN front page URL works fine, but right now it doesn't. This gives us a great chance to refine our discovery code a bit, so thank you!
<link rel="alternate" type="application/rss+xml"
title="Hacker News"
href="http://news.ycombinator.com/rss" />
was added to the <head> of the homepage, the discovery would probably work better.Wait, do feed readers really not do a HEAD request first and check Last-Modified ?
Not at all. If you have the RSS as a static file - and that's the best way for a resource that is much more often read than written - the web server can just stat() the file and get the modified timestamp instead of having to open and read it.
But there's probably a billion edge cases that make it too difficult.
That's a bad idea if you cache RSS in an xml file, i.e /rss.xml. It would make it more complex if you had to handle query parameters. The parser should just check for the version attribute on the rss node; <rss version="2.0">. Or, even better, it could just check for the availability of server nodes such as "<cloud>".
Or did you mean something completely different?
Every new feed reader on the block shocks me when I have to enter the exact URL to the feed (e.g., domain/xml/feed.xml, or what have you).