Built to Last – RSS, HTTP (2015)
blog.theoldreader.com
blog.theoldreader.com
For the corporate web RSS and HTTP are dead. But as a non-corporate human person I'll be sure to keep RSS and HTTP alive on my webservers.
I think as the nostalgia/retro factor kicks in, it will become cool again to support "Any Browser".
A few common questions about this are:
Security exploits: Yes, you do have to do this on a safe network, but at least over where I am, the Internet is safe enough for me to visit un-malicious websites, perhaps using a VM.
The difficulty of it: It's moderately difficult, but quite doable, with certain constraints combined with progressive enhancement.
(HTTP/1.0 browsers are slightly more difficult to support, because you need a dedicated IP address, as there is no Host: header yet. Netscape 1.x, IE 1.x and 2.x, Mosaic are quite usable if you can get this set up.)
HTTP-over-TLS works fine across all "versions" of HTTP, HTTPS has been supported by browsers since 1994; over half a decade before RSS was even a thing.
Your criticisms of HTTPS seem more targeted at HTTP3, but that's actually a separate topic (HTTP3 may be HTTPS-only in practice but that's not the specific aspect you're critiquing).
RSS/Atom also work fine over HTTPS.
It hasn't received patches in 10 years, so having it connected to the internet at all is a bigger problem than it's lack of tls1.3 support.
General-purpose browsing is a whole different matter. I wouldn't advise anyone to do that unless on a patched, up to date and supported machine.
Using a pre-set list of servers with many sites is not easy, because they'd use CDNs and such.
So basically you can only use it for your 'own' websites and then you can control crypto on server side.
A working MITM TLS stripping proxy deployed on a trusted server would be a better solution, but such options were complicated to deploy.
Either way though, I do personally think it's fairly reasonable for server admins to maintain webservers that target "General-purpose browsing" without also having to include support for odd individual long-tail edge-cases of people running restrictively firewalled instances of outdated insecure OSes.
> I do personally think it's fairly reasonable for server admins to maintain webservers that target "General-purpose browsing" without also having to include support for odd individual long-tail edge-cases
That's a good point
Many enjoy and find comforting the older OSes they're used to.
Research even shows that reliving sensory experiences from a younger age can have a measurable positive effect on a person's physiology and health.
I think the general trend of sacrificing everything to so-called security, which in reality just means being owned by the root certificate holders, all several hundred of them, and all their employees and affiliates, even for low-risk and low value targets, like someone who wants to blog with IE3, is overrated and stupid and will be disposed off once we feel more sure of the security of our underlying network...
HTTP/1x != Windows XP.
Advocating for older specs doesn't mean advocating for old insecure implementations.
If we imagine the HTTP+plaintext at one extreme, and HTTP+TLSv$latest at the other, a whole bunch of specs in the middle have effectively died out because nobody deploys them any more.
Therefore, it's unfair to claim that "HTTP-over-TLS works fine across all versions of HTTP" (because it doesn't, unless you also support modern crypto), or that "HTTPS has been supported by browsers since 1994" (because the 1994 HTTPS is not even partly forwards-compatible with the 2022 HTTPS).
E.g. today SSL/TLS almost universally means TLS 1.2 or 1.3. The SSL specs are no longer accepted by practically any server.
Not all older specs are worth advocating for.
Security is unfortunately an arms race, which means that broadly speaking software-updates are important for anything network-connected. Using old hardware is thankfully still possible, but the idea of expecting a 12-year-old piece of software to securely connect you to the internet today is disconnected from the reality of the threat landscape we live in.
HTTPS does not even work on devices from 5 years ago.
After non-removeable batteries, SSL/TLS implementations have been the single largest headache when you want to keep using a device for more than a couple of years.
Expecting network-connected clients to work securely without software updates for even a few months these days though isn't really possible, but that's down to the externalities of the world and has nothing to do with spec design.
One may argue that the problem is essentially unavoidable because the alternative is untenable (delivering any sort of communication without both encryption and authentication to ensure the veracity of the data) but that doesn't make the observation any less true. https is in fact a pain in the ass and introduces a lot of overhead and grief and shortened the useful life of countless things, relative to the time before https.
This is one of those many cases on the internet where people make an unfounded statement and then just move on as if it's accepted as true. How does the mechanism not change the fact???
The mechanism in this case is outdated software; upgrading or changing the software on the device makes the device operable (you absolutely can install modern TLS on Windows XP, alternatively you could also install a modern OS on most devices that came with Windows XP). That fact seems to make the mechanism extremely relevant in whether the device is operable.
Compare HTTP. Or even IPv4, which still works to this day, with the same software & hardware combination.
The only aspect of HTTPS that isn't backward compatible in the way you're talking about is, effectively, the algorithms keeping up with modern attack vectors. That's not an aspect of HTTPS the spec, it's an aspect of the arms race in network security. So pinning this issue on HTTPS is odd.
The reason you can't use outdated software to connect to modern servers isn't due to protocol design, it's due to threat actors.
What other definition there is? A protocol that has absolutely no backward compatibility, not even with itself, and by design, is "built to last"? Or that a protocol that is frequently rewriting its core algorithms in its base specs is "built to last" ?
> That's not an aspect of HTTPS the spec, it's an aspect of the arms race in network security.
It is an aspect of the spec. There is literally no HTTPS without SSL/TLS. If there was, you could have an argument there, but: the very definition of HTTPS is "HTTP + SSL".
Really, there is no color. Ethernet, IP, HTTP still allow an unmodified client from 25 years ago to work with practically full functionality. HTTPS does not. Whether this is because HTTPS has more ambitious goals (which it obviously has), or whether the designers were smoking something better, or because computer security requires an arms race (that's another debate), the effective result is that HTTPS is not built to last.
Everything you build on top of it will require continuous rearchitecturing every couple of years. And I'm not talking about changing some certificates or the like. I am talking spec changes. I'm talking most of the algorithms having changed significantly.
Or maybe TLS 1.3 will be the good one this time. Who knows.
This is largely not true, if you want to talk to today's servers.
Unlike HTTP, SSL is not cross-compatible with older versions.
I would be happy to be proven wrong.
And that is about as mainstream a device as you can get.
HTTP/2? Sure HTTP/1.1? Not a chance. It's too widespread and too valuable, and the floor is too low. Unless Google starts a singular effort against it and manages to pull in the other big cos while avoiding the ire of the liberal states, which seems unlikely.
It'd be like saying Firefox is deprecating JSON because it doesn't have a nice viewer for it...
IRC and USENET were built to last. Names in either network weren't tied to anyone but the collective "network." Neither network gets used much today, since names aren't tied to anyone in particular. It turns out that globally writable data stores are great vehicles for spam and fraud.
Content-addressed systems like BitTorrent and I2P can theoretically maintain content availability for as long as anybody wants to keep it available, not just whoever originally published it. BitTorrent is also pretty secure, but it's not truly fair since it's an immutable data store, and all the spam and fraud is just outsourced to HTTP instead of being eliminated entirely.
I'd argue that this is probably as good as we can get. Even if something is built-to-last, a person can still throw something away (and indeed many people find good stuff in dumpsters!) If nobody is willing to pay the cost to actually host something, then it goes "into the trash" so to speak. I realize that physical goods last even in the trash, but that's not something a digital good can ever be unless it's serialized to some form of storage.
VPRI (of Alan Kay fame) has published an interesting paper [1] trying to create a Cuneiform Tablet of code.
[1]: The Cuneiform Tablets of 2015 - http://www.vpri.org/pdf/tr2015004_cuneiform.pdf
1. The protocols you list are not strictly married to DNS in any way. e.g. https://<ip-address> works fine. Granted it's not very practical to use right now but there's nothing in the HTTP spec. preventing things working without DNS.
2. DNS isn't technically married to ICANN either. There are alternative domain name assignment proposals and systems that use the same protocols without ICANN. Again, not super practical today but theoretically possible to use.
So the fact that these systems are all so loosely coupled makes them pretty resilient and very much "built to last". Moving away from ICANN may be hampered by the inertia of a gargantuan network effect, but that's not any bigger than the inertia of moving off HTTP completely.
What is really funny about that sentence is that the word "http" links to the wikipedia article through href.li, in order to hide the referer, a built-in feature of HTTP 1.1 and the web. So much for elegance and simplicity when I am staring at a workaround using an external system for something as simple as a link.
There's a paid hosting version that's only $15/year.
I personally run tt-rss and I quite like it precisely because it's light, fast, has good keyboard shortcuts, and has a solid Android app.
I used to use Feedly and it's also an excellent option if you're not into hosting your own infrastructure.
Unfortunately neither meet your Win/Linux requirement, but just pointing out there are some excellent options out there.
There's little or no keyboard support at the moment, but if you're interested in helping to either develop or make some suggestions, I'd be up for improving it! Plugin: https://addons.mozilla.org/en-US/firefox/addon/brook-feed-re... Source: https://outgoing.prod.mozaws.net/v1/08a450f6d21388f6bfc51f2b...
[1]: https://feedbin.com/
I don't really like in-browser but could tolerate it if it could do all of the above
I use ReadKit on the Mac (has filters, but local to the client) and Reeder on iOS. I am using FeedWrangler as the back end; its filter language (at least as documented) is unfortunately inadequate for my needs. Otherwise it's been fine.
QuiteRSS (Win,Mac,Linux), RSSOwl (Win,Mac,Linux), RSS Guard (Win,Mac,Linux), Ravens.js (electron), TT-RSS (self-hosting), FreshRSS (self-hosting), LifeRea (Linux only), Arss (macOS).
If you are fine with self-hosting, then I recommend getting FreshRSS over TT-RSS and FreshRSS have support for RSS Bridge which will make FreshRSS a powerful tool. I prefers FreshRSS because I can set it up in my XAMP easily, TT-RSS moved to Docker only. Also FreshRSS provide their RSS feed API for other RSS software to hook into it. So you can use mobile app with FreshRSS feed directly to it.
Now if only I could find the perfect Android RSS client to talk to it. EasyRSS is the closest so far, but it has some bugs that occasionally annoy me.
Works great, fast, clean, simple. Highly recommended.
Disjoint set of platforms from the ones you mention in your second sentence, though.
1. Inconsistent implementation of standards: The implemented versions of RSS and Atom out there makes parsing more of an art than just throwing a library at it as no RSS parsing library out there can handle all the edge cases. (The last time I did a test a few years ago in an hour long sample from one of the RSS firehoses out there pulled up hundreds of custom namespaces and tag names.)
2. XML formatting: Consistently formatted, well formed XML is never 100%, even from by major news organizations. Embedded CDATA means parsing content is a quagmire of double escaping.
3. Inconsistent content: A RSS feed could just have the last few items that have been updated, with just titles or links, or it could be literally all of the content of a blog, jammed into some 20MB+ text file, double escaped and simply enlarged after every new update.
4. Inconsistent unique identifiers and HTTP header responses: Many sites will respond appropriately to requests with a 304 if there are no changes. Many will not. Many sites will give each RSS item a globally unique identifier, many will not. This forces every Reader to simply request the whole doc over and over again, comparing unique items with a blend of logic and magic.
5. Inconsistent support: Most sites that use RSS have no business model attached to it, so it's just sort of an afterthought and may be shut down at any time, and often is.
All this leads to: Massive amounts of wasted bandwidth as bots poll endlessly for updates, wasted processing time parsing unformatted or badly formatted content, wasted storage because of bad IDs and URLS, wasted effort on the user's part dealing with the inevitable errors, and wasted effort on the admin side dealing with an antiquated tech that should have gone away with MySpace.
RSS should be scrapped. Killed. Replaced. Forgotten.
It all depends on what you want to use it for. For my use case, I built a feed reader that just needs to know the titles, URLS, and publication dates of articles. That let me build an RSS/Atom reader that lets me curate my own news feeds and delegates everything else to the normal browser.
Are there inconsistencies and broken feeds? Sure a few, but for my purposes, I can ignore 99% of that. Is it wasteful? Sure, it. could be more efficient, but honestly downloading React and a hundred dependencies every time I visit a webpage is as well.
Fear not though, many sites built with tools like Gatsby aren't including RSS feeds, so your wishes may still come true.
[1] https://www.solberg.is/sveltekit-blog [2] https://github.com/jokull/blog/blob/master/src/routes/feed.x...
Fun fact on #3 that most people don’t know: Atom supports paginated feeds: <https://datatracker.ietf.org/doc/html/rfc5023#section-10.1>. No idea what library support is like. I admit it was defined in the AtomPub spec, but it should still apply to regular Atom feeds too. My own website’s feed is approaching half a megabyte with all my content ever; I’d like to make it paginated and only include the most recent ones in the first page (while still satisfying my unjustified thirst for the feed to still technically contain all the content), but as long as I’m using someone else’s static site generator that doesn’t already support that I’m probably not going to get round to implementing it.
is this some elaborate parody
also BGP...