Deep dive into finding RSS feeds
lighthouseapp.io
lighthouseapp.io
* Yes, I know the article talks about the RSS icon, i'm just soapboxing.
RSS is great but it has one great flaw in that it doesn't scale that well by itself. If 2 million people subscribe to your feed and try to update it once an hour, that is 48 million requests a day just for RSS.
What does work well (and how things have evolved) is to have a service that polls RSS on behalf of its users. This was the beauty of Google Reader but plenty of replacements exist.
Vivaldi browser still does that.
The readers often cache the result for all users of the same feed. The only exception are readers that are running 100% locally. In that case using Etag or Last-Modified will make each request cheap and manageble.
Who's updating their RSS feeds once an hour 24 hours a day?
Anyway, a very popular site with that many millions of visitors already has to handle extreme traffic, regardless of RSS.
By default, security rules are on, but I can disable security rules programmatically for a hostname too.
fwiw, I once got a ddos on a host running their pages product, and got no charge for it. it also stayed up and didn't give captcha pages
RSS is a pull-type system, no?. So the end-user is causing the overload problem by hitting the publisher’s RSS feed every hour? The problem is out of the hands of the publisher…
Yes.
The web in general is also a pull-type system.
> So the end-user is causing the overload problem by hitting the publisher’s RSS feed every hour?
There is no overload problem.
And again, this math is off, because an end user is not even awake 24 hours a day: "If 2 million people subscribe to your feed and try to update it once an hour, that is 48 million requests a day just for RSS."
Is that wasteful? Does that make me a poor Internet citizen?
I'd never considered using RSS feeds from an app and pulling it directly. I have too many devices for that to be a great workflow.
Normally I'd say I was an outlier, but my gut feeling on this is that the people still using RSS and people self-hosting would overlap in a big way.
I personally don't think it's a problem. As I said, "a very popular site with that many millions of visitors already has to handle extreme traffic, regardless of RSS."
> I'd never considered using RSS feeds from an app and pulling it directly. I have too many devices for that to be a great workflow.
A lot of apps have sync to handle that workflow.
> my gut feeling on this is that the people still using RSS and people self-hosting would overlap in a big way.
My gut feeling is that most RSS users, including myself, simply use client apps.
Sorry, I wasn't thinking about the end user literally pushing a button. I was reading into the discussion here what I imagined to be a gap in the overall design philosophy of RSS where the end-user, using their own app or preferred methods, poll all of their favorite websites selecting among common defaults (Select the frequency you like to poll this RSS feed: weekly? daily? hourly? hehe minutely? secondly?).
In this scenario the user doesn't need to be awake 24 hours a day. They just need to use software which stupidly hits someone's server like they're doing a denial of service attack. As your comment alludes no one needs to be that up-to-date.
In contrast, there was another solution: "...service that polls RSS on behalf of its users."
Again, reading the discussion I imagined the wild west of the internet has inexpert users choosing unsane defaults and overloading small self-hosting content providers. Compared to third-party aggregators who are trying to make money by solving a "Tragedy of the Commons" problem.
But the user's computer would need to be awake 24 hours a day, which is usually not the case.
And once per hour is not a DoS attack.
I use the Awesome RSS add-on for that in Firefox: https://addons.mozilla.org/en-US/firefox/addon/awesome-rss/
How long after Google Reader was canceled did FF remove RSS?
Firefox removed RSS support in version 64 on December 11, 2018, approximately 5 years and 5 months later.
And there are always CDNs if you really have a large, global audience.
I'm still continually astounded that the tech community appeared to absolutely swallow Google's utter bullshit answers about why they shut down Reader; that really marked a turning point.
I follow mostly RSS on non technology website, for instance road cycling. people that wouldn't care or know about RSS because they are not very techy, yet because they are normies that use WordPress for all their website it puts a page with RSS feed automatically. You got to find it with developer tool by searching RSS but 99% of the time if it's WordPress it got RSS.
Thank you WordPress you bloated piece of shit :)
I've been adding to my feeds.opml since reddit started dying in ~2015 and now I'm up to around ~1700 feeds and mostly independent from aggregators; though I still collect new feeds from HN/IRC/etc. Mostly I just always make a point to look for them whenever I read something cool on the web.
[1] https://chromewebstore.google.com/detail/rss-subscription-ex...
This causes the following error: TypeError: URL constructor: //matthew.science/posts/riscv/ is not a valid URL.
This was back before the Web became the one true way and is the reason it uses 2 slashes, to distinguish protocol local from site local.
Note: Hoarder can automatically hoard RSS feeds as part of its 'bookmark everything' functionality. Hoarder uses AI to tag all the content (URLs, feeds, images, notes) so you can then do full text searches on your personal archive of your bookmarks etc.
You might consider adding some of the key info from the first paragraphs in the docs (open source, self hosting) to the front page, above the fold. Github-link might imply it, but I was scanning for “open source” with my eyes and was initially disappointed not to find it and ready to dismiss the product right out of the gate.
What does your setup look like?
The browser as we now know it is mostly a static application that has long lost its user-centric mission. Websites might push some stuff but the user must do thinks manually. Its primary function is to provide a search window to external search. People even stopped using bookmarks and search for everything.
This hypothetical RSS-Browser could become the main organizational tool for the users web experience, integrating the use of bookmarks.
In fact even more "feeds" could be integrated like email and activitypub or atproto posts. It boils down to the fact that each person has a number of profiles/roles and within each they have a taxonomy of interests and we need a tool that integrates static and dynamic sources of information.
Turns out the feed finder couldn't find the feeds even though I've linked to them using clickable RSS icons.
I didn't know about the autodiscovery feature so I'll add that now.
https://github.com/begriffs/findrss
The combinations came from what I observed in the big list of blogs I follow. The script works pretty well for most sites.
The problem with the approach presented here is speed. Most of the web pages, especially smaller are really slow.
Crawling most of the web pages is pain, especially if you use selenium and small SBC.
Therefore either the page presents a clean nice RSS link, or get lost.
Most of the good, modern pages give you nice RSS. Even GitHub gives you RSS for commits.
For other pages I try openRSS.
For YouTube I use yt-dlp to obtain channel id, to establish RSS.
Algorithm is crude, but gets the job done.
https://github.com/rumca-js/Django-link-archive/blob/main/rs...
Or I suppose you could just find all "Content-type: application/rss+xml" in CC.
I know in the past, when I was looking for large lists of RSS feeds, I didn't really find what I was looking for.