After that introduction I was hoping Winds would be some kind of proxy that creates an RSS for sites that don't have one. But well, it's just another RSS reader. Those have actually never gone way, I'm using Inoreader every day.
After that introduction I was hoping Winds would be some kind of proxy that creates an RSS for sites that don't have one. But well, it's just another RSS reader. Those have actually never gone way, I'm using Inoreader every day.
Custom feeds can even be created for dynamic content, utilizing Chrome for full-rendering, and many other tweaks & techniques under the hood for seamless & scalable indexing.
https://www.kill-the-newsletter.com/
(I think a fellow HN user made it? Can't remember)
Note that many/most WordPress RSS feeds aren't that useful because they only show a snippet rather than the entire post. This dreary state of affairs is due in part to the fact that "checking" an RSS just means downloading the entire file containing all posts the author would like to make public, regardless of how many are new to the reader. This insanely unnecessary bandwidth usage penalizes sites that have long (>10) feeds with complete posts. (The single RSS file on my blog takes up the majority of my bandwidth costs.)
RSS needs to be replaced, not revived.
WRT scraping, it is getting harder because funny websites w/ no static content where everything is generated via Angular or Perpendicular or sth. are really hard to deal with. Recently my uni switched from an army of Wordpress websites to an homegrown AJAX-MVC-Reactive abomination where the links are reimplemented via some funny black magic where the actual link items (which are not anchors btw) don't have encoded in them the link targets, but only an onclick event that knows somehow where to go. And because they just killed all the RSS feeds, I wrote up sth. to revive them for me via phantomjs, but could not figure how to find the link targets, I cannot link the RSS items to anywhere but the main announcements page, and I can't add any description from the link target.
RSS should be kept, those who don't know the job they are doing should be replaced.
But people who don't use the right aggregators will not be able to read past 10 posts. And if everyone was using aggregators, I wouldn't have such a high bandwidth bill.
> WRT scraping, it is getting harder because funny websites w/ no static content.
But sites that can have an RSS feed necessarily need to deliver their content in a static form. What surprises me is that we don't have a tool for even those cases.
I do not think such aggregators exist, and you can just ignore people using software that do not comply to widespread conventions.
> And if everyone was using aggregators, I wouldn't have such a high bandwidth bill.
I don't understand this sentence. All RSS client software is called aggregator.
> But sites that can have an RSS feed necessarily need to deliver their content in a static form.
Not really. Many websites which are essentially blogs are transforming themselves into single-page web apps. My uni's websites included. Some do it for the $$$, some for reasons that I can not know (jumping the bandwagons with minds toggled off).
You can just set your blog software to truncate your feeds to a reasonable number w/o any worries. And I suggest you look at your logs, because some silly bots might be consuming your bandwith along with your normal traffic; there are some that like RSS feeds.
I'm trying to distinguish between (1) people using software the directly downloads the RSS feed from my website to their device and (2) people who use services that download the RSS feed to a server which can then serve a cached copy of the posts to many users. If everyone used (2), then I would only have my RSS file downloaded as many times as there are separate services, which is not very many. So apparently many people are doing (1), and if my feed is short then they are limited in the blog history they can read to how long they've personally been subscribed (or less, if they need to clear their device and can't re-download my old posts).
> Not really. Many websites which are essentially blogs are transforming themselves into single-page web apps
That's why I specified "But sites that can have an RSS feed...". If you have an RSS feed with useful posts in the feed, then you must be delivering static pages. (If you're just delivering snippets with links to a dynamic page, then there is no way for any service to cache the page either.) So my question is: why isn't there software to scrape webpages of blogs that offer an RSS feed with complete posts? This would enable me to comveneiently read post history going back more than 10 posts.
> And I suggest you look at your logs, because some silly bots might be consuming your bandwith along with your normal traffic; there are some that like RSS feeds.
All the bots put together make up 27% of my bandwidth. It's a lot, but it's not the root cause.
RSS is a means for people to follow new posts from you, in order to read new posts they are supposed to come to your blog and use your archives.
> it remains incredibly difficult to scrape a well-organized blog and turn it into something I can consume like RSS or a kindle book.
> Note that many/most WordPress RSS feeds aren't that useful because...
I don't think so, but maybe I missed. It is not the purpose of RSS anyways, though. RSS is for me to know when you post something new.
FTFY ;)
FTFY
Still, exceptionally low cost compared to running pretty much any website's massive CSS and JS files. A single image in most cases will take more than the entire RSS feed before compression.
That said, I would like to see a new standard (a new one would be needed [1]) that only gets the difference from what you last read - I think that would really take RSS feeds to new places of usefulness. There's no reason why you couldn't send the server an ID (not timestamp to avoid issues with timezones, clock stretching, forward/backward time setting, etc) of the last request and have it send back everything since (within reason).
My blog has both RSS and browser readers, but the bandwidth is dominated by the RSS feed. Since the RSS file has no images (just HTML pointing to the images), I think it's more likely that for my Wordpress blog the CSS/JS overhead is just not that much (as opposed to the alternative hypothesis that I have many many times more RSS readers who never end up downloading the images).
Exceptionally low cost per hit, as opposed to overall bandwidth. Overall bandwidth will probably fair off worse as you say due to the polling nature of RSS. I think in general RSS readers could do a better job of fetching heads and checking whether or not there is a change worth fetching, that would save a tonne of bandwidth.
Also in general, I would be tempted to make an RSS feed more minimalist in terms of content and markup. It should just be a short `<description>` and a link to the main article (which would still allow you to potentially monetize your content or gauge interest more accurately).
>I think it's more likely that for my Wordpress blog the CSS/JS overhead is just not that much
Also, bandwidth is just one resource - potentially each call to a page is a database read, whereas your RSS feed should be static (not sure about the WordPress implementation, but I would hope for static caching with something that doesn't change for long periods of time). I've seen WordPress database lockups with modest amounts of traffic (again, most of the time this could have been easily statically cached - but doesn't appear to be by default).
> Exceptionally low cost per hit, as opposed to overall bandwidth.
I don't understand. I'm telling you that my RSS file literally dominates my bandwidth usage in GB.
> I think in general RSS readers could do a better job of fetching heads and checking whether or not there is a change worth fetching, that would save a tonne of bandwidth.
Yah! Agreed.
> It should just be a short `<description>` and a link to the main article (which would still allow you to potentially monetize your content or gauge interest more accurately).
No, I want people to be able to read it offline. I'm not trying to monetize anything.
> Also, bandwidth is just one resource - potentially each call to a page is a database read
I have a simple website. The bandwidth is the dominant cost.
> Yah! Agreed.
You could also reduce the number of `<item>`s you keep in rotation to a more manageable number.
>> It should just be a short `<description>` and a link to the main article (which would still allow you to potentially monetize your content or gauge interest more accurately).
> No, I want people to be able to read it offline. I'm not trying to monetize anything.
It should still be much more lightweight than it's HTML counterpart. Should be almost nothing to it, next to no markup, no styling, no scripts and a highly compressible piece of data.
I don't understand how you're racking up massive bandwidth. Can you put some numbers to it:
* Bandwidth usage for RSS
* Bandwidth usage for webpage
* Hits to RSS
* Hits to webpage
* Links for both
Passing "the ID of the last piece of content I saw" would allow the server to return just the updated stuff, or abort early like an E-Tag. However, counterintuitively, as is often the way in computer science, I'm not sure it would be that big a win to be able to return partial content. The vast bulk of the win on most blogs will just be the ability to abort at all, provided just fine by E-Tags.
I would say that if your site is getting hammered by HTTP requests for your RSS, do double-check that you've got E-Tags set up and working correctly. It is in the best interests of the big scrapers to support that properly, as they are paying for that bandwidth too. RSS aggregators don't have to get too large before this becomes a top-priority feature request. Unless the feed is literally changing on roughly the same frequency as it is scanned, it shouldn't be the dominant factor in your bandwidth bill.
Since rss is text only the sizes are very small and compress well. Considering the average web page is 3M[0], pulling down 10-100k of every post ever doesn’t matter. And http takes care of not pulling the same file over and over.
It’s certainly a downside, but completely useable as is, and better than any viable alternative.
As far as standards go, I prefer simple, static file, over something requiring dynamic response.
There’s also nothing stopping the site from limiting the RSS feed to only 5 posts with a link for full.
My web pages are 100k and my RSS feed is 500k.
> It’s certainly a downside, but completely useable as is, and better than any viable alternative.
It doesn't fulfill the need I originally mentioned: making blog archives readable offline.
> As far as standards go, I prefer simple, static file, over something requiring dynamic response.
I agree static is better, but a dynamic response isn't necessary. You could just have a static file that listed all the blog posts, with a link to another static file for each post. This avoids having the user download 10 blog posts each time they want to poll if something new has happened.
You'll find your feed readers aren't particularly impressed, though.
The good news is that means you don't have to wait. You can write that now. It will work on existing RSS feeds, just not quite as optimally for your proposed use case as you might personally like, but it will still work. It will work even better on yours, which will also work in conventional RSS readers.
Now, you might have problems getting "the real page content" from your URLs, but that's a separate problem. (History strongly suggests the large-scale content producers would actively fight you if you try, because you'll probably be trying to strip their ad revenue either deliberately or accidentally as part of what you'd be doing. Which is, after all, the reason why RSS is already not terribly favored by that crowd and why they want you in closed gardens of their own devising... unfortunately getting around this problem is a great deal more difficult than hypothesizing that some sort of new standard could somehow deal with it....)
If you want to play semantics, I'm fine with rephrasing my complaint as: "We need to build on the flexible super official RSS standard -- which is little more than an XML file -- and actually agree on a way of delivering blog archives for offline reading. RSS in practice does not currently achieve this very simple goal." This is just different words to describe the same thing.
> You can write that now. It will work on existing RSS feeds, just not quite as optimally for your proposed use case as you might personally like, but it will still work.
Huh? Other website owners who would like me to be able to read their archives can modify their RSS file, but we have no agreement on how to do that in a standard way. Likewise, I could modify my RSS file, but since there isn't a standard my readers won't have software to take advantage of it.
> (History strongly suggests the large-scale content producers would actively fight you if you try,...unfortunately getting around this problem is a great deal more difficult than hypothesizing that some sort of new standard could somehow deal with it....)
The blogs I want to read offline do not have ads and do not care about this. I just want a solution that works for this simple problem, not a way to take content from people who don't want to give it to me without attaching ads.
On the other hand, if you want to say that your client supports this supposed RSS-ng standard they need to support those new features because they're part of the protocol.
This is to say, I don't like protocol extensions, but that's just me :)
I use rss-bridge in combination with rss2email to follow instagram feeds after leaving instagram.
(DMCA doesn't count, it's about circumventing things that prevent you from copying, not change your consumption.)
Not aware of any anti-user-made RSS laws, though.