Show HN: RSS feeds for arbitrary websites using CSS selectors
feed-me-up-scotty.vincenttunru.com
feed-me-up-scotty.vincenttunru.com
Anyway, here's another self-hosted open source RSS feed generator for arbitrary websites: https://github.com/hueyy/HungryHippo
So yeah, definitely straightforward enough for a case of NIH syndrome. I think putting together the website took more time than writing the tool itself...
See https://sfeed.org. In the spirit of the multi-meaninged RSS acronym itself, the S might stand for scrape, selector, speed, or of course Scotty.
Vinni, might you be interested in enabling the standard in Feed Me Up?
Happy to take a merge request that adds the option to set `extends = "sfeed"` or `extends = "hatom"` and then automatically sets the correct selectors though.
(That said, if a publisher goes through the trouble of adhering to those selectors, they might as well publish a feed while they're at it :)
Microformats, I see this is the proper way.
I do think it's (much) easier to add some HTML classes than output an entire separate file.
I've been using Nuxt + Strapi as my new CMS stack, and while it's a big step forward in so many other ways, outputting an RSS feed is far from automatic.
Ah well. What I do like about my current approach is that I have full control without having to run my own server, which is nice.
Feed Creator is in PHP and unlike yours doesn't produce a static file. The CSS selectors, URL and other parameters are all embedded in the URL, e.g.
https://createfeed.fivefilters.org/extract.php?url=http%3A%2F%2Fjohnpilger.com%2Farticles&item=.entry&item_desc=.intro&item_date=.entry-date
The biggest issue I've found is people struggling to work out which CSS selectors to use. I wrote a blog post not that long ago to help people use the browser's developer tools to do that. Might help too with anyone trying to use this (although perhaps most of the HN audience doesn't need help here): https://www.fivefilters.org/2021/how-to-turn-a-webpage-into-...I ended up writing my own CLI tool that similarly supports CSS selectors for feed generation: https://github.com/dayzerosec/feedgen
I did write it specifically for my use-case so there are some "warts" on it like custom generators for HackerOne and Google's Monorail bug tracker. But perhaps someone else might benefit from its ability to create slightly more complicated RSS, Atom, or JSON feeds.
Example config with date parsing: https://github.com/dayzerosec/feedgen/blob/main/configs/bish...
I've long been waiting for the rise of the unblocker. Community curated, like you suggest.
Ad blockers use filters to exclude. Like a blacklist.
My hunch is a strategy of using scrapers to extract OC, completely skipping over ads, would also be viable.
The output could be a bit more rich than today's Reader View.
For anyone curious: Free - 2 RSS feeds 24 hour refresh rate $9.99 - 40 RSS feeds 30 min refresh rate $19.99 - 100 RSS feeds 15 min refresh rate
My company uses RSS.app & I find it extremely nice to work with. You don't even have to provide selectors/click on elements for the vast majority of websites, they do that all for you.
We could roll our own solution using any of the above offerings (or something built in house), but it's cheap enough for our usecase that we don't see a point.
In these situations it might be best to rely on position of elements and other CSS selectors rather than attribute values. Unfortunately the :nth-child and :nth-of-type selectors still trip many people up. In Feed Creator we borrowed from XPath to make selecting by position a little easier. We've got a comparison here: https://help.fivefilters.org/feed-creator/css-selectors.html...
Unlike this project Feediron is only for modifying existing RSS feeds to extract the desired information. Typically uses xpaths to select content
Two things I am using:
Twitter to RSS: https://github.com/RSS-Bridge/rss-bridge
Arbitrary RSS feeds: https://feedity.com
We've since rebranded to New Sloth: https://newsloth.com, which is now a simple integrated feed builder, reader and clusterer/deduplicator, specially aimed for knowledge workers with hundreds and thousands of feeds to monitor daily.
Besides visual selector-based feed generation, our API can auto-magically detect relevant selectors (in most cases) as well.
Not to piss on his/her parade, though: this is a great thing. I am going to tinker with it tonight.
But it does.
Primary source: I experienced it on my own site.
Secondary source: https://developers.google.com/search/docs/guides/fix-search-...
(For my part, I disable JavaScript by default for various reasons, mostly performance, and it’s decidedly uncommon for a general-internet site to be completely broken by it. Sites that get posted on HN are disproportionately JS-dependent, especially if they’re new.)
Possibly SEO is less a concern for the type of website I initially made this for, i.e. Dutch real estate agents. Most people find their listings through funda.nl rather than through search engines; I was just hoping to see them listed before they got posted there.
Send me a message on Twitter or email me (hacker_news@ my domain) if you still want the URL of a failing website to play around with.
- https://vincenttunru.gitlab.io/feeds/funfacts.xml
- https://vincenttunru.gitlab.io/feeds/wikivoyage.xml
And the combined feed:
- https://vincenttunru.gitlab.io/feeds/all.xml
Add those links to your feed reader to see rendered examples while I update the site.
Edit: preview added to the website.
Not so; unlike the HTML <base> element which applies to a document, the xml:base attribute is applied to an element and its descendants. The typical pattern (as shown in the RFC 4287 example) is to put it on each entry’s <content>. In your markup, you’ll end up with each entry having its URL in three places:
<id>http://example.com/item</id>
<link href="http://example.com/item"/>
<content type="html" xml:base="http://example.com/item">…</content>Edit: the package I'm using to generate the feeds does not support that attribute yet, so it'll have to wait a bit for my PR to hopefully be accepted: https://github.com/jpmonette/feed/pull/158
Thanks for the pointers!
Similarly to my repository, I think I would suggest the option to fetch the configuration file from an external resource defined via an action secret. For my automation I'm using a Gist (not sure if Gitlab has same thing; also private but publicly accessible snippets).
At least that way you can keep your own feed configuration while allowing those that fork the repository to not have to manually fix conflicts within the feeds.toml config.
So far I've forked feeds, edited feeds.toml, checked it out as a branch gh-pages, pushed the branch up to github.
I can see the page at https://<username>.github.io/feeds/ but https://<username>.github.io/feeds/actions is just a 404.
Does that work?
ok I managed to access the actions route by a different URL to the readme. I copied the pages.yml workflow from your repo.
Few minutes later I could see my feed. Very nice! I need to clean up my selectors now!
https://github.com/DIYgod/RSSHub
This perhaps has more flexibility and can deal with almost any website.
https://www.notion.so/How-to-create-your-own-stock-and-crypt...
Actions should not be used for: - any other activity unrelated to the production, testing, deployment, or publication of the software project associated with the repository where GitHub Actions are used.
[1] https://docs.github.com/en/github/site-policy/github-terms-f...
Some design changes and the whole thing breaks. Maintenance nightmare.
This is why APIs exist.
It also begs the question of why they don't make their content that way, perhaps it was by choice. Especially if you use this kind of tool to syndicate things.
Not quite sure of the novelty of this one given that people have scraped things for decades.
There's probably a wiki endpoint for their example, maybe, maybe not. A lot of wiki's stuff is free to download and they have extensive endpoints for acquiring data.
Just to be clear, I'm not against scraping. Done it myself plenty, point was that it's hard to maintain and if I were to make a tool for the masses, it'd be in a UI that highlights elements and lets you select them rather than dig around the DOM. Pretty sure things like this have existed for 10 years +
Given that there wasn't really an alternative to scraping for me, I wanted to at least be able to pick selectors myself that were less likely to break with minor changes. Then I figured there might be more people who know CSS and have the same desires, hence my sharing it here :)
Purely IMO, a more friendly way to go about it to abstract from code and CSS knowledge is to run a UI that highlights elements, lets you select them, select the title, description, link etc and there you go. Same thing but without the knowledge of DOM/XPath/selectors.
And how would that software then remember your chosen elements if not by their CSS ids/classes?
Just kidding, that's horrendous.
Your comment is “scraping is brittle”, which is not helpful. Everyone knows scraping sucks. It’s why these tools are being made to make it less unpleasant.
Right click on the element in Firefox, click "inspect element" and it shows you the unique selector for that element.
Though Show HNs do not need to be novel to get to the top. And RSS related projects seem to go to the frontpage simply because of the fondness HN has for the technology. Also, don't forget https://xkcd.com/1053/ :)
But it can bring feed capabilities to simple, timeline like sites. Sites like most frameworks produce. And with the dev tools you can quickly find the needed dom path. It's limited, but easy to use. If it doesn't work, you need a real scraper which is an order of magnitude more complex. (I maintain some of them as well.)
Screen readers are scrapers too. "Scraper" does not mean "user agent I don't like".
Please and thank you.
neat
Would you understand it with those names?
I for one know most of my pals do know about RSS, but haven't heard the term 'Atom feed'.
But the problem didn’t exist in the same way with SSL/TLS: SSL was actively killed off in favour of TLS so that regardless of what you call it, you’re actually dealing with TLS. But with feeds, RSS is still supported, and so the mindshare problem happens: people hear about RSS and so implement the inferior and problematic RSS rather than Atom, because they’ve never even heard of Atom and so don’t know that it’s what they should have implemented instead, because it’s more robust. Calling feeds RSS perpetuates RSS, which is undesirable.
I say just advertise it as “feeds”, not “RSS feeds”. Atom is an implementation detail, just like RSS should be. (Well, RSS should be dead in general feeds, with only podcast feeds keeping it alive.)
I think that's even worse. For most ordinary people that has a bunch of unintended meanings. "My Facebook has a feed, do you mean subscribing to your site on Facebook? Why do I see a bunch of code when I click the link?"
This also leads to rather poor solutions and discussions which are just reinventing the old stuff again and again, instead of letting the domain evolve to match modern demand.
Take for example this project here. Is there any feedreader out there who has this built-in to make it comfortable for the user to use this? The state of tooling as I know it is still at the level were we have seperate tools which pushes out a newsfeed-format, which than must be maintainend seperatly in our feedreaders. Which is an extraordinary hazzle to maintain configurations on multiple places.
Maybe I'm misunderstanding the question. This project is easy to use in all feed readers precisely because it just generates a feed URL that you can subscribe to like any other feed.
On the other hand, maybe you're asking why the feed readers don't have this feature built in. Some of them do! The one I've used for years, Inoreader, has it, although it might be a premium feature.
On the other other hand, sometimes it seems from your comment like your issue is with the very idea of RSS. That a website can publish a feed in a well-defined / open format and that feed can then be consumed by any conforming reader. That's as I see it the beauty of the whole system. I don't need to use the site itself, if it sucks. I can subscribe to anything I like using whatever software I like, rather than relying on a centralized service. What's the alternative, something like Twitter where I subscribe to a bunch of accounts and they push whatever they like into my feed?
There never was a valid reason to define a separate format rather than extending RSS to clarify ambiguities - I think its time to drop this bikeshedding and just think of Atom as a slightly different type of RSS.