RSS-Bridge – The RSS feed for websites missing it
github.com
github.com
>>> Dear so-called "social" websites.
Your catchword is "share", but you don't want us to share. You want to keep us within your walled gardens. That's why you've been removing RSS links from webpages, hiding them deep on your website, or removed feeds entirely, replacing it with crippled or demented proprietary API. FUCK YOU.
You're not social when you hamper sharing by removing feeds. You're happy to have customers creating content for your ecosystem, but you don't want this content out - a content you do not even own. Google Takeout is just a gimmick. We want our data to flow, we want RSS or Atom feeds.
We want to share with friends, using open protocols: RSS, Atom, XMPP, whatever. Because no one wants to have your service with your applications using your API force-feeding them. Friends must be free to choose whatever software and service they want.
We are rebuilding bridges you have willfully destroyed.
Get your shit together: Put RSS/Atom back in. <<<
Not gonna happen but I fully agree with the sentiment.
I don't blame them for not maintaining it though. I grabbed the code to see if I could get it working and realized Facebook has made their newer pages fiendishly difficult to scrape. Assuming one does put in the effort to make it work, how long will it last till it's broken again?
I wish there was a viable alternative platform for organizations where "public" content is actually accessible publicly.
I'd personally be happy with an RSS feed of Facebook as just a series of images in a feed, but I guess once you've gone that far you could relatively simply run it through OCR to get most of the text.
Have you taken a look at https://mbasic.facebook.com ? It might be simpler to parse.
Data archival, retention and accessibility is absolutely fundamental and it's unfortunate to know that so many companies are hell bent on stopping individuals from using open-means to access such data (although I am not surprised, and from a rational perspective I can understand their reasons for making it difficult).
True. There is no maintainer for FacebookBridge.
As for now, the only way to fix FacebookBridge that I see is: 1. using existing Facebook account 2. fetch pages using Selenium (simply speaking - web browser)
This is mentioned in https://github.com/RSS-Bridge/rss-bridge/issues/2400
This has not been my experience. It is still as easy as ever using the mobile sites.
However I suspect one person's notion of "scraping" is not always the same as another's. I prefer to work on the command line, in textmode. Thus when I check Facebook I want text-only, no graphics. First I extract the all the story.php URLs and sort them by date. Then I retrieve the contents stored at these URLS and dump into a text-only format so I can read through comments. If there is something that looks interesting I can save the story URL and check out the photos later on a computer that has graphics layer loaded.
TBH, with Facebook I mainly just check messages and notifications. Saving story URLs in chronological order is useful for me because it makes it easy to go back and find items from the past. But if I were really serious about monitoring a "feed" I can create one myself, better than Facebook's, by retrieving the profile page of each friend and extracting story URLs from their source instead of relying on Facebook's manipulative algorithms that deliberately hide stories and re-order what they do show, non-chronologically, in a way that suits Facebook's interests over the user's.
I do not use any fancy software, just a local TLS proxy, netcat (or equivalent) and ubiquitous base userland UNIX text processing utilities. I use a text-only browser (not lynx) for reading HTML and a pager (less) for reading formatted text.
I use the m.facebook.com or mbasic.facebook.com sites, not www.facebook.com because the site on the www subdomain uses GraphQL instead of hyperlinks. Not to mention the mbasic subdomain has no ads.
The way websites use GraphQL instead of simple hyperlinks is unnecessarily complex, reminiscient of a Rube Goldberg cartoon.
[1] https://github.com/bAndie91/libyazzy-preload/blob/master/src...
Note I am referring to a local proxy, listening on the loopback device and under the control of the user, not one listening on a network interface connected to the internet and certainly not one operated by a third party as a "service".
Good luck.
There were a handful of pages on Facebook that I used to keep tabs on that would have had to have been public as I do not have an account, but now I can't get FB to let me see some of them without a login (for example https://www.facebook.com/MadisonAudubon). I have no idea if my IP as been flagged, which they appear to do aggressively, or the pages were made private or FB introduced some addition settings and the pages are now configured or what.
[0] Show HN: RSS feeds for arbitrary websites using CSS selectors
https://news.ycombinator.com/item?id=27739568
[1] All About RSS - RSS Feed Generation
https://github.com/AboutRSS/ALL-about-RSS#-rss-feed-generati...
I guess projects might have different goals in mind, but IMO would be nice to share the scrapers among the projects
Life itself works that way and it has proven to make life very robust, even when individual life-forms are very fragile.
In essence the project is a scraper that generates web feeds.