XPath Scraping with FreshRSS
danq.me
danq.me
(fwiw everyone I've wanted to follow has already had a feed; it's just that sometimes I've had to grovel around with View Source to find it)
That will retrieve the list of items using RSS but will fetch the article content by getting the URL and grabbing the HTML element you specified (like "article.post").
And that's not even getting into the raging tire fire that is akamai / cloudflare / whatever anti-bot technologies. I did see support for http proxies, but it wasn't clear if that was something one could set on a per-xpath-feed basis or whether the whole system had to run under one proxy (potentially $$$)
Proxies are actually pretty useless, you need to go one layer lower to be able to fake those fingerprints, what you need is a VPN instead (by VPN I mean an IP-level tunnel, not a public VPN provider - those IPs are already blacklisted and often are so for good reason).
For example, the xpath selector used in the article `//li[@class="blog__post-preview"]` would break if one more class is added which happens very often in the real world while CSS selector`li.blog__post-preview` wouldn't. (The correct xpath here would be `//li[contains(@class,"blog__post-preview")]` or even more accurately `//li[contains(concat(" ", normalize-space(@class), " "), " blog__post-preview ")]` - yeah it's realy ugly).
Either way both CSS and XPath selectors are actually really easy to learn and would be great if more tools adopted these methods like FreshRSS did! I made some interactive cheatsheets with all of the edge cases if anyone is interested in all of the web scraping parsing weirdness :) https://scrapfly.io/blog/css-selector-cheatsheet/ and https://scrapfly.io/blog/xpath-cheatsheet/
The only part that is still inconvenient is creating these xpaths which requires some trial and error.