Rdrview – Firefox Reader View as a Linux command line tool
github.com
github.com
For a quick preview of each library on a random sample of 16 articles posted to HN, see https://github.com/awendland/readable-web-extractor-comparis... (you’ll need to expand a row to see its results).
There's a lot of randomness in what happens to get noticed or get traction on HN. We have various tricks to try to mitigate that. One is that we allow a small number of reposts if an article hasn't had significant attention yet (https://news.ycombinator.com/newsfaq.html).
I'm the kind of user that just lurks in the front page, so I don't really have any right to complain if my work doesn't get noticed by others. I'm glad that it did though.
> we allow a small number of reposts if an article hasn't had significant attention yet
Does that apply to Show HN too?
Did you find a way to deal with this hindrance ?
It won't do anything for the TechCrunch case you describe, because it only fetches the one webpage you point it to (and any redirections).
Did you try pretending to be a search engine crawler, for an idea..?
So far, I've never run into this problem myself. If I did, I think I would use tor.
[1] https://old.reddit.com/r/commandline/comments/jaluzg/firefox...
(I understand that navigating some sites would be a challenge as Reader Mode doesn't always know exactly what is cruft, and what isn't. But I don't see it as insurmountable.)
Although it used to fully load the page then convert it to reader view which was a bit annoying on an old and increasingly sluggish iphone 6.
Slightly related: I'm always confused when I stumble upon an article that I can view just fine in Firefox's reader view, but doesn't get recognised as an article by Mozilla's Pocket, making me unable to highlight stuff.
So, granted that Mozilla owns Pocket now, but they were their own thing long before the Mozilla acquisition, and I doubt their codebases have really merged.
For example, the API has no way of extracting the highlights, I can't scrape them because their login form is behind Google's CAPTCHA, and at the same time I know of third-party services that somehow have access to the highlights (like readwise.io). Contacting them just resulted in "we're a small team and have no updates on when we'll expand our API".
Newsboat filters operate on the whole feed. So the script needs to read the input rss (cf feedparse) then retrieve the HTML (cf rdrview) and then generate rss errrr.... Feedparse doesn't do that last bit. And then add a layer of on disk caching because otherwise it would keep re-downloading all entries over and over again.
I'll get back to that tomorrow I guess. Now for shut-eye
I was just setting the BROWSER environment variable to a script that makes me choose if I want to open the article on rdrview (with lynx) or firefox.
Any plans to make this available as a library? Would be nice to use this as a C-library in other languages, instead of current approaches where it gets re-written everywhere.
Kudos!
Thanks! Yes, it was harder than I expected.
> Any plans to make this available as a library?
I've been thinking about it, though I've never written a library before. I think I would like to get the code to build on other platforms first.
Now, it would be great if there was a way to pipe the html from a webpage into another tool from inside lynx. Maybe that exists already and I'm just not aware of it.
For future reference, it appears that this can be done with elinks.
With the right configuration, it can also do stuff like automatically replacing urls, or youtube-dl'ing a page and piping its output to mpv.
Reader view is the exact solution you need to counter this.
Right clicking -> block element with ublock origin is also an impressively easy solution for those consent popups, since such a click does have value in court.