Reading the web offline and distraction-free
blog.owulveryck.info
blog.owulveryck.info
The take-any-webpage-offline need is also common in the education space (teachers want to save a webpage and send it to their students as part of a lesson and don't want to worry about availability or ads etc).
I used to work on tools for this https://github.com/learningequality/ricecooker/blob/develop/... and https://github.com/learningequality/BasicCrawler/blob/master... which worked quite well for most sites, but still very far from a general-purpose solution.
There is also more powerful/general-purpose scraper that generates a ZIM file here: https://github.com/openzim/zimit
It would be really nice to a "common" scraper code base that takes care of scraping (possibly with a real headless browser) and outputs all assets as files + info as JSON. This common code base could then be used by all kinds of programs to package the content as standalone HTML zip files, ePub, ZIM, or even PDF for crazy people like me who like to print things ;)
If I recall correctly the only gotcha is that the option to save in this format needs to be enable using the flags settings.
I use the image functional and only sites like Twitter fail to work correctly, although it's probably my CGI gateway timing out waiting for JavaScript or whatever.
My personal solution has been https://github.com/captn3m0/url-to-epub/ (Node/readability), which I've tested against the entirety of Tor's original fiction collection[0] where it performs well enough (I'm biased). Another tool that does this beautifully well is percollate[1], but it doesn't give enough control of the metadata to the user - something I really care about.
I've also started to use rdrview[2], which is a C-port of the current Firefox implementation of "reader view". It is very unix-y, so it is easy to pipe content to it (I usually run it through tidy first). Quite helpful in building web-archiving or web-to-pdf or web-to-kindle pipelines easily.
[0]: https://www.tor.com/category/all-fiction/original-fiction/
[1]: https://github.com/danburzo/percollate
But it has not been maintained, since the author joined Facebook.
It works alright, but it has many issues.
If I understand correctly, a full on replacement for newspaper is in the wings, seeking to offer a sustainable content extraction tool in Python.
But it isn’t ready yet. And some of the problems in this area mirror those faced by web scrapers.
I was working with Mercury Parser (pluggable parsing for different sites) in the past.
It connects to the Pocket API to get the parsed articles, pushes them through quite a lot of BS4 clean up, then renders them using paged.js. The resulting PDFs are then printed by Lulu.com, and they come once a month as a printed book to read completely offline.
I solved the Medium image issue with CSS as far as I remember. `.medium\.com svg:first-of-type` and then set it to `display: none`.