Readability.js
github.com
github.com
Running this in a terminal (after installing shot-scraper):
shot-scraper javascript \
https://simonwillison.net/2022/Mar/14/scraping-web-pages-shot-scraper/ "
async () => {
const readability = await import('https://cdn.skypack.dev/@mozilla/readability');
return (new readability.Readability(document)).parse();
}"
Outputs this: {
"title": "Scraping web pages from the command line with shot-scraper",
"byline": null,
"dir": null,
"lang": "en-gb",
"content": "... long string of HTML ...",
"length": 7104,
"excerpt": "I\u2019ve added a powerful new capability to my shot-scraper command line browser automation tool: you can now use it to load a web page in a headless browser, execute JavaScript \u2026",
"siteName": null,
"publishedTime": null
}The Alan Turing Institute maintains a Python wrapper around readability.js, too: https://github.com/alan-turing-institute/ReadabiliPy.
The first one use js lib but it's kinda limited. the second one use Go, compiled to wasm. Both deployed to Cloudflare workers
Ahaaa, so that's why sometimes I get the reader icon and sometimes not (especially on mobile).
Can also use `JSDOM ---> element.textContent` depending on your needs. Useful for snagging all the text content or a specific element's.
[1] https://smort.io
also 2nd version use golang and copmile to wasm https://github.com/tuananh/reader2
It works well with text-mode browsers like w3m.