I'm currently trying to bring the PHP port up to speed here: https://github.com/fivefilters/readability.php
We use an older version as part of our article extraction for Push to Kindle: https://www.fivefilters.org/push-to-kindle/
I'm currently trying to bring the PHP port up to speed here: https://github.com/fivefilters/readability.php
We use an older version as part of our article extraction for Push to Kindle: https://www.fivefilters.org/push-to-kindle/
Check the codebase of some popular parsers:
Firefox (already mentioned): https://github.com/mozilla/readability/blob/master/Readabili...
Google Chrome: https://github.com/chromium/dom-distiller
Mercury parser: https://github.com/postlight/mercury-parser
We use these in our own tools and also get contributions from others, including Wallabag users: https://github.com/wallabag/wallabag
Before it was sold, Instapaper used to have something similar. A public database of its site-specific extraction templates. We used that as the starting point for our repository.
What do you fallback to if the rule is not present or doesn't work?
Tidy determines if the source HTML should be cleaned up first with HTML Tidy - https://github.com/htacg/tidy-html5. If you're parsing the source HTML with an HTML 5 parser, as we are now, it shouldn't be necessary any more (I think we actually ignore it now). We used it more before when we relied on libxml parsing, which often trips up on modern HTML.
Edit: Oh neat it does actually. https://github.com/ArchiveBox/ArchiveBox/wiki/Configuration#...
> Archive method SAVE_READABILITY
> Extract article text, summary, and byline using Mozilla's Readability library. Unlike the other methods, this does not download any additional files, so it's practically free from a disk usage perspective. It works by using any existing downloaded HTML version (e.g. wget, DOM dump, singlefile) and piping it into readability.
ArchiveBox and the other stuff from the "DIY no-credentials don't-care-about-the-rules" web archiving community, like ArchiveTeam.... continues to astound me with it's quality and "professionalism" (as a credentialed professional in the field of digital library stuff... they are often outdoing the actual credentialed professional community).
Washington Post, I'm looking at you mofos. Chief reason I'll seek out any alternative news site for archival. It's been this way for about a year, if not more.
https://github.com/ushnisha/tranquility-reader-webextensions