Show HN: Automatic web page content extractor for Node.js and the command line
github.com
github.com
http://www.alfajango.com/blog/create-a-printable-format-for-...
The reason for my skepticism is Arc90's readability extension [1]. At the surface it looks more complex. I could of course be wrong!
[1] https://github.com/Kerrick/readability-js/blob/develop/reada...
The complexity of the readability plugin seems to be due to the fact that it actually does a lot more than just making something readable. For example, the point of my algorithm was to strip all style information from a page and show only the content, leaving it to be styled according to the global stylesheet. Notice that my script not only removes linked stylesheets and style tags, but also style attributes of all elements. The readability plugin actually does things like counting reference links and styling them a certain way [1].
It has 53 lines dedicated to both getting and normalizing the article title, when the vast majority of the time, it's just the first h1 or HTML title attribute (which could be a one-liner and is also outside the scope of the "meat algorithm" since it's just trying to get the body) [2].
It has 22 lines dedicated to injecting a custom readability footer into the page [3].
It has 74 lines dedicated to converting all inline links to footnotes [4].
It has 55 lines dedicated to injecting typekit fonts [5]. And on and on.
It also does things like figuring out when to float an image in the article and when to make it full-width [6] (as opposed to just leaving the image inline with no styling, as my script does).
And it even dedicates 333 lines of code to finding pagination links to build content from multi-page articles [7], which my script simply doesn't do, since it only cares about the content of the current page.
It also does things like computing a content-weight score for parts of the page, I'm guessing to determine a relative heuristic for which parts are most likely to be the main content [8]. This is actually the path I had started to go down, before I realized that my much simpler version solved 95% of the use-cases I had, and that, for my purposes, I didn't really care if it failed 5% of the time.
I think the discrepancy in complexity can be explained really easily:
a) The readability plugin does a lot of things not directly related to simply grabbing the content of the page.
b) There's a lot of complexity involved in trying to get it to work for that last 5% (those are the really weirdly-structured sites, for which you'd need to develop some sort of heuristic and/or learning algorithm).
In other words, no, I didn't exaggerate the success-rate. The readability plugin is just very functionally different. The algorithm in my article is also not the complete script; the algorithm was meant to be a base starting point to build from (certainly you would need to if you need greater than 95% success rate, which most applications would).
EDIT: Also should point out, my article was also meant to document my surprise in how easy it was to get a 95% solution. Most of the complexity I've seen in other scripts is in trying to figure out all HTML nodes on a page which could possibly be related to the main content, so that you can scan for them and reassemble only those nodes. The breakthrough for me was that if you can find one paragraph tag, then you can just go up a level or two to the containing node, and blindly grab all nodes, whatever they may be, within that containing node. The main pages this doesn't work for are pages that don't use paragraph tags in their article body (e.g. plain text with break tags all over the place, which were surprisingly few and far between with my sample set).
[1] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[2] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[3] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[4] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[5] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[6] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[7] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
[8] https://github.com/Kerrick/readability-js/blob/4596857da3cc4...
As an aid to extracting the right body content, have you considered comparing multiple pages from the same site, and giving greater weight to content that differs (the article) rather than content that stays the same (the navigation)?
Unfluff is a small, simple .js library that can be installed and used in seconds. It doesn't have any external dependencies on data files or other language runtimes. So it just depends on which tool is right for your job. If you are writing a quick script, this might be a lot easier to use.
it's written in CoffeeScript
(The transpiled JS files lack comments and meaningful new lines of the original CoffeeScript source.)
There's similar functionality on github (but I haven't found documentation for it)... Example: https://github.com/shirou/kptool/commits/master.atom
<link type="application/atom+xml" rel="alternate" href="/feeds/p/dcef3/gitchanges/basic">
which refers to an Atom feed for the project's commits.Same for Github projects, e.g. at https://github.com/petdance/ack2 you will find:
<link href="https://github.com/petdance/ack2/commits/dev.atom" rel="alternate" title="Recent Commits to ack2:dev" type="application/atom+xml" />Works as a SAX consumer, pretty fast, not sure how accurate.
There are of course other libraries like this (boilerpipe, Goose, etc), but they tend to be written in Java and Python. The very few existing Node solutions didn't fit my needs so I hacked this together. So for people looking for a quick and simple Node solution, this might be useful.