A New Instapaper Parser
blog.instapaper.com
blog.instapaper.com
We did some work in 2014 on an archive of all stories posted to HN, with the goals of (a) having a lightweight, readable version of everything quickly available and (b) doing analytics on the content. But we got bogged down on getting the actual content programmatically across the full spectrum of cases. This is one of those problems where not merely one devil is in the details but a whole legion of them, and not the glamorous kind. Getting it right would have sucked up all our resources, and the APIs out there (e.g. Readability) came with problems too, so we dropped the project.
But for a programmer who enjoys the snake-pit-of-corner-cases type of challenge, this would make a fine project, one with real public-service potential. We can't work on it ourselves, but we'd consider funding it.
Getting to 80% quality isn't hard. 90% is tricky. 95% incredibly costly.
Goose[0] for example was a project I last explored for article summarization.
I could see an algorithm merely trying to find the longest prose on a page, then analyzing it for content and then trying to find other content within other elements that are likely relevant. Armed with this and multiple articles from the same source, you should be able to correlate the two to get a smarter parser that can adapt to changing DOM on each site.
Interesting! Besides being useful for HN, I’m sure it could also be useful for quite a few of the YC startups.
I would assume this is built on some sort of headless browser implementation but who knows, maybe not. Hopefully Instapaper does a followup with more technical details.
I'm not sure how this article is supposed to make a paying user happy, it doesn't show any measurable metrics (neither any user should actually care about that, it should just work). I still wonder how hard adding a simple "if github then" check is.
MarkItDown is a "toy" implementation of a rich-text to markdown converter, that I wrote 3 years ago and still enjoys significant usage. Maybe it could be used as a start for a more full-featured parser.