We did some work in 2014 on an archive of all stories posted to HN, with the goals of (a) having a lightweight, readable version of everything quickly available and (b) doing analytics on the content. But we got bogged down on getting the actual content programmatically across the full spectrum of cases. This is one of those problems where not merely one devil is in the details but a whole legion of them, and not the glamorous kind. Getting it right would have sucked up all our resources, and the APIs out there (e.g. Readability) came with problems too, so we dropped the project.
But for a programmer who enjoys the snake-pit-of-corner-cases type of challenge, this would make a fine project, one with real public-service potential. We can't work on it ourselves, but we'd consider funding it.