The WARC format is extremely simple and yet so powerful. Most importantly though, there are already pebibytes of already crawled archives.
This is a fairly straightforward mapping of a .warc file to a .sqlite database. The goal is to make such archives SQL-able even in smaller pieces.
The schema I've come up with it's tailored around my requirements, but comment if you can spot any obvious pitfalls.
PS: I do believe that at some point .sqlite will become the defacto standard for such initiatives. Sure, it's not text... but it's pretty close.