(1) Used a preexisting aggregate web content format. Their ad hoc format is simple enough, but can't handle content with NULLs, and loses valuable information (such as time of capture -- you can't trust server 'Date' headers -- and resolved IP address at time of collection).
They could use the Internet Archive classic 'ARC' format (not to be confused with the older compression format of the same name):
http://www.archive.org/web/researcher/ArcFileFormat.php
Or the newer, more involved and chatty but still relatively straightforward 'WARC' format:
http://archive-access.sourceforge.net/warc/
(2) Explained how the 3.2 million pages in their initial dump were chosen. (That's only a tiny sliver of the web; where did they start and what did they decide to collect and put in this dataset?)
(FYI, I work at the Internet Archive.)