> Why not use a regular folder and file system with separate metadata files? It seems simpler and more user-friendly.First of all, no matter how easy it is to extract the files, the results will be kinda user unfriendly - unless someone goes through the archived html files updating all the URLs for images and javascript and so on, it'll probably end up looking pretty broken. For users who just want to view a few files, Wayback Machine is a better tool for the job.
With that said, some situations where warc is helpful include:
* If a website had a page at www.example.com/foo and a file at www.example.com/foo/bar.txt you can't express that in a regular filesystem, as you can't have a file and a directory with the same name.
* If a website uses some absurd 2000 character URL like https://s3.eu-west-2.amazonaws.com/document-api-images-live.... you don't end up with a filesystem-breaking filename.
* If www.example.com embeds an image from exampleusercontent.com you can capture the image in the same archive file.
* If for some reason you want to store daily copies of www.cnn.com in the same archive for comparison purposes - you can.
* It lets you store headers, 300 redirect messages, case-sensitive filenames, and all that sort of stuff.
* And it's an extremely simple format - basically human readable. So if you think you're archiving for the super-long-term and want to make really conservative choices, you can be pretty confident in plain text.