Does that 1.5m figure include redundancy (e.g. RAID)? What about if the data is compressed (either trivial gzipping each "file" the archive has, or using a compression window that spans multiple files)? I imagine the HTML/plaintext content of the Internet Archive would compress very well.