What do you think of this generalized architecture?
HARDWARE: differs depending on whether you want local search/analytics or just network storage.
For mobile use, either a VPN back to your personal home/cloud server, or a hackable wifi hard drive proxy, e.g. Seagate Wirless Plus + HackGFS.
For non-analytics home use, hackable router with USB3 storage and Linux software RAID, connected to a USB3 drive chassis with room for 2-4 disks.
For analytics home use, a microserver like HP N54L, Dell T20 or Lenovo TS140. Up to Xeon processor with ECC memory, plus 4-6 internal disks and up to 32GB RAM. Sold without a Windows tax, supports hardware virtualization and Linux. Possibly FreeBSD with ZFS.
SOFTWARE: generalized multi-tier cache AND compute. Camlistore and git-annex are tackling multi-device storage sync. For archives, we need a search interface that will query a series of caches, e.g. mobile > home > trusted friends private VPN (tinc overlay) > public paid cloud archive (pinboard et al) > public free cloud archive (archive.org).
It's important for usability to have a simple, local UX that will take a search string, propagate across all private/public federated tiers of storage and compute, then aggregate the metasearch results on the client.
With this approach, we can collectively pool resoures to improve on CommonCrawl.org, without locking up the 300TB index at AWS. This would turn web search engines into a secondary source, rather than a primary source. First search your archive + trusted friends, then trusted verticals (e.g. HN, StackOverflow), then a generic web search.
Let's be clear: the goal is not to archive "everything" in the world, only that which is personally important to the viewer. This attention metadata has long-term value. With this architecture, it is always optional to escalate a query to a public archive or search engine. Most importantly, there is technical autonomy and low-latency compute for local queries.
For web pages, wget of WARC formats (per HN advice on another thread) and wkhtmltopdf (available as Firefox plugin to print to PDF) will keep local archives. Recoll.org (xapian front-end with user-customizable python filters) on Linux will search full text and provide preview snippets, or lucene/solr can be adapted.