> designed to store data estimated to be on the order of exabytes or larger
While google says:
> The Google Search index contains hundreds of billions of webpages and is well over 100,000,000 gigabytes in size
https://www.google.com/search/howsearchworks/crawling-indexi...
Seems like the government already has the expertise and equipment required to handle data at Google scale :)
One of your suggestions was
> That's why the stipulation that other entities be allowed to mirror the index - they can optimize the index for their own purposes and rankings on their own hardware.
And the point is that there's nobody who can do this outside of Google, Microsoft (who also does), Facebook, and Amazon.
Not to mention the problems of actually getting the data. You're at the scale of data where trucks of disks are faster data transfer than cables unless you have direct fiber backbone connections.
You're just comparing two storage numbers without taking anything else that running Google at a global scale requires.
Regardless of that, the datacenter can be appropriated by the government too.
NSA Utah: 65 MW Google Pryor, Oklahoma: 340 MW
Every Google data center does...everything.
I can assure you that it's not mapping each query down to a single-sector disk read off an inverted index.
I think the query pipeline for NSA (relative to the scales of Google's query pipeline) looks like absense-of-query-pipeline. Hence NSA using less compute and thus (the reasoning goes) less power consumption.