This is a data center designed to, supposedly, store data on the scale of a yottabyte. I only say "a" yottabyte, because to assume even slightly greater than that is sheer lunacy.
That is freaking massive. If you took all terrorist cells and all terrorist activity for the history of terrorism and terrorist activity, you would not even touch a fraction of a percent utilization. We're talking rain drops in the ocean.
There is no way the NSA is merely watching the bad guys here. The data center is a few magnitudes too large for such a task.
I would assume right now they are merely recording all data, in hopes that one day they will have technology to quickly crack encryption. However, even without knowing what is said (the content), the metadata of connections gives plenty of information on what people are doing.
and yes, agreed - analogous on so many levels - a publicly admitted places where secret government things happen, which can be invoked to give an aura of reality to conspiracy theories true and false alike.
I expect these sizing claims (which presumably come from some sort of government statements about the facility) are based on a timeline on the order of ten years or more. A YB in 2024 is going to take a lot less physical space than a YB in 2014.
I tracked down what appears to be the origin of these yottabyte claims: http://www.nybooks.com/articles/archives/2009/nov/05/whos-in...
Based on that, it sounds like the yottabyte claims refer to raw, unprocessed data collected. Off the bat, I'm willing to believe in a 2 orders of magnitude decrease after data reduction techniques are applied to the raw data. My experience with data collected from telescopes was that we got about 100:1 reduction on the raw data versus what went into permanent storage.
The San Antonio site was news to me, though obviously no big secret since that book was published years ago.
We already know about Room_641A[1]
Exabytes are eminently reasonable; yottabytes are not.
Most of the large-scale sites are doing SSL offloading, so one of the first things that happens is the traffic is decrypted. Often this happens in the front end load balancer.
If the set up is as the WP described:
> “collection managers [to send] content tasking instructions directly to equipment installed at company-controlled locations,”
and this equipment is installed behind the SSL offload devices, it would see all the customer data in the clear.
Perfect forward secrecy in TLS is a bit different in that the ephemeral diffie-hellman key exchange sets up a shared key that is protected from a _passive_ attacker that observes the TLS encrypted communication and later gets a copy of the server's public key.
http://en.wikipedia.org/wiki/Secure_Socket_Layer
For instance, on HN, Chrome is currently doing this:
Your connection to news.ycombinator.com is encrypted with 128-bit encryption.
The connection uses TLS 1.2.
The connection is encrypted using AES_128_CBC, with SHA256 for message
authentication and ECDHE_RSA as the key exchange mechanism.
Using ECDHE_RSA, my browser and HN's server will agree upon a key to use for encryption using AES128 in CBC mode. Now in order to read what the server sends me and what I send to the server, you need to break the crypto:1. Brute force the 128 bit key. This is.. probably not going to happen?
2. Via a weakness in the AES128 algorithm or implementation, you can simplify a brute force into feasibility (AFAIK, no such attack currently exists).
3. Via a passive attack on ECDHE_RSA, you could potentially guess the shared key efficiently and decipher our communications (AFAIK, no such attack currently exists).
So it's not quite as simple as recording encrypted information and obtaining the SSL keys. You need the server to actively remember the keys used for every encrypted connection, and obtain those, too. Or MITM everything and record the unencrypted data.
Although your post did bring up another question that I never thought about. Does Google even "send" email when it goes from one Gmail user to another? That could theoretically all be handled internally, but it never crossed my mind that they wouldn't use SMTP.
Looking at a random email in my Gmail account from a different Gmail user, it looks like they do use SMTP, or at least they are adding headers as if it went by SMTP. But both ends of the SMTP are at the same IP address:
X-Received: from mr.google.com ([10.229.72.135])
by 10.229.72.135 with SMTP id m7mr3900891qcj.17.1370903118607 (num_hops = 1);
Mon, 10 Jun 2013 15:25:18 -0700 (PDT)