So what inference are we drawing? I honestly wasn't aware that the capability to
store all voice traffic was in question. That might be owed to my background in telecom though.
Telephone codecs are extremely low bitrate (relative to something like music, much less video) because you have very clear design constraints, and those constraints are forgiving. You just need to be able to understand the caller's voice, not accurately reproduce a live performance of Beethoven's 5th. I agree that the estimates to store the data are on the very, very high side, but I'm not sure that number is even significant in the whole scope of the challenge.
The challenges in obtaining every phone call made in the US aren't storage related; they're almost entirely collection and aggregation related. These ballpark numbers don't even attempt to factor in that portion of the cost. IMO, the collection, transport, and aggregation systems are easily 8/10ths of the problem, not storage. Storage is an extremely scalable solution.
If the government were going to do something like this, they'd likely tap in to the phone networks at the same places everyone else does. There are major telecom aggregation points - called "tandems" [1] - around the country. If you want to set up your own phone carrier with an actual physical network of your own, this is where you plug in. The government would have to do the same. They can't simply tap in at a single aggregation point, because not all phone calls pass through the same points.
From there, you face the choice of storing the data regionally in several data centers, or attempting to aggregate it all back to a central data center. IMO, a disaggregated approach makes a lot more sense. You could reasonably expect to transport all CDR [2] data back to a central location, but you wouldn't want to send all media (audio data) back to one place. It would be simple enough to only fetch what you need based on a query against call data. You'd want all the the CDR data in one place so you could perform your "big data" analysis on it efficiently, then cherry pick media to pull in for analysis.
I'm certain I haven't even scratched the surface here. This only gets the government long distance calls. It doesn't touch local, or even intraLATA calls. Maybe the government isn't interested in those calls though. Maybe they're only interested in international calls, which makes the problem simpler, not harder.
Others in the thread have mentioned that this should be treated as an "order of magnitude" estimation. I don't think we can draw any inference from this estimation at all, because it represents such a small portion of the problem domain. We also have no idea what the scope of the challenge is.
The entire exercise is pointless when you think about what we're asking. "Can the government actually implement a solution to record every phone call in the US?" I think that's a resounding yes. The solution would look a lot like setting up a tier 2 network provider [3] with extra investment in a storage back end. That's entirely within the realm of possibility given the size of the US Dept of Defense budget.
[1] http://en.wikipedia.org/wiki/Class_4_telephone_switch
[2] http://en.wikipedia.org/wiki/Call_detail_record
[3] http://en.wikipedia.org/wiki/Tier_2_network