The average call length is 1.8 minutes right now [2]. So we've got 240 mil person * 5 calls / person-day * 365 days * 1.8 min / call. So, about 788 billion minutes of calls / year at a flow rate of about 2 billion calls / day.
At MP3 compression of 128kbps, vocal data takes about 0.94 MB / minute [3]. So, close to 688 petabytes of storage data at a rate of about 1.9 petabytes / day. Seems within the realm of doable.
The problem of analysing this is in the "ridiculous parallelism" category, so they'd just be constrained by server farm capacity. Lets say they had a system with a conservative million nodes. Each day, each node would have to process ~2 GB of audio data looking for patterns. Not even challenging. If I were clever, I'd probably run a brute force audio to text on each node, then a text to symbol pattern analyser. I'd also have higher level net processes that look for patterns in calls spatially and temporally, but with far less processors.
[1] http://www.pewinternet.org/2010/09/02/cell-phones-and-americ...
[2] http://www.statista.com/statistics/185828/average-local-mobi...
[3] http://iaudiophile.net/forums/archive/index.php/t-2081.html