I use sampling a lot for this stuff because my work usually involves prototyping and for this I usually get a statistically-significant sample from the data. I use this also as test vectors.
At this scale you can use whatever you want (ipython+pandas+numpy/R/Matlab/Mathematica...). For visualization I like to use Qlikview, but you can use Tableau, or DS3.js.
When I have the prototype / proof-of-concept the next step is productivizig this.
You'll need:
- Somewhere to drop that data ( EC2 / Hadoop HDFS / NFS...)
- Something to get the data from its origin and put it into the storage, aka ETL (extract transfer load). Usual suspects here: Informatica, Talend, Sqoop, Pig, hadoop, spark...
- Some kind of "database". If you're using structured data you can use a "columnar" database like: Sybase IQ, Sap HANA, Teradata, Netezza, or MS SQL Server 2012/14/16 with columnstore index, Hadoop HBase, etc. If you're comfortable with it you can also dump the data into files with columnar formats (like Parket).
- Something that puts that data graphically in front of the user and it's easy to work with. Some examples: Qlikview, Tableau, Sap Lumira, or DS3.js. If you want to code your own stuff, take a look at the tools data journalist use.
You'll need some engineering to tune the data architecture to follow your users natural workflow. For instance, if they only need month-to-month reports you can partition your data to reflect that.