I worked on a similar system at somewhat larger scale (maybe 5-10x larger volume and dataset size). There's a lot of basic JVM tuning to be done. You have to monitor heap usage and GC pauses and keep manipulating heap size, perm gen size, heap usage that should initiate major GC (forget what variable this is) until you get things as smooth as possible and meet whatever resource constraints you want to meet on your hardware. In my experience, for large datasets ES really pushes the memory architecture of the JVM -- it doesn't really perform too well when your heap size is like 30g.
After that, you'll have to tune merging. The underlying Lucene storage engine chokes when it tries to merge segments that are large, so you have to tune the max segment sizes, etc.
Then, you'll have to tune your queries -- it's not feasible to do a full index scan on a large index so you have to get clever about how you pull data in chunks. If your data is timestamped and that timestamp is indexed, you can pull data for smaller time ranges which will be faster than pulling all data for an index at once.