1. Faster crypto impls for the platform (this will likely be the biggest win in terms of cpu usage per upload, if you're doing client side enc).
1a. Use Google ConScript as your JVM SecurityProvider, for instance.
2. Attach S3 Gateway to your VPC to avoid having to traverse through the internet to hit S3 front-ends.
3. Force resolve DNS per S3 upload request at request time. S3 vends different answers depending on the load at the front-ends.
3b. Assumption is that the cost of just-in-time name resolution + creating new connections would be offset by higher throughput due to connecting to S3 frontends with lower load on them.
4. Pre-warm your S3 bucket(s) to achieve higher throughput (Number of S3 partitions allowed per bucket is essentially infinity but carry a cap of 3000 write IOPS per partition, iirc).
5. Avoid tiny files like plague. You could try to zlib them up, but then retrieval isn't trivial anymore, and requires compute.
6. Use EC2 with enhanced networking that has dedicated 25GbE/40GbE links going out to S3, if you must use EC2.
7. Stream/Batch APIs wherever you can use them (delete, multi-part upload).
8. Latest AWS SDK.
Of course, there are new tools and services in town: S3 SDK TransactionManager API, S3 Accelerated Buckets, S3 DataSync, S3 SFTP, and S3 BatchOps that are much more hands-off and preferable to hand rolling your own highway.