Why wouldn't one use python for this job? Since the company already does a lot of python, and this job is actually easy to do in python (I've done something similar).
400 processes with python3.5 is way less than the 4gig on one medium instance (less than 2.5MB each). Just farm the work out to a ProcessPoolExecutor, and have a timeout on the S3 get requests. That would let you have enough resources to match the spike of 20000 requests per minute (334 per second).
A lot of an S3 GET request could be the SSL, AWS auth and such. All quite CPU intensive. So using an async framework that doesn't do SSL+AWS auth async is obviously not going to work well once the requests go up.
There's even an example in the concurrent.futures of downloading urls [3].
Made with the beautiful python3.5
import concurrent.futures
import requests
from awsauth import S3Auth
ACCESS_KEY = 'ACCESSKEYXXXXXXXXXXXX'
SECRET_KEY = 'AWSSECRETKEYXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX'
URLS = ['http://www.foxnews.com/', 'http://www.cnn.com/', 'http://europe.wsj.com/',
'http://www.bbc.co.uk/', 'http://some-made-up-domain.com/']
def load_url(url, timeout):
return requests.get(url, timeout=timeout).data
# return requests.get(url, timeout=timeout, auth=S3Auth(ACCESS_KEY, SECRET_KEY)).data
with concurrent.futures.ProcessPoolExecutor(max_workers=400) as executor:
# Start the load operations and mark each future with its URL
future_to_url = {executor.submit(load_url, url, 60): url for url in URLS}
for future in concurrent.futures.as_completed(future_to_url):
url = future_to_url[future]
try:
data = future.result()
except Exception as exc:
print('%r generated an exception: %s' % (url, exc))
else:
print('%r page is %d bytes' % (url, len(data)))
[0] http://docs.aws.amazon.com/AmazonS3/latest/dev/request-rate-...[1] https://coderwall.com/p/rlguog/nginx-as-proxy-for-amazon-s3-...
[2] https://nodejs.org/api/cluster.html#cluster_cluster
[3] https://docs.python.org/dev/library/concurrent.futures.html