How to Save 90% on your S3 Bill
appneta.com
appneta.com
I noticed that no such Issue exists, so I opened one. https://github.com/boto/boto/issues/2078
All they can really do at this stage is add a warning to the documentation and hope that new people using the library figure out the significance.
Totally agree, that this should be default False!
It looks like a HEAD request can fulfill the same purpose at a much lower cost (this is mentioned by a comment on your Issue: (comment by kislyuk)
http://docs.aws.amazon.com/AmazonS3/latest/API/RESTBucketHEA...
Shall I make the change and pull request it? I didn't want to duplicate effort.
> 09:00:04 < saurik> dls: I am performing millions of ListBucket request to this one amazon bucket every day
> 09:00:25 < dls> d'oh
> 09:00:28 < saurik> adding up to almost 1.5 gigabytes of data traffic in/out on just those requests
> 09:00:41 < saurik> I have NO CLUE what could POSSIBLY doing even a SINGLE ListBucket request on that bucket
> 09:00:49 < dls> LOL
Regardless, it's just a caching layer, and requests are passed through to the API in that case.
Send me an email nathan@nathancahill.com and I can show you how to set up a local cache. It's quite simple.
The tool they're actually referring to is Ice (also by netflix) [1]. Asgard is another AWS-based tool for managing deployments and auto-scaling.
(disclaimer: I used to work at Smore, and I'm friends with the author, but I've been burned by Boto myself)
https://github.com/boto/boto/commit/95939debc3813468264159d5...
EDIT: Looks like the original committer is an Amazon employee.
I wonder more if this was a product of Amazon developers using S3 (i.e. dogfooding) and not noticing the cost side effect because I'm assuming they don't get billed?
I also pay for my own personal EC2 instance and about 350 GB of S3 storage. Begin a genuine user and customer of AWS helps me to be a better employee.
You may want to check LIST request statistics over the next few weeks. Between this thread, an Issue for boto, etc. I'm curious if you see a noticeable decline in LIST requests with the attention this has brought. I'm just curious from a data standpoint.
It's a commit which specifically added validate to allow skipping validation. If you read the diff, the call originally unconditionally performed the validation call.
It is defaulted to true so that old code that called get_bucket() won't break. (Since old code would have used get_bucket with only one parameter.)
It was always true before that commit.
Therefore, in essence, an Amazon employee made it possible to save 90% of your S3 bills.
https://github.com/boto/boto/commit/95939debc3813468264159d5...
The boto guys have justified it in the past: https://groups.google.com/forum/#!topic/boto-users/1DVfbo4CD...
I still don't agree with their reasoning to leave it default, once you try to do something on a non-existent domain/bucket it will throw an error anyway, and I would argue the "extra work" is much cheaper than leaving these defaults on which I would expect are completely redundant for most users.
https://bitbucket.org/david/django-storages/src/cb7366693ce1...
The setting in question is: AWS_AUTO_CREATE_BUCKET.
And the conditional inside the except is useless, of course, since you can only be in the except if `auto_create_bucket` is set.
or you can say the fee is bundled, i.e.
(Rackspace) vs (S3 us-east-1)
Storage: 0.105/GB/mo vs 0.085/GB/mo
Bandwidth out: 0.20/GB/mo vs 0.120/GB/mo
I heard a talk by someone at a mega tech company that has their own internal cloud for their teams, and they "charge" each team based on usage. One team stored lots of file with 12,000 character filenames with zero contents. Since the company only "charged" for file size, that team had a tiny charge!
Yes, it is still a waste of money, but just make sure you understand that it's not actually listing your entire bucket.
I ran into this recently when making my own s3 sync tool, because the commonly used tool is completely broken (requires something called a 'config file' to function). But I didn't pay it too much mind, because I forgot the price discrepancy for ListBucket calls.
PS if you want to see what boto is doing do this: logging.basicConfig(filename="boto.log", level=logging.DEBUG)
It does not prefetch any key (maxkeys is set to 0), it performs a query on the bucket to validate that the bucket exists and blow up if the bucket does not exist. With validate=False, you can call get_bucket and get a bucket object where no remote bucket exists.
Two manual work-arounds that come to mind:
- store a list of created buckets as keys in another bucket.
- store a dummy file in each bucket you create.
Either method allows you to check the existence of the bucket with a GET request rather than a more expensive LIST request, but both are hackish. It seems like this is functionality S3 should already provide cheaply.
It looks like there now is: http://docs.aws.amazon.com/AmazonS3/latest/API/RESTBucketHEA...
I'm guessing (hoping?) that didn't exist back when the feature was added to Boto, 7 years ago: https://github.com/boto/boto/commit/8410c365ee0120e073bf00bd...
I don't worry about cultivating a connection between engineering and business needs.