Deleting an S3 Bucket Costs Money
cloudcasts.io
cloudcasts.io
https://stackoverflow.com/questions/59170391/s3-lifecycle-ex...
What I found out when I researched it is that there is a subtle difference between using lifecycles to move objects to other storage classes and for deleting objects: deletions are not transitions, they are expirations – and expirations are free. I submitted a clarification to the S3 documentation and now it says "You are not charged for expiration or the storage time associated with an object that has expired." (https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecy...)
If you have objects in IA or Glacier there is a minimum duration you're charged for, but there will be no extra charges for expiring these objects.
[0] https://aws.amazon.com/about-aws/whats-new/2021/09/amazon-s3...
"Minimum storage duration charge" along with "Minimum capacity charge per object" is a property of an S3 Storage Class, not a lifecycle policy. Therefore this cost needs to be considered before selecting a storage class.
The charges for deleting before the minimum storage duration are documented for each storage class at https://docs.aws.amazon.com/AmazonS3/latest/userguide/storag...
Refer to the "Performance across the S3 Storage Classes" table at https://aws.amazon.com/s3/storage-classes/
"If you create an S3 Lifecycle expiration rule that causes objects that have been in S3 Standard-IA or S3 One Zone-IA storage for less than 30 days to expire, you are charged for 30 days"
It goes on to say 90 days for Glacier and 180 days for Glacier Deep.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/lifecy...
I ran into this with a bucket full of EMR log files a few years ago and had to figure out some pretty crazy command line hackiness, plus running on a EC2 machine with lots of cores to figure it out. This a write-up I did if anyone else ever runs into this issue.
https://gist.github.com/michael-erasmus/6a5acddcb56548874ffe...
When we did this (a few years ago) it still took several days for it to remove all the files.
But yes, it can take a while to iterate through the objects.
It won't effect it, because there is no concept of nesting in an object storage engine like S3. Everything is a flat key that references an object, but we just abstract and conceptualize a directory structure because it makes it easier for us to manage our data.
But in reality you just have a really, really long list of keys and a big, flat file system of objects.
When this bit us on a project I made a tool to solve our particular problem, which tars files, writes csv indexes, and can fetch individual files from the tars if need be.[1] Running on millions of files was janky enough that I also ended up scripting an orchestrator to repeatedly attempt each step of the pipeline.[2] Not tested on data other than ours but could be a useful starting point.
[1] https://github.com/harvard-lil/s3mothball [2] https://github.com/harvard-lil/mothball_pipeline
AWS is designed to extract dollars from big enterprise contracts.
Also interesting from the article, this poor soul on StackOverflow was trying to figure out how to delete a bucket that would cost him $20,000 [2]. Can’t delete, can’t close.
[1] https://www.reddit.com/r/aws/comments/j5nh4w/ive_deleted_my_...
[2] https://stackoverflow.com/questions/54255990/cheapest-way-to...
When we reactivated the account a few yearss later, we were retroactively billed for all the files in the s3 bucket. We got the money back though.
1) Upload an encrypted blob to an S3 unique AWS account with a burner credit card.
2) Cancel the account.
3) If you ever need to restore the data, restore the account and pay the difference.
Since uploads are free you just do this to a new account every so often and you’ll only need to pay the time difference of the most recent backup!
Scummy company
But I had moved to Singapore. My PayPal account was in Australia. When trying to pay using my Singapore card it got flagged as fraud. I called and there was no way to pay. In the end I created a 2nd account in singapore. Added $2 credit to my account. Transferred it to the other account. Then closed both. Ive not used PayPal in about… 6 years now? Scam company with unhelpful support.
Amazon can keep trying to get blood from a stone...
Sorry to hear that you had a bad experience. That's definitely not the culture that I'm experiencing every day or that we hold ourselves to. Our goal is to deliver a ton of value to customers. If we are not delivering value then it would not be in line with how we operate to want to charge you for that.
Out of curiosity, did you try contacting AWS support? Almost all customers I speak to love our support and the fact they can speak to an actual human. Of course I can't speak to the actual case as I'm just providing my personal opinion here, but I would not be surprised if support could fix this for you quickly and zero out any small balances that you accrued because you weren't able to access/use your account.
Spending more to provide better customer service (explicitly or implicitly through bugfixes) is not generally how this type of need is. Folks shouldn't need to post their grievances to a very public site to be heard, and that it needs to happen to get traction on an issue should be taken as a clear gap in customer service provision.
Support will routinely give forgive an account balance given a good enough reason.
> I've had to ask my credit card company to block AWS charging an account that I can't remember the login to.
We don't know whether OP has reached out to AWS support, but one can assume so since AWS is charging an account the user cannot log in to.
Further, it sounds like the user has found their own resolution, so the user isn't asking for _help_ on HN, more that they are registering poor experience in agreement with others. So, it is a rough UX edge.
> Folks shouldn't need to post their grievances to a very public site to be heard, and that it needs to happen to get traction on an issue should be taken as a clear gap in customer service provision.
Yet by your very own comment you say that we don't know if he actually tried to contact support so there is no "clear gap in customer service provision".
The fact is, anyone who actually interacted with an AWS rep would tell you that they would delete the bucket and not charge you for it.
This is both on retail side and AWS side.
My only request - do some more automated discovery and remediation tooling for VPC-classic accounts! It's annoying to figure out what needs to be handled there - give me a dashboard for this.
I've been paying $10/month for over a year on an EC2 (I think?) instance I spun up, and I still haven't had the time to go in and find it and delete it. I've tried several times (and even disputed the credit card charge one month).
I'll probably just end up cancelling this credit card. That would be a lot easier than figuring out how to stop AWS from charging me.
Nooooo no no no that is not true. If you cancel, this debt will follow you on your credit history and via collectors. Canceling a CC does not waive you of responsibility.
Just contact AWS support. Tell them your story. They’ll get your account closed and they’ll probably forgive whatever outstanding bill you have.
Please don’t just run away from the bill, you’re just setting yourself up for pain later.
At this point, they're legally committing fraud. I don't see how closing a credit card that has regular fraudulent transactions is wrong, or likely to cause problems later on.
If some random wants to contact amazon to delete our business AWS account they need to reject that as well.
I would login, open a account and billing support case.
Do Account / Close and cancel my account
They will give you a chat or call back option in most cases. The call back option has worked well for me.
> They are absolutely not commiting fraud - that is a total lie.
Lie isn't the word you're looking for when you disagree with someone.
If they never respond they aren't engaging with the OP to figure out what this person want's. If the OP even contested the charges AWS has been warned that someone isn't happy with the contract. Both by personal email and by the credit card company.
If I write an email to them asking them to cancel an account, doesn't that constitute legal termination of my authorization to continue service or charge me?
Unless the service agreement I signed specifies a specific way to cancel an account - in that case my agreement to it constitutes authorization unless I jump through their hoops. Argh.
Remember, canceling your account means everything (EVERYTHING) is deleted. They are far far more concerned that someone will randomly email them pretending to be the CTO of some startup or you, and they will blow away 3 years of someone's work.
Canceling your account even cancels glacier, WORM records, object and compliance lock data etc.
Everything I've seen says that AWS rightfully biases towards retention unless the delete request is very clear - and email will never rise to that level (nor should it!).
So you have to login and close your account yourself.
How you get from a very reasonable business practice here (and they are frankly one of the easiest of the major players to actually talk to) to fraud... is a stretch.
Edit: For those not familiar with AWS account closure here are the steps:
https://aws.amazon.com/premiumsupport/knowledge-center/close...
That should give you a breakdown of where the charges are coming from, specifically which region an EC2 instance might be running in. Once you can figure out which region it's in, you can go to the EC2 dashboard for that region [2] and start terminating instances.
Having said this, AWS should seriously have a way to say "close this account and delete/stop everything associated with it." Having to spelunk through bills and the AWS console to figure out what you're getting charged for is a joke.
[1]: https://console.aws.amazon.com/billing/home?#/bills?year=202...
[2]: https://us-west-2.console.aws.amazon.com/ec2/v2/home?region=...: (update both us-west-2 with whatever region you're looking for)
Also, if they do send your debit to collectors, you can go to a court and ask for the debit to be dropped and for them to pay damages. That is a lot more work, but if they are clearly on the wrong will get you something.
Theoretically. I don't know in what cases Amazon might do either of these things. If you're not in the US, that's certainly another layer of barrier.
But yes, in the US anyway, whether you legally owe someone money or not and they can legally collect on it is not controlled by whether you've cancelled a credit card.
When does billing by the hour for storage/everything start to seem appealing?
In our case, it was a message saying something like ‘container timed out’ originally making us think it was a network issue. We ended up tracing it back to our app sometimes taking 21-30 seconds to fully boot, instead of the 20 second limit Fargate had for health checks. So even though the container did end up booting, Fargate was already waiting for it to start so it could kill it.
I mean, the issue was my fault, but AWS made it incredibly painful to figure out and fix and was not helpful at all. I decided that day that I will never use AWS again, unless I have someone with a lot of AWS experience on the team (and even then, only if I have a good reason to use AWS, in my case there is a reason why AWS would perhaps be desirable in the future but not enough to go through the pain again myself).
I haven't heard of Amazon specifically ever doing either of these things.
Or am I missing what you're saying is particular to Amazon and relevant here?
The alternative is the tiers which have their own problems.
Small Medium Big Call us
No one know if they are gonna have success in the beginning or need to scale up. Usually. I guess at a size like AWS it makes more sense to have those infinite pricing tiers and let the customers figure out how much they are willing to pay rather than try to negotiate with huge swaths of people that take sales staff salary to deal with.
It’s amazing how often it retries.
Example from go sdk: https://github.com/aws/aws-sdk-go/blob/main/service/s3/s3man....
If there are still those who do not use the cloud, it is because the big three have taken advantage of their position a lot.
The pricing of Hetzner, CloudFlare, Linode, OVH, ... seems to be cheaper and more transparent.
If you don't use special tools not available elsewhere, such as AWS' SageMaker, or Google's TPUs, ...., then it's probably not economically interesting to use the Amazon, Microsoft or Google clouds.
I spinned up an AWS instance to practice, and once I was done I thought I closed everything down.
Turns out I had just stopped my micro instances, and I didn't terminate them. I also hadn't released the my IP address. There was also a snapshot of the tiny db I had created still floating around. The documentation was a little confusing, so after I went through it I spent half an hour chatting with a support rep to make sure everything was completely good. After next month my last bill should go through and I should be free and clear. Unfortunately I have to wait for next months bill to go through as I can't just pay it all now.
This was mostly my fault for letting it go on for so long, but I hate how if you don't do some very specific steps you can still be charged. And I think if an account is closed, it should absolutely terminate all services that are still running on that account, and then send you the final bill.
I think in practice, S3 data is often indexed using other DBs e.g DynamoDB, Postgres, MySQL etc. Can't this index be used to enumerate all S3 URLs? I am off-course simplifying this a lot.
This specific issue probably isn't a very big problem.
The issue of Amazon repeatedly coming up on HN as a service that will bill you when you're not unexpecting it for things that are moderately hard to understand and might refund you later probably costs them tens or even hundreds of millions in lost revenue every year from developers being cautious about deploying things to their services.
My experience with AWS is that the pricing for each service is reasonably well documented and the calculator does descent job. The problem starts when multiple combinations of services are used and it becomes harder to reason.
With the advent of cloud, cost-modelling becomes an essential skill (which can be learned). One needs to be clear about total work that gets done and "how" that work gets processed. This in turn should translate to relevant cost metric (e.g PUT requests/s for S3 or IOPS for DynamoDB, amount of data scanned for Athena, etc)
This needs to be evaluated for zero load, normal load, 5x load, 20x load, etc. Zero load gives what is dead weight cost of the system i.e cost incurred when no work is being done (e.g EC2, EBS volumes, etc)
I see this as a good thing. They literally encourage you to be cautious with your pricing and resource usage, to the point where they put limits on what resources you can use without explicitly asking for more.
Developers should be cautious and aware.
People just playing around should be careful. There are ways to keep risk lower. But if absolute price caps are your priority you should probably be using a VPS of some sort.
In practice, activating that free tier requires a valid card, and I'd highly advise never giving them your own. Whatever alternative you can think of is 100% better for your sanity.
Correction: I misread - .5¢ per 1,000,000 items LISTed
.5¢ per 1000 LIST operations
LIST operations max out at 1000 items
Still a little pricey, but way less so than I'd imagined.Do they make a lot of money off of charging for basic operations? It seems like you could make the whole pricing structure a lot more friendly by only charging for bandwidth use. I guess when you're as dominant as S3, you don't need to care about friendly pricing structures.
Charging for basic operations like that is weird, it's akin to a service charging people per number of clicks on a website.
It sucks, but the fact is AWS/GCP/Azure aren't really designed for hobbyist, they are designed for massive corporations. Their free tiers exist merely as a service to help train professionals to use their platforms.
Luckily, there are still good, low-cost providers out there.
Note it’s $0.005 per 1k requests, not $0.05 per 1k items -- that’s an extra zero from what you said, and also important to point out that one request can list 1k items. So if you list in 1k batches, it’s $5 per million items listed.
It's $0.005 per million items listed. A thousand requests of a thousand items each is your million items, and a thousand requests is $0.005.
source : https://stackoverflow.com/a/67834172
2M LIST requests = (2B objects / 1000 per LIST)
$1 = (2M / 1000) × $0.005
At least that's my reading of the Pricing page for US East (Ohio).> Here's what it does: It calls a LIST on the bucket, pagination through the objects in the bucket 1000 at a time. It calls a DeleteObjects API method, deleting 1000 at a time.
> The cost is 1 API LIST call per 1000 objects in the bucket. Delete operations are free, so there's no extra cost there.
source: I read the article.
"S3 Batch Operations" sends S3 requests based on a csv file, which can but does not have to be from S3 Inventory. But S3 Batch Operations supports only a subset of APIs and this does not include DeleteObject(s). [0]
An AWS Batch job could run a container which sends DeleteObjects requests but only when triggered by a job queue which seems redundant here.
If I can't use an expiration lifecycle policy because I need a selection of objects not matching a prefix or object tags, I would run something with `s5cmd rm` [1]. Alternatively roll your own golang which parses the CSV and sends many DeleteObjects requests in parallel goroutines.
0. https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
https://docs.aws.amazon.com/AmazonS3/latest/userguide/storag...
So S3 inventory would be half price compared to LIST (or quarter price in IA storage class), but that's still small comfort if you're staring down the barrel of a bucket containing a large number of objects.
[1] Management & analytics tab on https://aws.amazon.com/s3/pricing/
> LIST requests for any storage class are charged at the same rate as S3 Standard PUT, COPY, and POST requests.
I read this as LISTs do not cost double for infrequent access, even though other Tier 1 requests do.
AWS's own pricing calculator doesn't split out LIST requests: https://calculator.aws/#/createCalculator/S3
Either way, the takeaway is that using LIST or a bucket inventory, will still be O(N) cost and there's only a factor-of-2-ish difference between the two.
Then again, a billion objects is $5 territory to delete, and if you have a trillion objects to delete and no pre-existing listing to go off of, then odds are you can stomach the $5000 hit more easily than you could stomach the staff time spent trying to reduce that cost!
Is also mentioned in the article though they don’t calculate the price.
I guess one could spam DELETE calls while bruteforcing filenames to make it free.
It's listed in the post as "extra credit", but its trivial having done it myself for a client. A bit disingenuous to state, "if you like to do things the hard way." It literally takes less than hour to do, and you can preserve the inventory report for housekeeping if needed.
[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/storag...
[2] https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
(OP: feel free to steal this comment's info if you want to update your post)
0. https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
[1] "Specify the operation that you want S3 Batch Operations to run against the objects in the manifest. Each operation type accepts parameters that are specific to that operation. This enables you to perform the same tasks as if you performed the operation one-by-one on each object."
https://docs.aws.amazon.com/AmazonS3/latest/userguide/batch-...
Yikes. An hour just to delete some files in a folder? It seems to me that S3 should be avoided unless you really, really, really need it.
If your use-case is storing random things you don't know the path of, maybe it's the wrong product to use.
You can retrieve all objects with a given prefix, which is great for storing content-addressed files, and being able to iterate on them. You can also partition on arbitrary prefixes too.
Nowadays, it (¿almost?) is. https://aws.amazon.com/s3/consistency/:
“After a successful write of a new object, or an overwrite or delete of an existing object, any subsequent read request immediately receives the latest version of the object”
I think that says that deletes are immediately visible, too, but they phrase it weirdly, as, after a delete, there is no latest version of the object.
Also, I don’t think buckets are objects in this sense, so the caveat in the article stands.
I think last time I did this, the wait time was pretty much exactly 60 minutes.
Based on some quick maths, deleting a million files would only cost you like $5.
P.S. Again its silly they do this and I'm probably greatly underestimating how these costs can add up for mid to large orgs.
As the article says DELETEs are free and you can do bulk deletes of 1000 objects at a time. However you need to have the object names. You get those using LIST, which gives you 1000 items for each request. LISTs are currently priced at $0.005 per 1000 List requests. So $0.005 to delete 1M objects. Using the “empty bucket” feature does this internally and charges you that exact same amount.
The only way you get near your price is if you try to delete by applying a new lifecycle policy to 4B objects that are not in Standard storage.
If you can select your objects with a lifecycle policy (by object tag or prefix) you don't need the LISTs either. The prefix can be "" to select all objects. Just be careful with that.
Right now, anything less than 400KB should really be plonked into its cousin, DynamoDB.
Other than that, I fully expect AWS to announce a new S3 bucket type (and pricing) for high-volume, small-size blobs. There is also a small matter of addressing Cloudflare R2, which should result in a Lighsail-EC2-esque fork of S3.
thats really really cheap.
This is partially incorrect. I can recreate it immediately in the same account, but in different account, I need to wait for ~1 hour
I work on https://www.vantage.sh/ which helps teams get visibility on their cloud costs which may be helpful to folks here as well on this topic.