Amazon S3 Batch Operations
aws.amazon.com
aws.amazon.com
What you could do is use s3's inventory report feature, give the manifest generated to batch operations and handle the delete logic in a lambda. A lifecycle policy with some tagging could also fit your needs here.
Also just a word of warning, if you do have a lot of files, and you're thinking "let's transition them to glacier", don't do it. The transfer cost from S3->Glacier is absolutely insane ($0.05 per 1,000 objects). I managed to generate $11k worth of charges doing a "small" test of 218M files and a lifecycle policy. Only use glacier for large individual files.
[1] https://docs.aws.amazon.com/AmazonS3/latest/user-guide/creat...
Edit: I ask because AWS suggests a key naming convention for large object amounts to ensure that you're distributing your objects across storage nodes, to prevent bottlenecks.
https://docs.aws.amazon.com/AmazonS3/latest/dev/request-rate...
Edit Response: I've always used the partitioning conventions they suggest so not sure what sort of impact you encounter without.
https://aws.amazon.com/about-aws/whats-new/2018/07/amazon-s3...
I have billions of files. What do?
e.g. if you write 1 million 10KB files per day to S3, you're looking at $150/mo in PUT costs. If you instead write 1,000 10MB blocks, you're looking at $0.15/mo in PUT costs.
Due to S3's support of HTTP range requests, we can still request individual files without an intermediate layer (though our write layer did slightly increase in complexity) and our GET (and storage) costs are identical.
If you buy a server and run a poorly architectured system on it you note that it does not perform and need to make changes.
If you use serverless and run a poorly architectured system on it you pay and you need to make changes (after someone noted the bill). Yes, there are cost reports but they are not easy to use and understand. With a performance bottleneck the system limits while you are trying to understand the performance measurements. In the cloud case you are paying while trying to understand what is wrong.
Of course in a big corporation money does not matter to a software developer. But in a small company the bill paid to the cloud provider might have a direct impact on the company being able to pay your salary in the near future.
The examples you provide are not equivalent. It’s more like “we have a poorly architected software, so we had to buy 200 dedicated servers, because we didn’t know/couldn’t make it work on 10 of them”.
On cloud you could simply update your software and then downscale. Of course you pay more for flexibility, but please stop with those strawmen.
Wow thanks!
You’re talking about API level batch calls. This is about simplifying workflows which rely on Listing every object in S3 and doing “something”.
Cool feature otherwise.