S3st: Stream data from multiple S3 objects directly into your terminal
npmjs.com
npmjs.com
export BUCKET=____; aws s3 ls "$BUCKET" | tail -n+2 | awk '{print $4}' | while read k; do aws s3 cp "s3://$BUCKET/$k" -; done
If you're just wanting to do a 'grep' style action on an S3 prefix, might be worth looking into "S3 Select"for your use case instead
Does s3st support tags or other ways of identifying which files to stream other than filtering by the content of the files? Asking because I didn't see this feature in the demo.
- categorising what's inside
- checking what's used or not
Thanks!
But usually you need to run run arbitrary code against the contents of a large S3 bucket, and that gets tricky. The main problem is tracking what you've done vs. what you need to do, because if you haven't categorized your data yet, you can expect that code processing it will break.
One technique is queues in SQS:
1. Keys to process
2. Keys that succeeded
3. Keys that failed
(Regular queues, FIFO queues probably won't be useful. A queue can have an unlimited backlog, but the maximum message timeout is two weeks. That's probably more than enough time to iterate over some code in Lambda.)
Your initial lambda should be triggered by KeysToProcess, which you can initiate off a developer machine and just run ListBucket and create a pile of messages.
When the lambda is done, it passes its information to KeysThatSucceeded. (Or possibly another S3 bucket, or Dynamo, or a database, or just drop its key if you determine you don't need it.)
Point your dead letter queue to KeysThatFailed. Let the messages pile up in there until you've figured out the errors and are ready to try again.
And then you can trigger off KeysThatFailed, point the dead letter queue at a new KeysThatFailed2, rinse, repeat until you're satisfied it's correct.
S3 Inventory is basically designed for your use case. It effectively writes a full-bucket index every day.
I've been using AWS CLI sync, but it's getting increasingly slow. To the point that it seems untenable.