Buckets and objects are not enough
sagi.org
sagi.org
In practice, many teams use S3 directly without any layer on top. So without better organizational capabilities, they can't keep track of what they have stored where, who created it, whether it is still used, etc.
And when teams do use a catalog, it's usually detached from the storage layer itself, so you can't easily view a dataset in the catalog and know how much it costs, who accessed it, and so on.
Have you seen better places that figured out a better way to handle this? Without a ton of custom tooling?
They should also accommodate my need for all POSIX filesystem API’s included cheap-moves and renames!!!!!
/s
S3 is an object store. Treat it more like a KV store. As other comments have pointed out, the solution here is pick-your-favourite-metadata-store, be it Postgres, or what iceberg does, and other data on S3.
But that's a different thing than what the post is about. Even teams that use prefixes for performance don't have an S3-native way to ask what a prefix represents, who owns it, whether it's still accessed, and so on. The semantic layer is missing whether you're hashing for throughput or just laying data out the obvious way.
[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimi...
Think of a team doing ML, for example. They work with data all day across many different tools, each reading some inputs from S3 and writing outputs to S3. They won't create a bucket for every output, that's not practical. So they write to a single bucket with outputs organized under prefixes.
Buckets are more of an administrative boundary (IAM, cost, replication) than a data organization unit. So even with more buckets, the dataset abstraction is still missing - there's no good native way to track what a prefix represents, who created it, whether it's still accessed, how much it costs, etc.
Just store such metadata in your database, where you can organize, index and aggregate it whatever way you like.
1. Invoice PDFs. Individually small, but there are a hundred million of them. Deleted after 10 years or when the tenant deletes their account.
2. Reports and exports. Few but potentially big files. If an export logically consists of multiple files, it's stored as zip file. Live 30 days or until the tenant deletes their account.
3. Streaming database exports using AWS Database Migration Service for replication into Snowflake
Every file has an entry in the database tracking its storage location and status.
Grouping them by tenant, (sub)type or time-interval makes sense for these. But "dataset" isn't an applicable concept.
The post is more about the pipeline / ML / log / export world where ownership isn't enforced by application code.
The DMS case sits somewhere in between - there's a per-table grouping that could be useful, but the files are usually transient enough that it doesn't matter much. Different problem from yours.
But that's just one slice of storage. Most teams also have logs, media, ML artifacts, raw dumps, etc., none of which fit into a table format. And even with tables, you often can't easily look at a Delta table and know what the underlying storage is costing you, whether it's still accessed, etc.
Another system might solve it for your media files, another for your log streams, and so on. That's the thing, you have a set of management nice-to-haves that are quite generic and aren't universally supported today, so you end up reinventing them separately across each domain. And even if you did, you still wouldn't have a central aggregated view across all your storage.
You would be appalled at the kind of stuff I have seen teams stuff into parquet and iceberg tables.
The solution is just to spin up the machinery you need for your solution, rather than making S3 cover all possible bases.
I could imagine it though that S3 could offer something similar. We can already list the bucket items, why not add some of querying ability?