This is what I was getting at, mostly:
> Any fragment which is missing the "ok" file, is ignored by the reader [3].
As a client, how do I detect this? Say I just wrote some data and now I want to generate a new total. Could my data be silently missing?
> TileDB handles corrupt or incomplete fragments by erroring out on the read. In the future, we could offer a retry mechanism for failed reads, but the important thing is we will never return corrupt or invalid results.
But now I have to handle this everywhere I interact with TileDB. Most users don't expect essentially random errors. Retries are a stop-gap, as it can take arbitrary time for consistency to converge. I've observed O(hours) regularly.
> Incomplete fragments can happen because of S3's eventual consistency; it is possible for the atomic ok file to show up before an object in the fragment "folder".
Indeed! And from what I understand, mostly from conversation with the Iceberg team last year, they solve this entirely by using a central ACID service backed by RDBMS (which, aside, I find very amusing!! They never listen to Stonebraker...).
Indeed, as far as I can tell, the only way to remove consistency issues is to: - store an index to objects somewhere central - never list the bucket (or speculatively stat, etc.)
Until either a) you support a central index or b) S3 has a better consistency model (not that DynamoDB garbage Hadoop uses) I would be very reluctant to use TileDB if it wasn't using my own disks.
I think for small stuff, it'll work great, but at scale you're going to hit some roadblocks. I recently encountered splitting a bucket into to many buckets and prefixing UUIDs at the root of the bucket to reduce consistency problems caused by S3's rebalancing, and it ended up being a very expensive marginal improvement.