To be fair to S3, Hadoop pretending S3 is a hierarchical filesystem is a bit of a hack. But I had cases where new objects wouldn't be listed for hours. There's only so much you can do to design an application around that, especially when Azure and Google Storage don't have that problem.
The first one happened often enough to cause problems, the second one was a fairly rare event, but still had to be handled.
It might have been more of an issue with "large" buckets. The bucket in particular where I had to dance around the missing object issue had ~100 million fairly small objects. I ended up having to create a database to track if objects exist so no consumer would ever try to see if a non-existent object existed.
Time to revisit all that mess, I suppose.
1. LIST 2. PUT 3. LIST
would trigger situations where (3) wouldn't include the object inserted in (2). This is well-known however.
1. HEAD key -> 404
2. PUT key -> 200
3. GET key -> 404 (what? But I just put it!)
This is commonly used for "upload file if it doesn't exist"i remember a time when if you were using the us-standard “region” and were unlucky it could take 12 hours for your objects to become visible
I've observed eventual consistency with S3 prior to this change where it took on the order of several hundred milliseconds to observe an update of an existing key.
I've observed this as recently as 2 years ago.
That's why manifest files became so popular.
like if you had tried to read a non-existent key, then wrote to it, it might continue to appear to not exist for a minute?
After a write, you would _always_ be able to read the key you just wrote.
After an update, you could get a stale copy of the key if your subsequent read hit a different server.
Anyway I am glad to see these gaps and caveats have been closed.