Tech lead of Cloudflare R2 if that's helpful as I've thought about this a bit.
First things first, a bit of a shameless plug [1] to answer this piece:
> Basically I want to be able to use object storage with simple `fetch` and few lines of code. Without all that SDK nonsense.
const object = await env.MY_BUCKET.put(objectName, request.body, {
httpMetadata: request.headers,
})
return new Response(null, {
headers: {
'etag': object.httpEtag,
}
})
It's called `put` instead of `fetch` but basically same concept.
> Buckets is unnecessary concept. Why use it at all, when we have domain and path. That should be enough.
The utility is because it's important for accounting, organization, ACLs, jurisdiction controls, analytics, different business units needing different properties, etc. If you use a virtual-hosted URL then you have something like https://<bucket>.<account>.r2.cloudflarestorage.com/<path>. That's your domain + path. Buckets is still a useful higher-order organizational concept. FWIW Azure Blob storage and S3 were released at roughly the same time. Both had the same organizational concept. The other thing to think about is that folders are tree structures and that just doesn't have the same scaling properties to billions or trillions of files that a flat namespace does.
> Authentication is too convoluted. Signings, etc. Simple `Authorization` header should be enough.
Yeah probably. 16 years ago, HTTPS was a lot less common. Even still, some users do have legitimate concerns about middleware injecting headers / intercepting signed requests & replaying them. Sigv4 does address that in a comprehensive way. It also ensures some level of protection against corruption that can happen before TLS encryption and after TLS decryption.
> Uploading a file should be as simple as PUT /path/to/file. Instead of multipart upload, header `Range: bytes=0-1023` should be used.
I agree that multipart has various misdesigns. The one that actually really gets me is ListParts which lets you retrieve the original parts of a multipart file - like wtf that seems legit useless. At the time though (16 years ago now), resumable uploads didn't exist as a standard (unless you count WebDAV which I don't because it failed miserably). There's now an attempt to standardize something using TUS a starting point but I don't see much incentive for Amazon to really update S3 here because that's not where they're investing their R&D budget. It's up to the rest of us (particularly browser vendors) to create a standard and enough demand that they have to update their API. FWIW the range thing has to be careful because aside from resuming the other piece that's important for resumable uploads is concurrent uploads of 1 part + integrity verification. One of the things I'm thinking about is whether to try to retrofit whatever the standard becomes into our S3 endpoint or put it somewhere else. It probably makes the most sense for public buckets (aka static website hosting) but that's a whole other ball of wax (security, etc).
To defend multipart as well, the other piece it does really well is the ability to track the state of the upload which regular "PUT /path/to/file" doesn't. Downloads can only publish the object at the path once they complete so then how do you identify ongoing concurrent uploads to the same path? An overwrite upload can be thought of as a strongly consistent "delete + publish" step.
> List is unnecessarily flexible. Why should users have the ability to use any character as path delimiter. `/` should be fixed. List operation should use ordinary REST conventions, like `GET /dir/?from=token` should return JSON with items inside `dir` and continuation token.
Eh. In terms of complexity, the '/' delimiter really doesn't add any complexity to the API and "users" here are application developers, not end users. It adds a bunch of horrible implementation complexity and it took us a while to get this piece correct, but the complexity was having any delimiter, not around letting the user pick a custom delimiter. And having a delimiter is definitely helpful because people like to view folders. Would I wish that they simplified and require the delimiter to be 1 ASCII character instead of supporting arbitrary UTF-8 strings? Sure, but only from a test coverage perspective.
My observation is that a lot of the complexity from the APIs stems from two things. The first is that REST is a terrible way for accessing and mutating a strongly consistent data store. 16 years ago it made sense though and even today if you want interop with the browser, it's all you got. The second is that the more general problem of figuring out a sane API for a distributed strongly consistent database is an unsolved problem. HTTP has stood the test of time on that front but I wonder if something like cap'n'proto might be a better fit.
[1] https://developers.cloudflare.com/r2/examples/demo-worker/