S3 is showing its age
materializedview.io
materializedview.io
Adding features makes documentation more complicated, makes the tech harder to learn, makes libraries bigger, likely harms performance a bit, increases bug surface area, etc.
When it gets too out of hand, people will paper it over with a new, simpler abstraction layer, and the process starts again, only with a layer of garbage spaghetti underneath.
Show your age and be proud, Simple Storage Service.
The CAS thing is also pretty essential. Everybody wants to build some simple storage system on top of S3 and without CAS it’s pretty damn hard to have any kind of consistency guarantees, even for very simple systems… you end up having to build something outside S3 to manage consistency. An easy to understand use case is backups. Suppose you are using S3 as a backend for a backup system which deduplicates backups. You want to make backups from multiple locations, you want to deduplicate because there’s a lot of duplicate data floating around (maybe you have terabytes of video files getting copied, ML model weights, something else big that you copy around), and you want to expire old backups. You can almost build this on top of plain S3, and the only reason you can’t is because it’s unsafe to expire old data in such a system if any backup is writing (because the other backup may add a new reference to data, racing against the expiration / garbage collection process). A simple CAS gives you a lot of tools to solve this. The alternative to CAS is doing something kinda silly, like running a DynamoDB table as a layer of indirection.
Neither of these things add much complexity to S3.
(I think append is less useful and potentially a lot more complicated, both in terms of its API implications and in terms of the underlying complexity. If I want “append”, then I can use multipart uploads, or just upload multiple objects and reassemble them on the client side.)
I did this, I just don't expire any object created in the past week.
If a backup and a GC race then a file can get both referenced and marked at the same time, but then a future GC will see the references and put the file back into a normal state. Assume other operations can still find the file while it's marked.
Are there benefits to CAS for this situation other than resolving faster?
> Are there benefits to CAS for this situation other than resolving faster?
I think this kind of thing comes up a lot, where you’d find it convenient to have a CAS update for your file. Like, maybe you should be using a database, but you’re already using S3 and having one or two CAS operation would mean that you can stick with S3.
Sometimes, the alternative is a little ugly. Like, “I’m going to create a DynamoDB table, and it’s only going to contain one row.”
What I’d really love, even more, is to have some kind of distributed lock service on AWS. Something like Zookeeper or Etcd as a SaaS product, where it’s cheap just to get a couple distributed locks. Feels like a gap in cloud offerings to me, but I can understand why it’s missing.
It's just that a backup tends to have mostly immutable files sitting around, so it becomes more niche. It's awkward to do a lock but you don't need a lot of locking.
LIST, GET, and PUT are strongly consistent, the file name is the lock name, write the owner id and expiry timestamp in the file, and periodically extend the lock expiry (heartbeat). If an other process finds an expired lock delete the file.
For CAS, one example is backup jobs. You can run backup jobs to S3, but there are some safety issues if you want deduplication and you want to expire old data.
> if S3 is too simple
CAS isn’t some kind of super complicated, technical thing.
It would be nice if S3 had this small, incremental additional feature. That’s all. It would mean that some people don’t need to fire up DynamoDB just to do something you can already do in, say, GCS.
The only essential thing it needs to do is store my files, with some assurity that they will exist x years from now in a cost efficient manner.
if we are talking about less than 1% of the applications, probably yes.
What does 'simple' even mean when you talk about authz?
I'm pretty happy that there are S3 compatible stores that you can host yourself, that aren't insanely complex.
MinIO: https://min.io/
SeaweedFS: https://github.com/seaweedfs/seaweedfs (this one's particularly nice and is permissively licensed, in contrast to everything else)
There was also Zenko, but I don't think they gained a lot of traction for the most part: https://www.zenko.io/
Of course, many will prefer hosted/managed solutions and that's perfectly fine, but at least when you run software yourself, you are more in control over it and for the most part can also make the judgement on how hard it is to operate and keep operational (e.g. similar to what you'd experience when running PostgreSQL/MariaDB/MySQL or trying to run Oracle).
That said, my needs (both in regards to features and scaling) are pretty basic, so it's okay to pay any of the vendors for something a bit more advanced and scalable.
Here's simple protocol:
PUT /my/key
Content-Type: plain/text
Hello, world
GET /my/key
Did you ever tried to use S3 without libraries? Did you ever checked size of AWS SDK? It's incredibly overengineered.$ curl \ -H 'Content-type: text/plain' \ --aws-sigv4 'aws:amz:eu-west-1:s3' \ -u "$AWS_ACCESS_KEY_ID":"$AWS_SECRET_ACCESS_KEY" \ -H "x-amz-security-token: $AWS_SESSION_TOKEN" \ -XPUT --data 'hello world' \ https://mybucket.s3.eu-west-1.amazonaws.com/my/key
$ curl \ --aws-sigv4 'aws:amz:eu-west-1:s3' \ -u "$AWS_ACCESS_KEY_ID":"$AWS_SECRET_ACCESS_KEY" \ -H "x-amz-security-token: $AWS_SESSION_TOKEN" \ https://mybucket.s3.eu-west-1.amazonaws.com/my/key
just works.
I don't like putting words in other people's mouths, but that really does seem like a fair paraphrasing of your comment.
https://github.com/Peergos/Peergos/blob/master/src/peergos/s...
Because EC2's hypervisor was (when it was launched) lacking features (no hot swap, not shared block storage, no host movements, no online backup/clone, no live recovery, no Highly available host failover ) S3 had to step in to pick up some of the slack that proper block or file storage would have taken.
For better or for worse, people adopted the ephemeral style of computing, and used s3 as the state store.
S3 got away with it because it was the only practical object store in town.
The biggest drawback that is still has (and will likely always have) is that you can't write parts to a file. its either replace or nothing.
But that's by design.
So I suspect it'll stay like that.
would it be better to merge ddb and s3? maybe, maybe not.
gcp/azure/etc are there to provide fancy cloud solutions. no need for aws to serve that market as well.
The blocks of a filesystem can now be objects, replicated or erasure-coded - like Ceph running filesystems on top of its low-level object storage protocol, which is done on raw disks, not a filesystem.
This can't be done for something like Minio just running on your filesystem, but if you're building the storage system from the ground up it can.
We will see more and more of these products appearing.
Vast Data is another interesting one. Global deduplication and compression for your data, with an S3, block, or NFS interface. Storing differential backups for thousands or millions of VMs? You'll only store the data from base Ubuntu image once.
>> Because EC2's hypervisor was (when it was launched) lacking features (no hot swap, not shared block storage, no host movements, no online backup/clone, no live recovery, no Highly available host failover ) S3 had to step in to pick up some of the slack that proper block or file storage would have taken.
I don't want any of that nonsense in my compute layer or an application (at scale) that relies on shared block storage or host movements or live recovery.
I'm sure S3 append would be super handy.
transparent HA means that I can fail over services to other regions without having to get the programmers to think about it. Most of the busy work at scale is managing state or, more correctly recovering state from broken machines.
If I can make something else do that reliably, rather than engineer it myself, thats a win.
So much of the work of standing up a cluster (be it k8s or something else) is getting to the point where you can arbitrarily kill a datastore and it self heal.
If you're talking about s3 partial updates, its about cost/and or performance. If you dealing with megabyte chunks, and you want to flip a few bytes over hundreds of thousands, thats going to eat into transfer costs.
Sure you could chunk up the files even smaller, but then you hit into access latency (s3 aint that fast. )
Reliability is always your problem not something to be punted to another layer of the stack that lets you pretend stuff doesn't go wrong.
yup, which is why relying on devs to engineer it is a pain in the arse. Having online migration is such a useful tool to avoid accidental overloads when doing maintenance, its also a grat tool to have when testing config changes.
Currently I work at a place that has its own container spec and scheduler. This makes sense because we have literal millions of machines to manage. but thats an edge case.
For something like a global newspaper (when I used to work) it would be a massive overkill, we spent far too long making K8s act like a mainframe, when we could have bought one 20 times over, and still have change for a good party every week. or, just used hosted databases and liberal caches.
But that's not "at scale" that's just some great plains accounting app that's been dragged from one pickle jar to another.
In 2016 we had a 36k cluster. There was something like 2 PB of fast online storage, 48pb of nearline, and two massive tape libraries for backup/interchange.
The cluster was ephemeral, and could be reprovisioned automatically by netboot. However the DNS/DHCP + auth servers were on the critical pathway. So we dumped them on a VMware cluster to make sure that we could run them in as close to 100% as possible. Yes, they were replicated, but they were also running on separate HA clusters, with mirrored storage. This meant that if we lost both of them, we could within a few minutes run them directly from a snapshot, or if it was a catasrafuck reload the config from git.
Now we could have made our own DNS+dhcp server, and or kerberos/ldap/active directory. but that cost money and wasn't worth the time. Plus the risk of running your own with a small crew (less than 10 infra people) was way to high.
VMware was almost mainframe level of uptime, if you did it right.
You can emulate it though via UploadPartCopy in many cases, no?
If you are spooling through a file and changing a few chars at a time, you'll be generating a boat load of writes (and or reading from the wrong file too. )
[1]: https://learn.microsoft.com/en-us/rest/api/storageservices/u...
Sadly, Azure's implementation of its blob-store, is kind of underwhelming — especially for any kind of infrastructure-level use-cases.
For example, while there is a change feed for blob events, akin to S3 lifecycle event notifications, it stops exactly where S3's API stops; so there is no event generated by an append to an appendable blob, nor a write to a page in a page blob or a block in a block blob. (And even if there were, they make no guarantees of the change feed being linearized — saying that changes to some resources might arrive out-of-order or not at all; and that if you want a linearized change feed, you need to read it out of a log-multiplexer, which puts a several-minute delay and multi-minute step-granularity on reads from it.)
As such, you can't use Append Blobs as the storage layer for a Kafka-alike; and nor can you use Page Blobs as the transport to enable an embedded LMDB-alike to be network-replicated. (Or rather, you can, but in both cases you won't receive timely notifications that new data has been added / that pages have been invalidated at the origin, so unless you're operating with zero caching, your cache will end up stale and your state from successive reads will end up incoherent.)
Multipart upload allows you to upload a single object as a set of parts. Each part is a contiguous portion of the object's data. You can upload these object parts independently and in any order. If transmission of any part fails, you can retransmit that part without affecting other parts. After all parts of your object are uploaded, Amazon S3 assembles these parts and creates the object. In general, when your object size reaches 100 MB, you should consider using multipart uploads instead of uploading the object in a single operation.for multiple writers, you have to use uuid and pay the search cost.
S3 launched in 2006.
EC2 didn't go GA until 2008.
Elastic Block Store (a proper block store) showed up in 2008 too, actually went GA before even EC2.
S3 was for the longest time pitched as storage for the internet, and remains so today.
I super doubt these limitations are there because of any EC2 requirements.
I think you are misreading me, but to be fair, I was being vague about timelines.
my main thrust is that S3 is dominant because EC2 was/is lacking. S3 is optimised for uptime and consistency, which means that its brilliant at mostly static file hosting. It will work for more dynamic state type stuff, its just not really designed for it (see https://xeiaso.net/blog/anything-message-queue/)
making your own object store that is fast, durable and available and has another feature is really really hard to do at scale. Its far easier to put up with s3 than make your own.
It also can't read (seek) into the file (or couldn't when I last looked at it).
It is not a file system, but is often mistaken for one.
https://stackoverflow.com/questions/26964626/specify-byte-ra...
i wish it had a CAS mechanism, though. (google's does, with the x-goog-if-generation-match header)
Is that an ETag-alike?
maybe because you're more likely to have the same generation for different paths, because it's not a hash? an ETag also isn't guaranteed to be a hash, though, so idk
Adding to the thread: EBS was kind of created to make up for EC2's ephemeral storage issue, but in 2024 I'm pretty sure AWS would design S3 + EBS + EC2 in a very different way.
For that purpose, I think S3's age is a killer feature. I won't be surprised when we see the HN post "google cloud storage has been sent to the graveyard"
GSuite is the most profitable part of GCP, and it's totally propping that division up. It'll be the last to go - it's a lot easier to recognize bad management.
Domains probably was expensive to maintain or understaffed for some reason, and some exec knew a guy a Squarespace.
https://cloud.google.com/iot-core
... shows this:
Google Cloud IoT Core has been retired.
Several years ago a startup owner I was friends with was doing IoT stuff. If they'd ended up choosing the Google option for their device comms they'd have been in a bad place from this. Pretty sure they went with the AWS IoT stuff though.At this point, I wonder if there's anyone left on Earth who Google haven't screwed over in some significant way?
For cloud backup I use Arq backing up daily to AWS (no affiliation with Arq other than being a happy customer). You get client side encryption and the daily backup directly mitigates your concern, if there’s an AWS account issue you will know immediately and can fix it. For my storage amount and use it only costs about $2 a month.
I have about 80 GB or so of data being backed up. The daily backups upload only new files plus files that are changed. The largest monthly AWS bill I ever got was $6. The next month it went back to the usual ~$2 range.
While a scanned copy of something like a birth certificate or deed might not often work in place of a physical copy they're nice to have on hand. Plus they're a part of a family history.
i use spinning rust in x3 single usb external enclosures[1].
i mirror s3 data, both as an extra backup and to prevent s3 egress billing for random access.
s3 egress is for disaster recovery only.
Understandable concern, but AWS being targeted at enterprise gives me confidence they don't do any funny stuff like that. On and of course, they bill you monthly, so ideally you'd have a tell if something goes wrong.
I also store my files locally on a USB drive in alternative locations, but I conceptually trust S3, but maybe actions speak louder.
Using file operations for mutexes makes sense in Unix because of the filesystem semantics there but it makes less sense in a distributed object store.
Even if there is some technical reason why the data needs copying, S3 could at least pretend that the file is in the new place until it’s actually there.
This is true, but S3 does support replication (including deletion markers), and even 2-way replication, between two regions. Definitely not the same thing as a dual-region bucket, but it can satisfy many use cases where dual-region bucket would be used otherwise.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/replic...
You pay for durability and availability in the form of disk overhead, CPU, and network. Each encoding scheme has some expected cost. If your overhead for one-region is $X per gigabyte, then generally speaking, the overhead for two-region is going to be less than $2X—each region is more durable + available because of the copy stored in the other region.
Also you can then create a single multi-region endpoint for those so I don't know why this person says it doesn't exist: https://docs.aws.amazon.com/AmazonS3/latest/userguide/MultiR...
You can totally build stuff like Megastore or Bigtable (or Spanner) on top of S3. You use a log-structured merge tree. That’s how these systems work in the first place. In the log-structured merge tree, you have a set of files containing your data but you don’t modify them. Instead, you write new files containing the changes (the log). Eventually you compact them by writing a complete copy and deleting the old versions.
This works just fine on S3, and there are even some key-value stores built on top of S3 that work this way. Colossus is cheaper for short-lived data.
Your database may use multiple underlying storage layers anyway.
At a high level, yes, you can implement systems like Megastore or Bigtable over S3. However, there are many details you must take into account. You cannot simply wave away the complexity and potential failure scenarios.
For starters, how are you going to create the newest SST?
If you keep it in memory or on disk, it must be replicated to prevent data loss if a machine fails. This approach could lead to losing the most recent changes. Additionally, you end up with a hybrid system that needs to read data from multiple sources, which adds complexity. If you essentially reimplement the system, why use S3 at all?
What if the data volume gets too low and you end up writing many small, expensive files?
Using something like Kinesis for batching might work, but the data won't be visible for N minutes.
Merging partial tables also requires maintaining an external index to track availability. Transactions would be helpful, but how do you handle failures?
And we haven't even mentioned managing garbage collection. It would require an external lock or reference count system.
Maybe I am thinking more broadly when imagine what it means to implement something like Spanner on top of S3.
We know that Spanner on top of S3 is not going to give you the same price/performance as building Spanner on top of Colossus while giving you the same semantics. You either relax the semantics a little bit, you pay out the nose for a lot of little files, or you find a durable place outside S3 to store the newest data.
> At a high level, yes, you can implement systems like Megastore or Bigtable over S3. However, there are many details you must take into account. You cannot simply wave away the complexity and potential failure scenarios.
Most of the complexity is the same whether you implement Megastore on top of S3 or on top of GFS. You can’t handwave it in either scenario.
> If you essentially reimplement the system, why use S3 at all?
It’s highly durable, highly available, and cheap (under certain usage scenarios).
GFS is not available for anyone to use, inside or outside Google. Its successor, Colossus, is not available outside Google. They’re just not available.
People are asking for create-if-not-exists specifically to be able to add objects to an ordered log, without needing a separate service for coordination. S3 cannot be used for this. GCS, for example, can.
The “append” feature isn’t necessary for functionality, it just improves the cost / performance.
The idea of running your own database on top of S3 is, well, it’s gonna be janky. It’s not ideal. You do end up seeing databases running on top of S3 (I’ve seen some), and sometimes it even makes sense.
https://www.databricks.com/wp-content/uploads/2020/08/p975-a...
It's the actual federal hate crime that are the egress costs I could do without.
Being able to reuse the s3 command-line and existing S3 libraries has made the migration painless, so I'm thankful they've created a defacto standard that other S3 compatible object storage providers can implement.
Yes yes yes. However, DynamoDB can be expensive very quickly :]
Cloudflare supports CAS on copyObject and on putObject [1]. It doesn't support CAS on deleteObject.
I don't know about the others though (ABS, GCS, Tigris, MinIO).
https://developers.cloudflare.com/r2/api/s3/extensions/#cond...
[0] https://docs.aws.amazon.com/AmazonS3/latest/API/API_CopyObje...
[1] https://developers.cloudflare.com/r2/api/s3/extensions/#cond...
EDIT: This was based on my experience from ~2 years ago. The Blob storage team has reached out to me via email and let me know that both the key partitioning and throttling issues have been fixed since then.
From what I can tell, the S3 request rate is about 9,000 requests per second [1] split between reads and writes for a single partition. From my perspective it really just depends on what you're trying to build but I don't see the performance of Azure Storage as being an issue in any way for a typical application.
Partitioning will also depend heavily on what kind of application you're building, but the documentation does point out that load balancing will kick in once it starts to see a lot of traffic on a partition [2]. Since you have to use partitioning for S3 in order to get better performance, I don't really see how that's a point against Azure.
As for SDKs [3] I have no idea how good support is, but they all have commits within the last day.
[0] https://learn.microsoft.com/en-us/azure/storage/common/scala...
[1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/optimi...
[2] https://learn.microsoft.com/en-us/azure/storage/blobs/storag...
S3 does a bad job of describing what a prefix is, but for most[1] practical intents you can consider the entire key of an object the prefix.
[1] The actual behaviour treats prefixes like partitions, but it’s completely automated and as long as you don’t expect an instant scale up to very large request rates, S3’s performance is basically unlimited. There are no per account hard or soft limits that need increasing or limit scalability.
the ideal design is to use s3 for immutable objects, then build a similar system on ec2 nvme for mutable/ephemeral data.
the one i use by default is s4[1].
- Write to a unique temporary path (call this temp_path)
- Commit a record to DynamoDB with put-if-absent semantics on the key (dest_path) and the value (temp_path, "incomplete")
- Three possible outcomes:
1. Your write succeeded; proceed to issue a CopyObject call from temp_path to dest_path and if successful mark as complete in DynamoDB.
2. The row already existed in DynamoDB as "complete" -> do nothing
3. The row already existed in DynamoDB as "incomplete" -> another write has been committed, but is not complete; attempt to repair it by issuing a CopyObject for the path in DynamoDB. On success, mark as complete in DynamoDB. ("repair" step)
Reads could also hit DynamoDB before S3 in order to perform the "repair" step if applicable.
(But this rarely happens with S3, because it's so simple.)
I just implemented a "posix-like" filesystem on top of it. Which means large object offload to s3 is not a problem or even an "ugly abstraction." In fact it looks quite natural once you get down to the layer that does this.
You also get something like regular file locks which extend to s3 objects, and if you're using cognito, you can simplify your permissions management and run it all through the file system as well.
Actual storage cost is "meh" but...
My god the absolute highway robbery of the bandwidth costs.
https://www.cnbc.com/2021/09/05/how-amazon-web-services-make...
R2 still has per-request fees which, given PUTs are still $4.50/million and GETs $0.36/million (only about 10% off S3’s), they still have significant margin there to cover egress for objects of reasonable size.
One other benefit of paying up front is that I think you are less restricted on S3/Cloudfront in what you can host. I think Cloudflare has some restrictions (video files?) to keep costs more sane. Otherwise, you could spin up a YouTube clone extremely cheap.
And just to reiterate, AWS pricing can be absurd.
Is there a data center with many solid state disks?
Are tape drives used?
https://medium.com/@ivaramme/systems-design-notes-aws-s3-6ef...