Amazon S3’s 15th Birthday: 5,475 Days and 100T Objects
aws.amazon.com
aws.amazon.com
[0]: https://vimeo.com/7330740
[1]: http://itc.conversationsnetwork.org/shows/detail5273.html
https://cacm.acm.org/magazines/2021/3/250706-a-second-conver...
https://news.ycombinator.com/item?id=26365873
[1] which you commented on, but others might not have seen.
I have a tough time taking him seriously reading this - I'm not sure if he's being intentionally ignorant or intentionally misleading.
I don't know of a single enterprise customer that doesn't have their data replicated to two datacenters, and then backed up to some other medium (tape or disk-based backup appliance) for anything business critical.
Their goal is 5-9's of UPTIME - the expectation is 100% data durability. You would get fired if you architected a solution keeping two copies of data in one datacenter.
Object storage absolutely has a place and I'm happy amazon was able to push a quasi-standard for the industry. Object storage prior to S3 was a hodge-podge of proprietary plays (EMC Centera, Bytecast, etc) and sort-of standards that nobody really used (CDMI). I just wish they did a better job of fairly representing it vs. turning every opportunity into a sales pitch spreading FUD about the alternatives.
He said:
> Most of our customers, if they have on-premises systems—if they're lucky—can store two objects in the same data center
You said:
> I don't know of a single enterprise customer that doesn't have their data replicated to two datacenters
His customers don't equal your enterprise customers. You are both right most likely.
Longtime consultant here with BCP/DR insight into 20+ large F500 companies... I think you would be seriously surprised at how common it is for even enterprise customers to not bother with multi-DC or even offsite backups. And even among the ones that do, many of the ones that are non-cloud-based are so immature at it that I would not put money on their backups being restorable if needed. And this is even more true for your “we’re a startup, we don’t have time to worry about backing up our data!” companies, which are in abundance.
For a small peak, go look at how many people were freaking out about losing their entire business due to the loss of a single OVH data center.
Of course, as a consultant I do naturally skew towards customers that need help with this stuff, so my perspective is probably biased towards the companies that are worse off in this regard. But they’re definitely out there.
You snipped out the part of the quote that answers your question.
"Most of our customers, if they have on-premises systems—if they're lucky—can store two objects in the same data center, which gives them four 9s. If they're really good, they may have two data centers and actually know how to replicate over two data centers, and that gives them five 9s. But eleven 9's, in terms of durability, is just unparalleled. And it trumps everything."
> I don't know of a single enterprise customer that doesn't have their data replicated to two datacenters...
Except, in AWS' case, each AZs (Availability Zones) is made up of upto 8 DCs (Data Centers), and each full-region has at least 3 AZs and 2 Transit Centers. Amazon S3 replicates data to 3 different AZs (which, I am guessing, is in addition to replicating it across DCs in a single AZ for 'eleven 9s').
With S3 cross-region replication durability may shoot up to 'sixteen 9s'? 100% durability, if it exists, is something the major Cloud providers are yet to offer?
ref: https://maisonbisson.com/post/object-storage-prior-art-and-l...
Replicating across two datacenters isn't "really good" though, that's considered table stakes.
>Except, in AWS' case, each AZs (Availability Zones) is made up of upto 8 DCs (Data Centers), and each full-region has at least 3 AZs and 2 Transit Centers. Amazon S3 replicates data to 3 different AZs (which, I am guessing, is in addition to replicating it across DCs in a single AZ for 'eleven 9s').
In AWS' case, the 8 DCs in an AZ are directly adjacent which isn't really useful for anything beyond metro (active/active) availability. I haven't seen a DR plan that doesn't have a hard requirement for the secondary copy of data to be outside of the metro loop/blast radius generally in a different state.
That’s mind boggling to think about. I wonder how much paper would’ve been needed to store 100T objects on paper. It’s like the new Library of Alexandria.
Presumably the Library of Alexandria didn’t have access control lists
Would be very unfortunate if the S3 story ends like the Library of Alexandria.
Has anyone lost data on S3 or know anyone who has?
> To ensure that data is not corrupted traversing the network, use the Content-MD5 header. When you use this header, Amazon S3 checks the object against the provided MD5 value and, if they do not match, returns an error. Additionally, you can calculate the MD5 while putting an object to Amazon S3 and compare the returned ETag to the calculated MD5 value.
It’s not as obvious with multipart uploads where each chunk is (iirc) individually checksumd and stored, but only a single md5 is returned in the “etag”
https://www.quora.com/Has-Amazon-S3-ever-lost-data-permanent...
This is just a guess based on my own usage of it over the last 15 years at different companies.
https://www.allthingsdistributed.com/2010/05/amazon_s3_reduc...
aws s3 sync ...
And you are in business.But I think the real lock-in lies in services like S3, making so many workflows simple and intuitive (while controlling your data, which is where the real lock-in lies)
How does something like that get handled? At some point isn't contacting this service really just contacting a handful of IP addresses? Are those just "regular" computers? Specialized hardware?
Also with the amount of centralization here, presumably they are the target of bad actors. How do they defend against things like DDOS's to provide their claimed uptimes?
A service that can handle tens of millions of requests per second can probably outscale most DDoS attacks without further engineering.
"Handful". S3 has 283 subnets allocated to it, biggest one being a /15, so there are some few hunderd thousand IP addresses reserved for S3. While impossible to say how many of them are actively in use, I imagine large fraction of them.
The AWS IP ranges are documented in https://ip-ranges.amazonaws.com/ip-ranges.json
We will start having trouble with random UUID collisions when we have made 2^64 of them, which is a huge number. It's about as many iron atoms as are in an iron filing.
It is also about 200,000 times 100 trillion, so AWS is just 18 doubling times away from the UUID system breaking down.
UUIDs are usually not simply random numbers. So it's possible to tell when a particular standard of UUID is used. Amazon's code can branch based on this detection, or even something simpler like "UUID-type", a new field that would be initialized by default to the old UUID type.
I simply don't see this as a problem for anyone.
Also as another commenter mentioned, if they are tied to a namespace prefix, that delays collisions.
And this is why GCP will never come close to AWS.
But it would be because 1) they probably don't have many duped objects to begin with 2) the system to de-dupe items would be complex, error prone, increase latency, and just not worth it
That's crap reasoning. You have no way of knowing that.
> 2) the system to de-dupe items would be complex, error prone, increase latency, and just not worth
Would it really be that complex? I kinda doubt that.
Read latencies wouldn't be affected - except being improved.
Write latencies - the deduplication can happen after the write has been confirmed, so no extra latency added.
Whether or not it's worth it - I think it's very hard to to estimate the deduplication factor. My guess is 1.5x. At that point it would save 33% storage. That would instanly make a whole lot of extra complexity worth it.
I dunno, I kinda think you might be out of your depth on this one, based on that response.
Also: none of your few previous posts are tech-related. They're all about personal finance.
All things considered, I’m much more concerned about the centralization of eg browser technologies than this.
If you try to decentralize the data, or keep your data yourself, you end up with something that uses more resources, has worse durability, or is worse along some other axis.
I am not so worried about centralization yet. There is still healthy competition between different cloud storage vendors. My impression is that cheap storage is used to make other cloud services more attractive—which means that Amazon’s incentives here are somewhat aligned with mine (as a customer). Personally I am much more worried about the centralization of the shopping experience.
I don’t think filecoin will ever demolish a fantastic business like S3 based on stealing away customers. The interesting value proposition from the protocol is probably more about censorship resistance and adjacent issues.
I think it's clear as-is.
No, it doesn't. Not more than it reads terameters or teragrams. SI prefixes were not invented specifically for computer memory.
Further, '100 Terabyte Objects' wouldn't really make sense as a sentence anyways.