Amazon S3 – 2 Trillion Objects, 1.1 Million Requests/Second
aws.typepad.com
aws.typepad.com
That actually sounds underwhelming! IMHO our brains have an easier time thinking "hey, 1 every 60 hours that's not much" compared to figuring out the universe is really incredibly old ;-)
How about comparing to the lifetime of one person? With the US average life expectancy (78 years) that means you'd have to upload 800 objects/second for your whole life to get to 2 trillon :)
Edit: One S3 object for each fish in the ocean (3 to 4 trillion) could also be a nice future milestone (if also slightly underwhelming :)
Edit2: I also love the eye blink as a unit of time. Each time you blink (average is around 15 times a minute), XXX more objects will have been uploaded to/requested on S3
Now imagine putting 9 999 people behind every person.
And finally, imagine stacking 9 999 on each of these million people's head. That's twice the height a normal airplane flies at.
Then double that number. That's how many objects have been created so far.
---
A bit long, but I think the roughly human scale numbers, 999s and even more effect makes it almost imaginable.
EDIT: Another one - you could fill the whole of manhattan with cents and still have money to spare.
(Manhattan is 87.46 km^2 and a coin is less than about 40mm^2)
Not least people will have seen the many, many visualisations of that figure. :P
If you had $8 for every object you could wipe out US National Debt.
A typical grain of beach sand – which must be about the lightest solid thing that’s easy for most people to envision – is 3 mg. That many grains of sand would mass 6e6 kg (6000 metric tons), which in turn is near the upper limit of masses that make sense to most people in everyday terms.
So you could say that if every S3 object were a grain of sand, S3 would be 3× the launch mass of the Space Shuttle, or somewhat more than the gold held in Fort Knox, or as much as the heaviest living thing: http://en.wikipedia.org/wiki/Pando_(tree)
And the the 1.1 million requests per second would be about 3.3 kg, or as much as a gallon of milk/water.
In fact, even if you and all your friends spent your lives reading the lists of object names, you wouldn't have time to get to the end.
In fact, even if you and all your friends and all of their friends devoted your lives to reading lists of S3 object names, you wouldn't even make a dent because new objects are arriving faster than you could collectively read their names.
(1 card is 0.11cm x 0.15cm x 0.01cm, 32GB. 32k objects, 31.2M cards needed. 165 cm x 165 cm x 165cm = 110 * 150 * 1650 cards = 27.2M. Instead of 165cm which is 5 and a half feet, say 6 feet, so try 166 * 122 * 1829 = 37.04M, enough to use some error-correcting codes just in case.)
Anywhere from 20 to 200M per year depending on how many people use RRS...
2000000000000 - (.99999999999 * 2000000000000) = 20
Makes me feel pretty OK about having backups there.
Assuming each object is 100KB (generous estimate, after compression) that would be 270GB per day -- or assuming ten levels of redundancy and striped across three RAID storage devices (per level of redundancy) then 8.1TB per day.
I'm not familiar with their hard disk procurement policies but it wouldn't be difficult to assume they've been purchasing 1TB drives, so 10 new disk drives per day just for keeping ahead of growth. Furthermore let's assume their disk drive failure churn rate is 10% per day so another 1 new disk drive for parts replacement (so 11 disk drives per day).
These are really loose numbers not based on any actual data (or any personal experience at all) but just napkin math, so take it all with a grain of salt.
After applying the same shoddy math with each object being 100KB -- 270TB with 10 levels of redundancy across 3 RAID drives resulting in 8,100TB per day. This would be 8,100 drives (at 1TB per drive), or 8,910 drives after 10% being dead-on-arrival.
The math is sketchy, so let's cut it down by 10x (10KB per object): 891 drives per day. Keep in mind this is just for S3 and it doesn't account for existing drives failing, growth, or what other services require (eg: EC2, RDS, Cloudwatch, Cloudfront, etc).
9 months ago, they announced they stored twice this amount - 4 trillion objects. A year before that, 1 Trillion. Given that previous rate of growth, we can expect they have a lot more than this now.
They also announced peaks of 880,000 requests a second. Whilst Amazon wins here, I'd say its fair to assume this number has increased in those 9 months.
http://blogs.msdn.com/b/windowsazurestorage/archive/2012/07/...
https://github.com/kennethreitz/elephant
It's going quite well so far :)
Jk... but seriously, good question! it's gotta just be a massive feet of engineering
Dropbox does use S3, but I think you're probably doing the fairly standard human mistake of not realizing just how big a trillion is. Dropbox hit 100M users in November of 2012, so for Dropbox to use up two trillion objects each user would need to have 20,000 of them. Dropbox does deduplication, has lots of inactive/minimal users, etc., so they're probably a percent or two of S3 objects.
(At one point, the stuff I was storing/logging for Cydia actually represented over a percent, maybe it was even over two percent, of all objects in S3; now I'm between 0.1% and 0.5%.)
http://demonocracy.info/infographics/usa/us_debt/us_debt.htm...
The average server can serve more than 1,000 static objects / second easily.
I was assuming that the number of requests referred to read requests, and I guess that the system is designed in a way that makes read requests very cheap, maybe even cheaper than reading a file on a usual filesystem, at least for hot data.
Just because it's huge and complex doesn't mean it's slow and that requests are expensive.
1.1M RPS is the amount they actually serve, not how much they can serve. Just because your single server can serve more than 1,000 static objects/second (in fact, that number should be much higher), it doesn't mean you need to.
It depends what you call a static object.
If the content of your object is stored somewhere and you can just send it without transformations, it's some kind of static object.
Now, looking at a single S3 bucket as a key-value store, with some kind of routing mapping an object's URL to a set of shards each containing the object, one could argue it's serving static objects.
> require a lot more computation to work out where they are and where they need to go
I hope not. I would bet it's not very far from serving a file from the filesystem. There may be a lot of i/o contension though.
> in fact, that number should be much higher
Yes, very probably, and that only makes the 1.1M RPS number seem even less impressive - for amazon.
> 1.1M RPS is the amount they actually serve, not how much they can serve
My whole point was this number of requests seem low for amazon, not that they couldn't handle more.