AWS Identity service handles 400M API calls every second
aws.amazon.com
aws.amazon.com
Disclaimer: I'm a founder of authzed (YC W21), a productized Zanzibar implementation
The computational complexity of the policies wasn't high (in fact, they were pretty much key-value things, as the only time we had something more complex was in AssumeRoleUsingWebIdentity calls - btw, that API sucks, embarassingly so - everything else was plain list of what a role could access).
Otoh, it bothers me that every single service call needs to go to IAM to check for permissions. Has anyone explored other architectures/designs to circumvent centralized auth?
"400m operations per second? Wow, that's almost 5% of the number of operations per second a typical consumer processor can do!"
I realize this is an apples to Buicks comparison, I just get bothered by how millions of anything might be impressive when we have PCs that are designed with billions of everything.
That is a strange thing to be bothered by. Are you similarly unimpressed by a $400M lottery jackpot, since CPUs can execute billions of instructions per second?
As people have pointed out, something like Dynamo is likely FAR more complex - but my understanding is that it's also 3-4 orders of magnitude smaller.
Boundary is mainly self contained but there can be interactions outside, sometimes significantly.
Having said that, under the hood who knows .
Thought experiment: if it suddenly took 10 seconds to respond to the calls, the number of responses still equals the number of requests, so the request velocity mentioned in the headline remains the same.
Used to work on 20M+ user identity service that didn’t see volume anything like this, but when utilization went up usually it was just a matter of adding more replica databases and auth servers.
The load isn't the only issue either you also have availability SLAs you need to meet that force you to design the system to tolerate faults at that scale or else half the internet gets taken down. Your system has to be broken up at a granularity where many engineers can work in parallel to diagnose and mitigate an issue when your service does get taken down to minimize downtime. “Scale breaks everything” as they say.
If you're using caches, you're deliberately sending old data to the replicas. I can't think of why you'd do that. I guess you could have replicas of your caches, but I'm referring specifically to data store replicas.
A cache stores the output of an operation to avoid needing to perform the same operation again. You deliberately put data into a cache to avoid needing to create that data a second time—it won't magically appear there.
But perhaps the permission updates are slow. Like, you change a permission, wait 5 seconds, and the system tells you that they are updated across all servers.
The difference between mostly static and completely static is where the interesting bit is. The two categories beget completely different solutions, especially when serving stale data is not acceptable, as in the case we're looking at.
I suppose this comment falls in the same category as "never underestimate the bandwidth of a truck full of tapes".
Yes, the numbers are huge. Whether it is a technical achievement is another question entirely.
This is a globally distributed service operating at massive throughput and ultra low latency scaling requirements where if it fails millions of customers and a majority of online commerce as a whole comes to a halt. What you’re referring to is just the very tip of the iceberg.
The hard problems stem from how the system deals with failures and how the system propagates writes across the replicas while meeting latency and consistency SLAs. On top of that the system needs to be built in a way that it can be maintained by many developers each working on a small piece of the system without knowing the ins and outs of the system as a whole. In addition, when the system fails debugging and mitigation needs to be able to be parallelized across many developers so that availability SLAs can be maintained. You can read about this in “Designing Data-Intensive Applications” by Martin Kleppman where he discusses the complexity involved in building distributed systems.
That's the point: nobody said permission updates have to happen within a second. They are performed infrequently, so the system can afford to make them slow.
https://aws.amazon.com/iam/faqs/
P90 values provided to us in 2017 (large enterprise considering AWS) was under 1 second.
Edit: I understand how every machine needs to invoke IAM APIs and how temporary credentials and other uses increase the number super-linearly with every active user. Still, 400M RPS (nearly 35B requests/day) could be reduced significantly by improving the underlying object model so it scales down better. Right now, even a simple Lambda function that needs to access other AWS resources requires 3 API calls: create a policy, create a role, and connect the two.
The thought process that would lead to such a statement intrigues me.
IAM activity scales with total aggregate activity across all workloads that run on AWS, which I would assume scales with the total activity of the billions of people who interact with AWS directly or indirectly each day.
3.456 × 10^13
34 560 000 000 000
~35 long-scale, european billions (10^12)
~35000 short-scale, american/english billions (10^9)