How Amazon uses chaos engineering to handle 80k requests per second
community.aws
community.aws
an advert in disguise then?
for that matter, any blog post by a company is an ad. maybe not to sell product, but to at least build exposure and familiarity with the brand.
As a stand-alone article it is fine, and is likely to trip the more fluff than stuff alarm on many people's bs-detectors.
"Stress testing" might be more intuitive but that's already established as simply testing under high traffic
It's just in this case the system is the company itself instead of just a suite of software services
If it's the real-deal, and not like people saying "Bun.js can serve 65k req/s+ (cough cough to localhost,)" that's impressive.
But I never see anyone talk about real-world numbers. Just synthetic poopoo.
I think I read the article correctly, but I think it only talks about how they introduce "chaos engineering," I didn't recall them talking about how they actually handle a volume of traffic like 80k req/s.
The number probably changes all the time based on load. They'll never release these details because it's a competitive advantage to have the "how popular are they in $place at $time-of-day" data private.
When they do share numbers, it'll always be the most flattering and devoid of any context beyond the "wow" factor.
You could take low-end instance specifications and standard industry stacks and extrapolate forward how many instances they might need to maintain at a maximum, but those numbers are going to be off.
Are they running 400 low-end front-end instances across the globe? Probably, (plus the 40 or so other services they claim to need, multiplied by region count at a minimum) and that would actually be well below realistic and reasonable for a company like Amazon. You can take a bunch of regional instances that handle roughly 200 req/s and make that work.
There's also a common misnomer that Amazon.com is somehow just this one giant app running on a set of servers, which isn't remotely how it's actually deployed, and that's before we spend time arguing whether a team's instances even count as "primary e-commerce front-end stack" or not. :P
Not to mention all the data center logistics
Source: Worked at Amazon/AWS for almost a decade.
Everything was Dockerized and I think we were using Docker Swarm for container orchestration. I don't remember the specs for each box, but we had auto scaling set up so at peak we'd hit a little over 200 containers.
Looking back now, I'm sure we could have gotten much better performance out of that service, but the team was young and inexperienced and throwing money at the problem was an easier solution.
Not all req/s are made the same.
Amazon search is made of 100s of services, and Amazon's search page loads 20 products per page, that means 80k search req/s translates to 1.6 MM product API req/s for example.
FWIW a search request at Amazon hits roughly 100 unique search clusters (think of this as ElasticSearch clusters - but its not ES), with different product groupings in each cluster. Each cluster is made up of 1000s of nodes running Lucene (think similar to ES shards). This is just for the "match set", i.e. which products to return.
Then there are services to re-sort those matched products based on popularity, likelihood of purchase, etc. Think giant ML models. Then there are product lookups. Before all of this, there is Query analysis to simplify/improve the query (think giant ML models) to classify "Apple" into electronics vs groceries based on the other keywords and your current context.
Meanwhile, Bun.js is talking about 65k "hello world" type req/s. The compute per req is magnitudes different.
That's their problem, no? Nobody's forcing them to have an architecture where a request propagates to hundreds of services.
Well, Jeff Bezos circa 2002 or so did:
https://news.ycombinator.com/item?id=3102800
It's hard to remove that ethos now
Scale issues are emergent issues. It's a phase transition.
I don't get the hostility around this article. Nobody is forcing you to read it or to do it this way. If your system is architected in a different way where you can run your whole system on a single instance, then good for you! But Amazon presumably doesn't have that luxury, and others may not either.
I mean, physics kinda is.
Just the volume of data that needs to be hosted, needs multiple nodes. ElasticSearch has some good general documentation of search engines if you want to learn more. In Amazon's case, a general query will hit a fanout of about 10,000x nodes - I wasn't even counting that fanout because it is all technically 1 service.
Meanwhile, Bun.js (no offense to Bun) is built for IO-bounded workloads and 65k req/s is great for that. However, executing the required Natural Language Processing (i.e. multiple ML models) would result in an exhausted CPU (even if Bun could distribute the compute across cores or compute the ML inference on a GPU). I'd be willing to bet even on a great CPU it gets throttled at 10 req/s (at most 100 req/s).
This really isn't some "They should have just done it all in Postgress, Bun.js can just use a nice ORM and be IO-bounded" type situation - which is a philosophy I very much agree with for 99.9999% of usecases.
So, instead of that, what would you have? A single, binary that makes multiple DB calls? Hmmm, let's see, some people have tried that and written extensively about the problems they faced doing that. Wait, one of them is actually a small e-commerce firm named after a larger river. Wonder what the issues were ...
Observability was our secret sauce. We would monitor everything. Our caches, NICs, our load balancers, etc. Cache hotspotting and DB problems were the problems that kept us up at night, though my teams didn't deal with much stateful data.
Like, you could hit a cache and serve 95% of page from it. Or hit some long path that will burn half a second on 20 servers in the backend to serve some big query
But yeah title is useless
Back in 2014 they were still sharing a few years old (at that time) detail-level service call graph that had so many nodes and lines it looked like string art.
https://aws.amazon.com/blogs/industries/building-blocks-for-...
80k RPS is not just for a single database lookup, don’t be silly.
req/s from localhost to localhost, and req/s from the Internet to any user.
The latter is actually interesting. People saying you can get 10k req/s from Node.js is stupid. You're not actually getting that on say, a single low-end instance over the Internet, which is what most developers are actually going to do.
Instead, you'll get two orders of magnitude fewer requests per second.
What Amazon is talking about here is most likely non-synthetic, real-world 80k requests per second. Which is actually a decent job.
No, it's not, for exactly the reason you state:
> You're not actually getting that on say, a single low-end instance over the Internet
Some languages are, of course, more efficient, but it doesn't matter - you can get very good performance out of any language/runtime - it's all about your architecture and infrastructure.
From what I've measured, code I've written performs around the same in production as it ran locally given similar hardware. If you're deploying to a VM with 3000 IOPS and 1/2 a CPU core, obviously it's going to run like garbage. If you wouldn't run your business on a raspberry pi 3, you probably shouldn't be running it on an AWS xlarge instance either.
I imagine search is more complex and expensive than CRUD, but 80k isn't something you can only do with a "hello world" tier application.
There is surprisingly little seasonal variance. You have your weekly/daily traffic rhythms based on when your users are awake/active based on their geographical distribution and that's mostly it.
"World events" also have very little impact - they tend to barely make a dent in that massive background noise.
Before we had large volumes of traffic I thought we'd be seeing all sorts of unusual peaks, after a few years I realized growth at scale tends to become boring (but in a good way).
About a decade ago Opera Mini did 150k transcoded full pageloads/s (times about 30 inlines per pageload that was the average back then, so about 4.5 million requested/loaded/processed/compressed HTTP resources/s).
(All of the public Google Search numbers I've seen have seemed one or two orders of magnitudes too small. Or maybe most people don't use their search engine/browser as much as I do, so my perspective is skewed...)
https://cloud.google.com/blog/products/databases/youtube-run...
For example, it is just as true for this title to have said "How Amazon uses ... to load 1.6 MM requests per second, from just the search page."
Each search page load, is 1 request to the search backend, but 20x request fanout to the product's key-value store to render the images and titles, etc.
I'm at the point that I rarely ever buy products on Amazon anymore. It's a total disgrace. On an ethical level, I wish I had the ability to say "I only want to be presented with results that weren't made in China or other slave societies".
Contrary to popular belief Amazon actually does put energy into making sure products are responsibly sourced. Products are de-listed if they’re found to come from unethical sources.
To take that even further take a look at Climate Pledge Friendly. Those are products with (at least one) third party certification. These certifications don’t just further climate goals. Social responsibility is also considered. Including worker conditions and product durability. You can filter search results by this attribute. Admittedly it can be hard to filter for specific certifications.
Funny story, internally Amazon Search doesn't consider the ads products to be part of the "search results". It is tracked and accounted to Ads.
The way Ads are handled on Amazon is really poorly done. The Ads teams claim to make a lot of money (and based on the internal accounting tricks they do), and as such have been pushing Amazon's leadership to go more into Ads, even tho every person I've met that worked at Amazon also hated the prevalence of Ads.
Literally, Directors and VPs at Amazon are afraid to step on the toes of Ads' leadership team because of how well they have told the story about "Ads is excessively profitable".
Meanwhile, all of us in the thread can easily say, even if it is short term profitable, it most certainly is not long term profitable for Amazon.
Both from internally and externally it has been very disappointing to watch actually.
Which pales in comparison to the problem of counterfeit goods, IMO.
I can at least somewhat comb through the reviews to look for outliers of well written reviews. Getting something that's obviously a fake (has happened to me multiple times) is completely unacceptable.
Newegg has this issue too, I got a knock-off Intel CPU there once, I was furious.
If you have to use 'Chaos Engineering' to experiment your way into innovation, this is a sign you built your service wrong. What will Amazon Re-Invent next!? I am guess the wheel. Well written article though.
However, it won't necessarily help you know how your system will behave if S3 kicks the bucket in us-east-1 (again), your image host for that super-cool Kubernetes cluster suddenly throttles you during a critical restart, or your other service of choice went down due to an expired certificate.
If you however mean to use it to perform a denial of service on an endpoint you don't own, you're more hard-core than I thought.
How does a bash script simulate network failures, connectivity drops or a gray-failure in one of your dependencies? [2]
A lot of thought has been put into this domain. Dismissing that without understanding any of the complexities is just showing your ignorance.
[1] - https://brooker.co.za/blog/2023/05/10/open-closed.html
[2] - https://docs.aws.amazon.com/fis/latest/userguide/fis-actions...
With you so far
> A bash script utilizing curl will suffice.
Lol hell no. Yes, AWS/amazon does require a "GameDay" before launching a service which will execute an mcm (managed change management) that's basically a runbook of (way more in depth and comprehensive way to test your service than a single bash script with curl), but chaos engineering is a great additive, additional verification mechanism that really helps with service outages.
How are you going to test a thundering herd with a bash script executing curl?
How many machines are you running this simple bash+curl script on anyway? Using a single node to generate requests isn't going to do much in testing a service's reliability.
Also, they literally did invent a wheel