How capacity planners credibly estimate application performance
queue.acm.org
queue.acm.org
The first step is not to use technical measures like throughput or response time, but to use business units (like ATM account balance checks or withdrawals), so you can forecast capacity according to business plans.
In modern web apps, that's served by issuing a transaction id to user actions, and tracing that through all the systems and sub-systems required to service that transaction. That's your data plane.
Unfortunately, the next step of applying queuing theory to your software model is almost moot in a multi-core, multi-machine world with GPU's on a modern memory bus, because the model cannot be both abstract and accurate.
But the real problem is cultural: IT is willing to set targets for measures it can own and manipulate, but it's much, much less willing to commit to business unit measures that senior management can see clearly. Indeed, one of the best ways to understand how IT fits into the organization is to ask what's being measured.
We actually do capacity planning in dollars to this day, and did do queuing models of both batch and multi-core, multi-machine transaction processing.
The simple fact that a large company is able to liquefy a non-linear large-blocks market into dollar amounts that you can measure, is the ultimate dream of any 1800 economist.
For example, businesses in HR might have metrics like average case correspondence count, open rate per day, idle status, etc. The dev team would commit to certain of these metrics, and optimize both costs and these meteics. This, to me, would signal an IT organization with high status (autonomy delegated to tech leaders).
Whereas a different company might have IT committed to X TPS without the ownership of business metrics
I sort of lump the "do a benchmark" advice in with "look under the lightpost, it's much brighter there" when you've lost your car-keys in a dark garage (;-))
They also tend to sort of make you want to optimize the way the code already works on some level (a function or a service), as opposed to making larger scale optimizations that often yield much bigger performance improvements. Often the really big real world improvements are architectural rather than the result of tweaking the access order to get better CPU cache access patterns. Like the latter happens too, but you run out of those types of optimizations relatively quickly.
Like if you have some small piece of code you want to optimize (for some reason), a benchmark is invaluable; but if you have a system you want to optimize, it's less useful.
Benchmarks are well-known, well-understood and popular. They aren't what capacity planners and performance engineers use, though.
By determining when and where the bottlenecks will occur, teams can plan for resource allocation ahead of time.
Disclaimer: I'm not involved with GCP in any way, besides being a satisfied customer.
(The exception being for a request that requires no asynchronous handling from the point the GET is received on the server to the point the first byte is written to the response, but that's rarely the kind of request that's ever a bottleneck in the first place since it's basically limited to just static html.)
In a sufficiently complex system you will have many different queues and other bottlenecks that affect throughput at scale. Some requests will also require more compute than others.
I like to plot the correlation between utilization and latency. Since utilization will typically vary during the day, you get lots of datapoints without running any benchmark. That lets you make more informed decisions regarding target utilization and when it becomes necessary to scale up.
You also have to adjust your thresholds based on your DR/BCP plan, your fault tolerance design and operational requirements. If you were hot/cold between 2 sites or AZs with a 95% SLA but you could run things hotter than if you were hot/hot between 2 sites with a 99.9% SLA. It wasn't unusual for our CPU usage threshold to be 30%.
Ultimately the critical piece of knowledge is when your utilization vs latency curve turns into a hockey stick and knowing what the bottlenecks are that drive that. For us we had to learn the hard way exactly what kind of throughput we could expect out of each layer of our infrastructure because in many cases, the latency would spike but the utilization didn't correlate with that under stress conditions but it did under normal conditions, learning how to pick out those early warning signs was an art.
We did some analysis to determine the mix of transactions for different types of business activity (BAU, annual big event that changed a profile permanently, periodic big event that represented an impulse/one time change) and could project expected CPU utilization based on a model for each mix of transactions based on business metrics. We based all of our capacity recommendations on the business volume projections by event type so that we weren't asking our BAs to tell us things like transactions per second by API which in most cases would have just gotten us blank stares, we'd ask them to give us a sales projection and apply our model to it.
Where businesses often seem to fall down is that, if they round-trip lessons from capacity planning to the business strategy, they sure don't look like it. They don't talk about it, and they make decisions that don't align with the capacity planning numbers.
If you know the cost of 1000 new customers is $M of new hardware, and revenue from the new customers is $N, then if you spend more than $(N-M) on acquiring the customers you're digging a hole, and the more successful you are the faster you dig. And then there's the grey area where you actually make money, but so little that any other initiative would have been a better use of your time.
https://en.wikipedia.org/wiki/You_Don%27t_Know_Jack_(franchi...
[0]: https://mitsloan.mit.edu/teaching-resources-library/mit-sloa...