Taking too much slack out of the rubber band
rachelbythebay.com
rachelbythebay.com
One example is a recent project was a very cost-sensitive machine in which a small heater was copied over from another product, but no one actually verified that it was good to the required limits (just the default use case). Well, turns out it wasn't quite powerful enough but it is way too late and expensive now to fix it at the end! Also, all the engineering time was wasted to figure this out (but it often seems management doesn't count engineering time the same way as parts cost)!
I've since learned that in the beginning of a project it is critical to identify the riskiest parts of the design and try to isolate that to a module and over-spec it, hopefully with a path to reduced cost later on. But the most important thing I've learned is don't try to solve tomorrow's problems today!
because the IRS doesn’t either
And while they're looking at that heating coil, they're not doing something that could be generating more value for the company. But opportunity cost is extremely hard to measure.
increase in COGS looks very bad for the business.
https://www.investopedia.com/ask/answers/101314/what-are-dif...
Some service has three units of capacity available (e.g. VMs). This is the minimum amount allowed, on the theory that things won't break too badly if one of them happens to crash. You target 66% CPU utilization. Suddenly, one goes down, and the software sees 100% CPU utilization on the other two. What should the software do?
Well, the obvious thing is to add one more instance, assuming that one of them crashed and its load shifted to the other two. However, what if the thing that actually happened is that the demand doubled, and the load caused the crash? Then, you should probably add six more instances (assuming that the two remaining live ones are going to go down while those six are coming up).
If you look at only CPU utilization, it's impossible to tell the difference between these two situations.
The reality is, we have a much more interconnected web of dependencies with little capacity to absorb disruptions. We'll almost certainly see some much more significant consequences when those now low probability events finally do occur.
Say I'm an e-commerce site and on Black Friday I can see historically (or just make an educated guess if it's your first holiday sale) I get "n" requests per second to my service.
I'll set my autoscaling group the day before to be able to handle that "n" number of requests, with the ability to grow if my expectations are exceeded. If my expectations are not met, then my autoscaling group won't shrink. Then the day after the holiday sale, I can configure my autoscaling group to have a different minimum.
This solves the problem of balancing between capacity planning and saving money by not having idle resources running.
If you're the type of person who hates human intervention for running your operation, then fine. Put in a scheduled config change every year before a sale to change your autoscaling group size.
It's pretty rare to have enormous spikes in application usage without good reason. Such as video-game releases, holiday sales, startup openings, viral social media campaigns.
Do you really think people do things because it makes sense to do them for their particular situation or because those things are "the thing to do(tm)"?
Most people go to see Mona Lisa because that's what people do when in Paris, not because they care about that particular piece of art.
Same with automation. It really makes me sad when I see people "automating" things they barely understand how to manually do, let alone the "when" to do it.
Yes, your example is perfectly valid, but that means one understands the system they are working with and generally people have no bloody clue about what they are doing.
[1] https://www.usenix.org/conference/srecon19emea/presentation/...
Maybe off-topic, but what are some good strategies for the kind of "self-healing" being talked about here? If a service needs to be restarted, how could you automate the detection and restart process?
Supervisors like systemd also have a watchdog that will force-restart a service that hasn't checked in for some time.
For a service that manages its own network connection, implementing auto-reconnect can be a form of self-healing (and surprisingly hard to get right in all edge cases).
The key is, as Rachel wrote in the OP, to get a good signal. You need to be able to distinguish a working from a non-working service to implement reliable self-healing.
I think this is the crux of what I was trying to get at. Curious to read how others have approached this problem.
What she's saying is that if you configure scaling such that it'lll scale down when demand is unusually low, and then demand returns, the spike may be a difficult one to handle, particularly if your services depend on each other but each scales only based on its own history.
If A needs B needs C, and demand suddenly returns to A, does that cause C to scale up? Or will A scale up first, and but C stay low for another half-hour because it recently scaled down?
Having C stay under demand for a half-hour after an outage ends wasn't anyone's intention when the autoscaling was configured. But as I wrote, don't confuse intention with effect.
In a fully optimised setup each service image is itself 100% preconfigured, and only provisions node secrets during the boot. Even one of these types of nodes takes easily 30-40 seconds from launch event to actually serve traffic: it may join the load balancer just 25 seconds in, but the load balancer will want to see at least two good health checks before allowing any traffic to it.
The problem with aggressive upscaling in the depicted scenario is that your plumbing layer is also likely scaled down. Hitting it with a cascade of new nodes has the risk of going all thundering herd, crippling the system for both existing and new nodes.