Port-scanning the fleet and trying to put out fires
rachelbythebay.com
rachelbythebay.com
This is very much tickling me as one of the truths about operational work. I can't count how often I have said something like "Yes, technically we don't need all of that.... but..." Systems tend to become weird in production if many people use them.
Like we have a central component at work and most things have good client side load balancing and retries in place. Thus we decided against adding a server-side load balancer - more software and systems on the path are slower and more pieces that could fail.
Except, as we found out during an outage, ~2-3 critical requests into this central component in a third party component... aren't and just fail hard if they hit a bad node. And then everything escalates into a mess very quickly.
So now we have both kinds of load balancing. Isn't that wonderful?
Minor note (of the literally/figuratively type): shouldn’t that be ‘yes, theoretically we don’t need all of that, but technically [or “practically”] we do’?
I think that ‘technically’ connotes actual correctness.
I think this is the right way around, because I lean towards the "technically correct" side here. "Technically correct" always carries a bit of an ivory tower snobby attitude to me.
After all, the tech I'm working with has proven and working client side load balancing and retries, so there is no server side load balancing necessary. That's technically correct if you observe the system in isolation. And it's also correct that the bad software interfacing with it should fix their implementation.
It's just not correct in practice if terrible software, implementation bugs, or all manner of things beyond a simple, correct implementation start interfacing with the system. Then even a technically or clinically or theoretically correct system will fall apart because there's mud in places where mud shouldn't be able to get to. After all, a patch in a month won't help me with a preventable outage tomorrow.
We refer to "unexpected critical things" as "load bearing flasks" - you never know they're load bearing until you get ye flask.
> since they were so low in the stack, when they broke, lots of other stuff broke
It is interesting that while engineers are far too happy to complain about dependency hell in terms of libraries and languages they don't complain enough about dependency hell when it comes to systems. Engineers should take a more proactive role in identifying and pushing whatever changes are required to prevent mass outages due to single points of failure. In other words - if more than a significant percentage of your business requires that X system be up and running, then your business has made a grave error. Engineers should be the first to point this out.
"That's the architect's job"
My thoughts on the usefulness of architecture aside, it's also the engineer's job. Not only should an engineer identify and be aware of system dependencies they should also build software that allows, prevents, and gracefully handles system dependency breakage. This is why the engineer should understand the business as much as the technical. (this is especially true for leads)
"But I inherited this system"
Yeah, it sucks. But now you've got work to do to mitigate risk. Risk mitigation outside of the context of cybersecurity is often overlooked. This is not an advocation for deploying to every data center/cloud provider in the world in the name of high availability. You'll need to do the math to make sure it works for those critical pathways.
EDIT: clarity
In my experience even if you've eliminated single points of failure you can still get failures.
Sure, my server's got dual NICs and dual power supplies. But if the guy sent to replace server 12 in rack 345 accidentally gets server 12 in rack 346 I'm going to lose both NICs and both power supplies at once.
Sure, the network is fully redundant. But we naturally need to keep the settings on the two sets of kit in sync. We've had outages due to them getting out of sync in the past, not going to make that mistake again. So of course if there's a change to the firewall rules it automatically gets rolled out to both of the firewalls.
And so on.
Eventually we built this into the application we were deploying, so that it could monitor its ZooKeepers, raise alarms, and give us insights into aggregated reliability data.
It seems like the most useless piece of software. Like someone wrote a concept down but said it was software.
Apache Curator builds on top of ZooKeeper and turns the primitives into usable recipes, which is what I recommend to most people these days.
One concrete example: Consul can provide you with locking and thus leader election with relatively little effort. Patroni, the postgres manager, can use a consul lock on a key to elect the postgres leader and consul ensures mutual exclusion.
Or in a similar way, I needed to run some data collection once a day, but the cluster consisted of 3 identical nodes and I didn't want to start treating one node specially.
I could use consul to hold a lock by using <consul lock whatever/lock collect-data.sh>. Once I had mutual exclusion, I could store the last time data was collected in a consul kv entry with <consul kv put whatever/last-run $( date +"%s")>. Then each script just grabbed the last time data was collected or 0 if missing and checked if it was 23ish hours ago. And there we go, all systems identical and fault tolerant once-a-day data collection without additional config to forget on top.
Nobody chose to have Zookeeper in their stack, but they needed it for these other services and so eventually they were stuck with it.
The problem I see often is that it's wedged into places for which it's absolutely overkill. As in clusters with less than a dozen nodes, no service level determining such a need, etc. - it's often just a shiny object for architects.
as a note, both ZK and etcd have key watches such that one need not poll the key-value store, it'll push out notifications for changes to subscribers, making for super quick reactions in any such distributed applications
Kinda weird how often that sort of conclusion pops up, in otherwise completely unrelated articles...
/real-mode-on
i am slowly (few decades) giving up and having to accept that "software IS about human in the loop".. at least culturally. It wasn't back in the beginning maybe.. but now it more or less have become an institution, The_Software, with usual Shirky's principle applying etc..
that shuttle blew up because of (mis)culture.. and that wasn't (invisible) software there. i guess we are up to interesting times ahead..
Yes, this describes a management problem. Sounds like shoot the messenger would be an easy de-facto choice for middle management, given how many years and spent dollars creating that situation must have taken...
For rapid use next time you can run lsof with various arguments to obtain a list of listening services on UDP and TCP ports. Cluster health and service status check wise, I have had good experience with https://clusterlabs.org/pacemaker/ + https://github.com/corosync/corosync
though it seems like a good step to run in a CI system, to try to preclude invalid configurations from ever showing up
I’m sure Rachel had her reasons but jikes
That service (Chubby) is a dependency of Borg in the same way that etcd is a dependency of Kubernetes. Except the development of etcd got to skip making all the same mistakes that were made when Chubby was developed a decade earlier.
Oh hey, I didn't see you there. Welcome to earth.