An update on recent service disruptions
github.blog
github.blog
- The team was not able to sufficiently diagnose the bad query program, and/or devise emergency mitigations, by 1400 UTC March 17 (24h after the first incident).
- They were not able to reduce the load on the db more quickly (eg; cranking up rate limiting, shutting down async or noncritical services and features, etc). At $prevjob we had the ability to do things like this at the push of a button, and generally would have within 90min of incident onset to help systems heal. Here they did eventually figure out they could throttle webhooks but it took several days of repeated incidents.
- They did not see connections creeping up towards their limits and have emergency mitigations prepared in case of spikes like this.
- The simple failover described on March 22nd (for example) took almost 3h to complete.
- They decided to enable profiling during the approximate time window of their load spikes, without sufficient resources.
- They lept to failover their db when the proxy had issues, when it seems the failover can take multiple hours to successfully complete(??).
- They did not have sufficient sampled trace/log data to form a hypothesis without this profiling.
- New services like Packages and Codespaces rely on the `mysql1` primary db.
Certainly I have much empathy for the responding teams – I have scars from responding to some very painful degradations myself – but as a Github customer this does leave me concerned.
Most NASDAQ tickers are 4 upper case letters. Most NYSE tickers are 3 upper case letters. I've never seen a 7 letter lower-case ticker. Tickers aren't unique across exchanges, so you need a pair of <ticker, exchange> to avoid ambiguity, but then you have to deal with shares being fungible across primary and secondary exchanges. CUSIPs aren't memorable and only apply to North American securities. ISINs are even less memorable. CUSIPs and ISINs include Luhn check-digits, which help catch typos. RICs are pretty good, except that they're copyrighted by Reuters, so if you build your systems around them, you may be obligating yourself to buy their market data services. Conversions between different types of symbols are sometimes many-to-one. (For instance composite RICs like AAPL.OQ and their primary exchange RICs like AAPL.O would have the same CUSIP and ISIN.)
In most contexts, you can get away with treating tickers case-insensitively, except (last I checked) 1 pair of case-colliding tickers on the Toronto exchange and 3 colliding pairs of tickers on the Bangkok exchange.
If you're dealing with financial symbols, implement them as an interface with methods to perform conversions, fungibility checks, etc. Don't pass around financial symbols as strings if you can help it. At my previous job, we had our own internal hierarchical symbology that covered rates, credit, currencies, commodities, equities, and more. Unfortunately, some systems pass these around as underscore-delimited case-insensitive strings. Externally, some symbology conventions distinguish day count conventions case-sensitively with an m or M suffix, which then had to be translated to ^M and M, respectively (and confusingly, since ^ looks a bit like an up arrow but marks the lower-case version). They've luckily replaced most usages of the hierarchical symbols with objects that pretty-print to the older textual representation and have constructors that take the older textual representation.
(For more than a decade, I maintained/improved the functional reactive time series subset of a domain-specific financial programming language and integrated globally replicated distributed NoSQL DB for a large global financial firm. By default strings are treated case-insensitively. The case insensitivity helped lazy typists when the system was first used in New York and London, but is now illiquid technical debt. The ability to put spaces in variable names feels odd at first, but you quickly get used to it. The pain of case insensitivity by default never goes away, especially when someone from New York or London modifies symbology code. On the other hand, the lazily evaluated dataflow subset of the language with close database integration is really elegant, and it deals with collections much more elegantly and uniformly than Java/C++/Python/JavaScript.)
In a previous job, after a few years I settled on (Type, Expiration, ISIN, Listing CCY, Listing MIC, Trading CCY, Trading MIC, Fudge) as The Tuple that could handle almost everything…. Once in a while I would still get a messed up case of non-uniqueness, that’s what the “Fudge” is for :)
> New services like Packages and Codespaces rely on the `mysql1` primary db.
A lot of this is due to that DB handling core data to GitHub (notably organization and repository data), which the integrated nature of the new feature offerings forces them to interact with (syncing permissions from repos for packages, publishing from actions, etc.). The link between repos and codespaces is even more unavoidable.
We were careful with Packages to keep as few service dependencies as possible in the critical path, especially for things like the anonymous read path, so serving project packages to users for open source projects or package managers such as homebrew are as insulated as they can be (and I suspect were unaffected by this incident).
But at the end of the day, there is some data that is central to most everything at the company, unfortunately.
And of course it's worth applauding that I think most read traffic (whether to Packages or other services) worked just fine through these incidents, if I understand correctly.
E.g. codespaces is literally a web IDE, so extensively interacting with repositories and accounts would be involved in pretty much all if its tasks.
While we obviously have our own unique issues, this could never happen (in its entirety) at our company.
Like you would have all of that "core" data as Kafka topics and can safely interact with them without affecting the core services?
I know the answer is always "legacy" and "its been designed that way from the beginning" but I was wondering what do you think would have been the right way to design something like this to mitigate risk of current / future problems?
> The simple failover described on March 22nd (for example) took almost 3h to complete
and
> They lept to failover their db when the proxy had issues, when it seems the failover can take multiple hours to successfully complete
Our failovers are very fast, take a couple seconds most, and work with near zero down time. I'm sure there were also other issues at play here.
At some point it's actually cheaper for a coalition of international organizations to fund inventing a backwards-compatible GitHub replacement that is very resilient to failure, rather than wait for GitHub to get a measely enough budget increase from Daddy Micro$oft to shore up their legacy MySQL database.
I had to look it up. Wow... just wow.
https://bitbucket.org/blog/sunsetting-mercurial-support-in-b...
Today I learned that all my old code from university is just… gone.
Is it too much to expect for them not to delete data that we’ve given them for safekeeping?
I only see one mercurial repo (yamlconfig) from you, though: https://archive.softwareheritage.org/browse/search/?q=https%... But you can search them directly on https://archive.softwareheritage.org/ if you used a different username.
But they have a real competitor for Jira. Check out the recently beefed up GitHub Projects, my previous (scrum-ish) team was running on it and we've liked it more than Jira. Much simpler, but just enough for a dev.
Didn't check out whether it's available on enterprise github though (i.e. local instances).
There are so, so many and there have been for years. Folks use Github because it's what they know, not because an alternative is hard to find.
It's nice to finally get some comms, but this is incredibly late and incomplete.
We're doomed >_<
You would think it wouldn't be THAT hard to shard something like GitHub effectively.
I mean, all user accounts/repos starting with the letter 'a' go to the 'a' cluster and so on seems not exactly science-fiction levels of technology.
The Mythical Man Month has a few things to say about that.
(It's tempting to feel that the information is outdated, but in my experience it still seems true.)
This is an architectural problem, which even if they had the massive expensive brains behind something like mysql on their team they couldn't fix it.
(at least, I'm guessing, I think this kinda architecture doesn't scale even if they could kick the can down the road a few times..)
livejournal, facebook, twitter, linkedin, tumblr, pinterest all use (or formerly used) sharded mysql and most of these are at larger db size than github
i will also repeat my comment from another recent thread: i just cannot understand how 20+ former github db and infra people recently left to join a db sharding company. this makes no sense whatsoever in light of github's lack of successful sharding. wtf is going on in the tech world these days
I believe you have the chain of causality backwards here. In fact, I think it suggests that talent that went to planet scale is perhaps not the issue.
this is like if you were building a high-rise condo, would you hire the architects or management company from the building that collapsed in surfside florida? sure, they know what NOT to do next time, but that doesn't mean they do know what TO do
Like, there’s another angle here: management, yeah?
Another way of reframing it is “maybe the folks hiring at planetscale know the inside baseball about GH infrastructure”. For example: https://www.linkedin.com/in/isamlambert
github should have sharded years ago, every other large mysql user did so much earlier in their growth trajectory
Platform migrations take a very long time and it's very complicated especially with decade old codesbases. I will say the current team at GitHub are nothing but outstanding people and engineers with a difficult task of managing a very large deployment.
hundreds of engineers have worked deeply on sharded mysql at massive scale, many of us comment on HN!
but this is different. github is a 14 year old company, with annual revenue in the hundreds of millions USD
I literally cannot think of any other comparable size and age mysql user who has not successfully sharded long ago and avoided outages of this magnitude. and in return we get hand-wavy excuses of "complexity!!!" from their former vice president of engineering who was previously also their first DBA.
criticizing these excuses is not trivializing the complexity, it's more of a "we all did this at our respective companies, who also had a lot of complexity! why can't you? we are github users, we rely on you and are very unhappy, we want to know how this happened" and just getting excuses as an answer.
How about a little humility? You have no idea who else has similar problems out there. I'm sure Percona and other DB consulting companies would tell you otherwise, as would PlanetScale. (If this weren't true, they wouldn't be viable businesses.)
> github is a 14 year old company, with annual revenue in the hundreds of millions USD
You're making the classic mistake of thinking that having money means you can just wave your hand and hire whoever you want with whatever talent you need to solve your problems. The world just doesn't work that way.
As for the "hand-wavy" excuses, well, the people who know the issues don't need to disclose the dirty details in a public forum, nor are they often able to because of NDAs and other legal encumbrances. And it can be a career-limiting move to throw your colleagues under the bus.
You can choose to assume people are incompetent, or instead choose to assume that people are working as best they can under the constraints placed on them. I think it's better to assume the latter.
i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago, because that's what literally every other large comparable company has done.
there are front-page-of-HN threads about multi-hour github outages every single day this week. show me an example of another similar-size company having an equivalent meltdown please. only real equivalent is twitter during fail-whale days when they were only a few years old. they solved it early on, as should have been done.
i am not saying the entirety or even majority of github is incompetent, but i am saying there is clearly something extremely wrong and extremely unusual that led to this, and citing "complexity!" is just a pile of BS.
As I said, a lot of the problems organizations have are not publicly disclosed, and the people who know aren't authorized to talk about it. So I'm sorry, you're not going to get the examples you seek here. Companies like to keep their weaknesses close to their chest.
> i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago,
If not through money and hiring, how should they have accomplished it? You propose a goal but no actionable plan. That's worth diddly squat, both in engineering and in business.
if you think Percona and Pythian implement sharding solutions, you are deeply mistaken, this is not what they focus on at all. advice, sure. more on the perf and ops side. but not sharding implementation, and definitely not application side of things.
furthermore i am saying if other major companies were having this type of issue, everyone in the public would know about it! because the product/site/company would be down all the time. this isn't some thing you can hush hush. close to the chest? how do you keep daily outages close to the chest? completely absurd
there is literally no analog in the US tech world to a large 14 year old company having daily outages for a week, preceded by multiple outages per month consistently for the past year+.
i will stop replying now because it is clear we live in different realities or smth
I would also back otterley's position that there exists many other companies with larger non-sharded MySQL instances than github's. They may or may not have better reliability than github.
"They probably outsource the lock queuing logic to their Memcached layer." you are grasping at straws. the wrong straws. go attend some facebook conf talks, meet their engineers, read some arch papers, and stop speculating this nonsense. largest db tier there does not even use memcached, hasn't for 9 years. read the TAO paper!
horizontal sharding works great for many many companies with databases several orders of magnitude larger than github, why would it not work for github?
you are also saying there may be large companies with lower reliability than github? ok name them. again, if this was the case, everyone would know! their availability at this point is like what, 1 nine?
And frankly, given your attitude on this, I think you’re going to need to present some bona fides for anyone to take you seriously. Maybe you can even convince GitHub to hire you and solve their problems. I bet they’d be happy to listen to you berate them on how pathetic they are.
i have posted many correct technical details in various subthreads of this post. i don't care if you believe me, i do not live in your reality where constant downtime is totally fine and major companies are down all the time but no one knows about it
No, this performance cliff does still exist in the current mysql 8.0 branch. The Contention-Aware Transaction Scheduling added to 8.0 doesn't help the queuing problem with hot row at all. So facebook either doesn't suffer from the problem, e.g. all their writes are append-only so there is no hot row, or prevents the problem from happening, e.g. with API rate limiting built on top of something like memcached.
The improvement in innodb lock architecture mostly comes from splitting the locks into more granular levels. However, updates to the same hot row would require the same lock at the page-level. It is already as granular as it can be. So the tag of the bug report is correct, this is a problem of "lock queuing".
The hard part isn't solving the scaling problem, but the uncomfortable conversations with leadership about how chunks of roadmap will be drastically delayed. Inevitably, leadership makes a decision to kick the can down the road, further exacerbating the problem.
I think the complexity of GitHub’s data management requirements would surprise you. Better to withhold judgment until you possess all the facts.
this is classic Microsoft. spend a ton of money for something very valuable -- in this case virtually all developer marketshare -- and then casually pedal it into the ground while you lie about the KPI's to C levels (IIS marketshare on netcraft as a function of parked websites at GoDaddy to dominate over Apache) and keep it on life support with other revenue streams (XBox) for the next 16 quarters until it becomes a repulsive enough carbuncle to shareholders that it gets the axe (Microsoft phone.) then in a year, limp into the barn with another product nobody else but you could afford to buy (minecraft) and slowly turn it into a KPI farm for Microsoft account metrics to drive some other failing product (Azure) and keep the C level happy while you alienate virtually every player with mechanics or requirements they hate.
galera has lower max writes/sec than a traditional async single master because it's a cluster. the other members of the cluster need to ack the writes, and all members are doing all the writes, so adding machines does not increase your max writes
They don’t know Microsoft for anything other than ruining Minecraft. They didn’t know Microsoft made the Xbox or even Windows.
They made this statement after Microsoft forced them to migrate their account they’ve had for 5 years to a Microsoft account. That broke their computer for a few days and reset their games. For no useful reason.
I doubt they have enough to prove wrt SQL Server to make it worth going through that again.
If you meant that and I misunderstood, the other issue is that you don’t migrate to a different RDBMS. You rewrite half of the app and then spend a couple years fixing issues.
Nothing I have seen about Azure SQL fills me with confidence about their technical capabilities.
Performance at all tiers is woeful, the networking connectivity is madness, and disaster recovery doesn't.
"Ping" it using a trivial query such as "SELECT 1" a thousand times in a row and draw a histogram. Or just eyeball the numbers. I did this recently and had response times over 12 milliseconds regularly. For comparison, a small and cheap IaaS VM running SQL Server can get down to the 150 microsecond range and stay there.
For this 100x degradation in performance you get the privilege of paying several times the cost of an IaaS VM + SQL license.
Azure SQL proxies connections. I mean sure, the documentation says that they can do a "redirect" instead of a proxy, but not if you use any of their "private" network connection options. You would think that is some sort of simple IP header change implemented by the switching gear in hardware, but you'd be wrong -- it tunnels through what is essentially a VPN -- with all of the predictable issues. These tunnel VMs don't have accelerated networking turned on, for example. So you can have "business critical" tier on one end, and huge VMs with accelerated networking on the other end, but the traffic in between is being routed through some 1 vCPU virtual router appliance processing packets in software.
Anyone with database admin rights can alter firewall rules via SQL commands. These are via SQL Server accounts and hence they're just a username and password, no MFA or anything. If you can find a SQL injection vulnerability you can punch a hole through the "firewall". Brilliant.
The firewall supports IPv4 only, and uses different CIDR syntax to everything else in Azure. It doesn't support Service tags, or logging, or monitoring, or anything really. It exists only to tick a checkbox.
Unlike most other Azure services, SQL doesn't integrate with Azure Active Directory or RBAC properly. So for example you can have one AAD group as the SQL Admin. No other rights, no list of principal IDs... just one admin group. All other delegated permissions must be done through SQL commands, blocking the use of ARM templates, Policy, custom roles, etc...
If you delete an Azure SQL Server instance, it deletes all backups associated with it. Sure, they recommend that you put a "delete lock" on the server, but then most administrative operations become impossible because you can't then delete any child objects. And even if you do create a delete lock, that can just be deleted. There is no way to say "backups cannot be deleted for at least 'x' days", even though the underlying storage accounts support this feature.
DISCLAIMER: Some of the above may have changed since I last checked, always read the documentation and/or verify with support if your data matters to you.
That... is a massive footgun. Essentially a 1-click method to destroy your business.
I'll set up delete locks ASAP, thanks for the tip.
I would argue it’s more performant than vanilla MySQL and it supports multiple write masters using peer to peer transactional replication an enterprise license feature (https://docs.microsoft.com/en-us/sql/relational-databases/re...)
There are certainly some very niche use cases where you need it, but there’s a reason why tech stacks don’t use sqlserver. I think there are better ways to handle durable transactions and redundancy than using sql enterprise with sql‘s replication.
functional partitioning is a band-aid. you do it when your main cluster is exploding but you need to buy time. it ultimately is a very bad thing, because generally your whole site is depenedent on every single functional partition being up. it moves you from 1 single point of failure to N single points of failure!
> In addition to vertical partitioning to move database tables, we also use horizontal partitioning (aka sharding). This allows us to split database tables across multiple clusters, enabling more sustainable growth.
I disagree, functional partitioning is not a band-aid, but an architectural changes that in the end can reap much more benefit than simple data sharding.
>> your whole site is dependent on every single functional partition being up. it moves you from 1 single point of failure to N single points of failure!
Not necessarily, it can also be that only some parts of your site are dead while others work perfectly fine.
then they fix it in post-mortem but pattern just repeats. i have seen it so many times! used to be much worse in the earlier days of the cloud when VMs would go poof more often
Tech debt has to be paid in full, whether it's person-hours, or downtime, or both.
A DHT isn't a great fit for really rapidly changing data like metrics and such, but I figured that everything required to keep the basic git clone, commit, and web serving functionality is a pretty natural fit for a DHT.
I figured even code search had a relatively small amount of storage for indexing recent commits, with compaction of that data resulting in per-token skiplists stored into the DHT. That way, failure outside of the DHT serving path still allows code search for all commits older than the latest (incremental) compaction. Distributed refcounting or Bloom filters could be used for garbage collecting the immutable blobs in the DHT. The probability of a reference cycle in SHA-256 hashes of even quadrillions of immutable blobs is vanishingly small.
Some maintainer is acting in bad faith, someone else quickly locks them out of all the repositories that they have not yet thought to corrupt, then it looks really really bad if they find out that they still have permission to run the CI/CD scripts, maybe with malicious substitutions, even though you blocked them from being able to push commits up.
The other part of that is, it's not too hard to get right. See what GitHub was doing. People don't change ACLs that often, nowhere near as often as you read them. If they did, you could rate-limit them. And the rest of the problem is, you have one writer and many read copies in one place and you can enforce boundaries on how stale the data gets on the read copies.
Two hard problems, first is cache invalidation.
Of course, they have found out that they are big enough to shard, one would have hoped that they would have found that out in a gentler way, much sooner. But let's not pretend that their architecture makes no sense, it's a fine architecture, they just happened to outgrow it in a bad way.
If your CI/CD pipeline uses capability-based permissions, utilizing delegatable expiring cryptographic tokens (e.g. Google's Macaroons), then a stale read of the ACL (presenting a token incompatible with the latest ACL) would allow the CI/CD pipeline to start running. Completing all non-local I/O (i.e. observable side-effects) would require synchronous reads of a cryptographic hash of the latest repository ACL and comparison with an ACL hash in the token (with a slow path of re-verifying the token could still be issued in case of ACL changes). Presumably, this ACL hash would be a part of the repository's immutable root node, so in the common case of CI/CD pipelines storing outputs into a repository, this synchronous read is "free" in the sense that it's part of the transaction to change the reference to the repository's immutable root node.
Granted, you have to be careful and make sure all of the I/O that's not local to the CI/CD pipeline (i.e. all observable side-effects, e.g. releases, sending emails) goes through through an ACL check that involves a synchronized read of the repository root reference. The fast path is just checking the cryptographic hash of the ACL in the capability token against the hash in the repository's root node. The slow path (when the ACL has changed since the capability token was generated) involves re-checking the ACL and verifying that the user named in the token is still allowed to create the presented token.
The downside is a racy read of the ACL (in the form of presenting a cryptographic token based on a revoked permission) when starting the CI/CD pipeline results in wasteful resource usage. If this is a worry, an additional synchronous read can be added at pipeline start time to avoid this race. If the cost of this synchronous read is prohibitive, a cache of <user, revoked_permission, revocation_time> can be added to the CI/CD start logic.
The upsides are that you can cache ACL reads client-side (in the form of your cryptographic capability tokens), needing only synchronous ACL reads for non-local I/O (i.e. observable side-effects) and if an attacker starts a malicious CI/CD operation, locking out the attacker before the malicious I/O (releases, emails, etc.) completes results in the malicious non-local I/O (i.e. observable side-effects) getting discarded.
It is honestly a bit humiliating for a company like github to both says :
We have had issues with our database for years and have still not found a solution. I mean what?
We have been down several days in a row and we have no idea yet how to solve the issue apart from throttle limiting webhooks
- Was it DNS?
- Was it a bad config update?
- Was it an overloaded single point of failure?
There's rarely a #4
Build&Configure server, A few hour drive it down to the DC, rack the server. Get back home, try to access. No luck. Turned out I had forgotten I had forgotten to connect power and turn it on.
Although most of those fall under “bad config update” (although likewise that applies to DNS).
[0] https://blog.cloudflare.com/october-2021-facebook-outage/
A more accurate checklist item would be “is it the network?”
However I think it’s hard to ignore in light of recent issues that there has been a large amount of attrition in the last year or so, specifically to PlanetScale by some of the most senior/database focused engineers.
Edit: just noticed that you’re the CEO of PlanetScale. I don’t mean this as a slight against you or GitHub or the folks that moved. I genuinely do just think that the number of people who shifted over left GitHub with a fairly large experience and knowledge gap. I don’t think anyone is to blame, it’s just an outcome of organic evolution and changing of a technical organization.
Others have said the only recourse is to contact support. I finally pulled the trigger today.
I know git does not need to consult any database but its own when committing. Auth is solved with tokens and url matching.
I get the need for the fancy database for the all the non-git stuff, but it is very concerning this stuff is in front of the meat and potatoes. Sounds like Microsoft is making another Skype/Teams disaster...
> We were able to identify the load pattern during this incident and subsequently implemented an index to fix the main performance problem.
What actually is "load pattern" in these cases?
Technical debt - here's a real world example of what those words really mean!
I don't think we are going to see GitHub be up for a full month without an incident anytime soon and I guess my entire comment chain [0] on the whole situation has aged for two straight years in a row, especially yesterday's one from [1]:
>> Until the next time GitHub goes down again (hopefully that won't be in another month's time).
*Goes down the very next day*
That says it all really. Lets reset the counter and try this again.
ResourceExhausted desc = transaction pool connection limit exceeded
My unsolicited advice: this looks like a classic case of https://bugs.mysql.com/bug.php?id=53825, which remains unfixed after 12 years.Here's the solution space:
* Oracle added NOWAIT and SKIP LOCKED to MySQL 8.0, and called it a day. This is a bad solution because it just changes from one extreme (unbounded queue) to another (no queue).
* AliSQL added Statement Queue via hint (their older implementation extented the UPDATE syntax instead), which should work great, but do require modifying the write queries: https://partners-intl.aliyun.com/help/en/doc-detail/144127.h...
* At $prevjob we added block-level queue within InnoDB, based on an older experimental patch from AliSQL which they didn't actually bring into production. No app change required.
* That bug was reported by Facebook, but there is nothing relevant in their MySQL fork. They probably outsource the lock queuing logic to their Memcached layer.
The problem with this bug is that, without a correct analysis, the knee-jerk reactions are typically:
* adding more CPUs
* throwing in even faster disks
* bolting on a sharding layer like vitess
* blaming your DBaaS providers (well, they are to be blamed if they intend to stick with "vanilla Oracle MySQL")
Question to Github: I see you running pprof on vttablet, but have you run perf on mysqld?
Is all cyber warfare completely underground and invisible?
You could in theory pull off an attack like that but it seems sort of dumb because it's not destructive/permanent so unless the timing it really critical it's probably a huge waste of a (likely expensive) exploit chain.
If you have deep enough access to degrade performance of the database you likely have deep enough access for exfiltration or other more nefarious activities.
If you want your enemy to fear the repercussions of attacking you, you need him to understand the strength you have to fight back. And in the case of cyberwarfare the only way to show strength is by hacking stuff... so I imagine there's some value in smaller not-so-destructive attacks.
a: building new systems on an internet scale non-relational datastore b: migrating relevant features onto an internet scale non-relational datastore
Are there a lot of folk who are operating at Github's scale still deploying core services to platforms like MySQL?
I say this as someone that despises MySQL.
Same obviously applies for PostgreSQL but is actually slightly harder for a number of reasons with less out of the box solutions like Vitess available.
I blame the program providers.
Some Debian maintainers are trying to do this simple querying of complex configurations (dpkg-reconfigure <`package-name>`). And I applaud their limited inroad efforts there because no else one has seem to bother.
I have made a bash script to configure for each Chronyd, named, sshd, dhcpd, dhclient, NetworkManager, systemd-networkd, /etc/resolv.conf, amongst many. They try and ask simple questions and glue appropriate settings then run their own syntax checkers (most are provided by the original stream).
Postfix, Shorewall, and Exim4 remain a nightmare to my evolving design.
CISecurity and other government hardening docs were applied as well and then some I took even further like Chrony had its file permissions/ownership even further and MitM block feature as well.
These are dangerous scripts where it can write files as root but as a user, you will instead get configuration files written out in appropriate directories under `build` subdirectory.
If these designs work across Redhat/Fedora/CentOS, Debian/Devuan, and ArchLinux well, I may forge even further.