Don't Make My Mistakes: Common Infrastructure Errors I've Made
matduggan.com
matduggan.com
> Don't migrate an application from the datacenter to the cloud
Good advice that follows the pattern of don't do a rewrite with a plan to cut-over. Instead make two and gradually transition to the newer one.
> Don't write your own secrets system
Follow-on to 'don't do your own crypto'
> Don't run your own Kubernetes cluster > Are you a Fortune 100 company? If no, then don't do it.
> Don't Design for Multiple Cloud Providers
This seems like an extension of YAGNI. If you're not using it, don't build it. If you really will need it, use it as you're building it. A good example is if you're doing multi-datacentre redundancy using standard tech, it might be worthwhile to do multi-cloud. But have an actual threat model in mind that this mitigates, e.g. cloud provider 'A' may become a conflict-of-interest competitor in the foreseeable future.
> Don't let alerts grow unbounded
> Don't write internal cli tools in python > Nobody knows how to correctly install and package Python apps. If you write an internal tool in Python, it either needs to be totally portable or just write it in Go or Rust. Save yourself a lot of heartache as people struggle to install the right thing.
+1 This one rarely gets mentioned.
One that I would add, that I haven't experienced but can certainly foresee is "don't switch major datastores". Switching between MySQL and PostgreSQL might be rough but doable. Switching between MySQL and CockroachDB for a large, heavily-used app could stall developing new features for a long time or make everything take 5x longer. The reason isn't the query syntax, it's the query characteristics. Using a std rdbms gives you the ability to do many relatively quick round-trip queries (though you should avoid N+1's) some can be tolerated. A distributed DB can have high throughput but will have high latency for simple queries.
This seems an infrastructure extension of treating compile warnings as errors.
At that point I think it's pretty competitive with "go get". In both cases users have to install one thing, then use it to get your thing.
Haven't seen an internal rust tool yet.
> Users who can’t do that shouldn’t be using CLI
Users who can do that will still find it an odd departure from convention.
It'll never be as clean as a static binary build, but it saves us from having to build out two language ecosystems when the rest of the company uses Python for everything.
Python is awful for this stuff
It's never fun. It's never pleasant. But to be fair, if I have a CLI tool that needs a deployed SSH client, or Tensorflow, or SDL or Qt or something else, I'm not convinced packaging gets much easier no matter what language we're talking about. If your use case is simple, Python is easy enough to deploy, and Go is even easier. If you can't disable CGO or need a third party component, I imagine the fun is just getting started anyway.
As a counterpoint, awhile back, discovered that Golang had a minimum kernel version requirement. That pretty much eliminated it as a possibility for writing tools for legacy systems. Python was viable though, Bash moreso. :) Couldn't tell you if that was still a requirement for Go today.
Agree with you strongly. Also by committing to a database you can take advantage of things it does for you rather than trying to stay "generic" and portable. Pick a database (I like Postgres), and then wring everything you can out of it.
> Don't migrate an application from the datacenter to the cloud [..] Instead port the application to the cloud.
You can certainly do it without an app rewrite if you have good infrastructure engineers with the support of a small competent dev migration team. The real questions is when and why you should do it. One valid scenario is: you want to sell the app and an AWS setup is a lot more appealing to a buyer than a custom self owned or collocated setup.
> maybe even doing something terrible like connecting your datacenter to AWS with direct connect in an attempt to bridge the two environments seamlessly
You can certainly do that too, and AWS can be used as a cold disaster recovery option.
> Don't write your own secrets system [...] how do you keep from hitting this service a million times an hour but still maintain some concept of using the server as the source of truth
Simple, you get the secrets very rarely: at deployment. If you need to change anything you redeploy that part which should be very easy and fast if you got your orchestration/config management in top shape. Why would you do it? To avoid paying the Vault enterprise license and still get a highly available, version controlled and even more simpler and stable service.
> Don't write internal cli tools in python [...] Nobody knows how to correctly install and package Python apps.
This one is particularly sad snapshot of the state of the industry expertise. Debian packages solve this easy and completely. Oh, you don't understand your distro packaging system? Stop wasting time on blog posts and start reading documentation!
> Simple, you get the secrets very rarely: at deployment.
Do you mean writing custom scripts to get secrets? Where do you store secrets in such a scenario? What if you need to change secrets at runtime.
Vault is nice in that you can fine grain access:
- ACLs which allow some team members to be responsible for setting secrets and other team members for using them,
- temporary credentials for well known DBMSs which helps secure logs from leaks,
- PKI Management,
- temporary SSH keys.
And more. Not sure how you'd get that with a deploy time script.
Anywhere basically, a HTTP server with directory listing is enough. You don't need to roll your own security management, you can just encrypt secrets with ssh or pgp keys.
> What if you need to change secrets at runtime
I certainly don't need that as an infrastructure engineer. If you want that for development be my guest, pay the exorbitant Vault license and don't come crying to me Vault is down. Vault is always down, you just introduced a critical runtime dependency on an immature solution. Your problem, I have the email to prove you took responsibility for this decision against my recommendation.
Also, what's immature about Vault?
I had read them, and I'd anything just to be able to not deal with that mess. I believe there should be special people with a proper mental constitution and a high salary to deal with it. Not me.
The problem is complex so the solution is not one youtube view away from being understood. Deb packages helped Debian to ship tens of thousands of software packages with frequent updates for years and years with few maintainers.
I am sorry to say this, but calling it a mess says a lot more about you than about the package formats.
So it is a mess of a lot of solutions condensed over a period of 30 years.
I did look into two of such package management systems (deb and ebuild), and I'm happy to use any. To keep my system up to date they are fine, but not for anything else. It needs a lot of domain specific knowledge to make a package. Knowledge that is useless outside of the world of packaging. Knowledge detoriates when not used, and every time like the first time. Grr.
> I am sorry to say this, but calling it a mess says a lot more about you than about the package formats.
You shouldn't feel sorry saying it. I said it already. Though I used different words, but it is essentialy the same idea: there should be special people to make .deb packages.
Probably you tried to say that it is me who is special in that regard? Can I advise you to can check your intuition? How many developers bother with preparing .deb packages? Is it 10% or 90%? How many github repos contain .deb? Most of developers who doesn't prepare .deb packages are like me. They would use flatpack, or some other format allowing to capture an environment easily. But they wouldn't use deb.
As such I will stop arguing and offer something that you might find useful in the future: arch linux PKGBUILD (https://wiki.archlinux.org/title/PKGBUILD) - an order of magnitude simpler than debs. You can learn this in a day and you can use them on any distribution.
At $previousjob I watched in glee as a conference room full of sales people smugly told us the whole product was "cloud native" on AWS, and then the realization slowly spread from face to face that they were pitching Amazon's biggest competitor and we wouldn't allow our data to reside on their systems.
https://news.ycombinator.com/item?id=29434740 - 130 comments
There’s a temptation, when the sun is shining, to make extra work for oneself. Let’s add a compatibility layer around our app so we can port it from AWS to GCE on a whim.
Then the storms come and you realise the last thing you want, when under pressure, is to have added-complexity to your business logic. Oof.
On the final point, pip install . and a tiny setup.py has worked really well for me. Putting these all in one place and having a house style is nice too. There’s probably even a debhelper to turn them into native packages, though again, that seems like extra makework compared to just throwing your junk onto automatically provisioned production hosts. Vive l’/opt!
If you don't completely bake your app into AWS by using the AWS SDK all over the place or using a database that only exists in AWS or something, moving individual apps is just not that bad. You still gotta solve the cross-cutting things like logging and metrics, but you gotta do that anyway, and that shouldn't (!) require code changes to your app.
To be fair, that's all moot if you're not using Kubernetes in the first place. As well, things like EKS pod roles add great value that you'll have to sacrifice to truly call your app "portable".
Some industries legally require you to plan for the case where a cloud provider cuts you off. So sometimes you don't have a choice.
In those cases, it wasn't "This app needs to be AWS/GCP active/active". It usually meant: if you're building out core functionality to use AWS, you should build out on one of GCP, Azure, etc, too so there's an alternative available. The long term strategy was always to just stick things where they were cheapest to run (on prem, AWS, etc)
(So the same thing came up with physical equipment like Dell and HP servers)
From the business side, there was also concern about Amazon becoming a competitor in the Fintech space but that wasn't regulatory related
I agree with the spirit of what you wrote, but I think that not allowing it is ultimately even worse then the unnecessary dependencies.
Limiting your dependencies should be a choice the developers themselves are meant to make, otherwise you risk estrangement and having unmotivated workers that don't feel responsible for their own work. That ends up costing more in the long run I think.
Note that there is one big Kubernetes consultancy (namely, Flant) that requires you to run your own Kubernetes (managed by them) and not the managed offer by the cloud provider. They do it because they know how to run Kubernetes and don't know quirks of a zillion managed Kubernetes providers.