I made a mistake with Terraform and Azure made it worse
craigstuntz.com
craigstuntz.com
There is also the prevent_destroy[0] meta argument for resources but afaik it has no effect when you remove the resource from your .tf files[1], so it would not have helped in this case.
[0] https://www.terraform.io/docs/language/meta-arguments/lifecy...
In addition to that things can be made even less error prone. Ive done this using yaml pipeline in azure devops. The plan task can be used to set an output variable which indicates if the generated plan contains any changes. That boolean value is used as a condition to trigger a manual verification task which basically prevents apply running if there are any changes without reviewing it first.
As the op mentions, the generated plan is an artifact itself that is used in a following apply task
Unfortunately, I've never been able to get much mileage from the Terraform prevent_destroy lifecycle option because it can't be set from a variable. Most of my configurations use a module and pass different variable values per each environment. I'd want the lifecycle flag in production, but maybe not dev.
There is a trick around this. Code 2 mostly identical resources, one with prevent_destroy true and the other with prevent_destroy false. Give both resources a count. You can make the count dependent on interpolation. And the count can be zero. So depending on the interpolation result you have resource(s) with prevent_destroy or without. For the "wrong" ones you just create 0 of them.
https://stackoverflow.com/questions/53727357/terraform-how-t...
https://stackoverflow.com/questions/53744441/how-to-refer-to...
Plus the combinatorial explosion of all possible conditions.
Not a good trick.
When I want to deploy a test system the resource can be destroyed after I am done.
Of course this is not really elegant code. It violates DRY, I need to make sure that my test resources have identical configuration as my production resources, except this one attribute.
Maybe your use case is different. I never change the prevent_destroy attribute during the lifetime of the resource.
It's perfectly OK to have a completely separate Terraform project that just configures the DB initially (or even manually, I see lots of places running DB's that predate Terraform with immutable infrastructure for everything else), and applies minor non-destructive changes in the future. This way you get the benefits of IaC, but the DB plan doesn't participate with the rest of your infrastructure that IS ok to blow away and re-create at will.
BTW, Amazon RDS backups work the exact same way: Destroy the database and the backups are also destroyed. Therefore, same region automated RDS backups are fine for day-to-day, but in a true "DB goes poof" disaster you should expect that you WILL lose them too! You need cross-region, or even better, cross-account DB replication or snapshots to survive this.
(It is possible to create a custom RBAC role that excludes zone deletion only, but this is very fiddly and not-quite-the-same in complex ways.)
It allows for running a plan or apply on the whole site, or each individual module in the site.
So in this case it's easy to just apply the database changes, and then run the whole site plan to make sure everything is indeed in the right state.
APP-LOC-SHARED -- wildcard certs, deployment scripts, etc...
APP-LOC-ENV-Common -- gateways, etc...
APP-LOC-ENV-Data -- databases, storage accounts, pets.
APP-LOC-ENV-Web -- cattle
APP-LOC-ENV-App -- back-end cattleDatabases (or cattle, how or what ever they may be) are always created as a separate stack in your IaC of choice. The values you need for this stack are always outputted for consumption into other pipelines.
The benefit being you can make lots of small changes to your upstream stacks without breaking your database. And you can place additional controls around the database stack to prevent instance replacement from occurring.
It's a reference to the idea that you should treat your servers as "cattle, not pets". In other words, they're disposable and you should be fine with destroying one and creating a new one to take its place.
The parent is saying that it's fine to treat databases more like pets, where they're pampered and looked after carefully, and fed treats - much like DBAs are, I believe.
Excellent point.
AWS provide a couple of decent options for cross account backups, the Aurora Snapshot Tool [1] and the AWS Backup service. I've used both successfully.
The author, a "Director of Consulting" might want to do some training (Both Hashicorp and Microsoft have free training).
Why in the hell would you have the same statefile for different environments? Why would you have these environments in the same Azure subscription? Why would you run your terraform so infrequently that you'd forget about a botched statefile move? Why would you not read TERRAFORM PLAN (It's LITERALLY WHAT TERRAFORM DOES)?
I also suspect that while the author's probably heard of a CI/CD pipeline, they're running their IaC from their local machine, given the tone of the article.
Why, for a production database that contains information not contained elsewhere, would you not configure an Azure Recovery Services Vault?
Like I get it, people make mistakes. I once did something similar with an overenthusiastic use of terraform destroy, but this guy just seems like an absolute cowboy who doesn't really know what he's doing.
It’s the same with something like a programming language where there are thousands of best practices and foot guns, but infrastructure as code is more dangerous and much newer. There are also less books, venues and even training courses to learn these practices.
There are also a vanishingly small set of engineers who have actually done this in production at scale so it can be hard to find experienced people.
In several firms I've worked in now someone missing so many things wouldn't be above consultant level, and wouldn't be approving PRs let alone deleting a live prod off from their local laptop without some significant questions being asked.
I get everyone must learn things, but when you position yourself as a technical expert (Director, in this case) you should have enough experience to be a bit more thorough with your work so if a mistake happens, there's a way out, or just not make what amounts to several design and implementation mistakes.
Part of what I think makes this a little egregious is the author didn't f** up his own systems, he f**'d his clients. I'd understand a little more if it were "I'm the in-house guy upskilling" rather than "look at this mistake I made as a (presumably) highly paid outside consultant literally brought in to make sure stuff like this doesn't happen".
However, I still must give absolute kudos for sharing mistakes publicly. We all do make mistakes, and most people try and hide it. When the author realised there was an issue, every step after that was handled like a pro.
I do however think that Terraform should have a first class support for different environment without the need of a wrapper to facilitate that.
https://www.terraform.io/docs/language/state/workspaces.html
> Certain backends support multiple named workspaces, allowing multiple states to be associated with a single configuration. The configuration still has only one backend, but multiple distinct instances of that configuration to be deployed without configuring a new backend or changing authentication credentials.
You might want to actually read the entire blog post before going off on a tirade.
I am not OP btw.
> Why, for a production database that contains information not contained elsewhere
* This is a dev environment - the one that he uses to to dev work where its perfectly okay to corrupt, which is why dev environments exists.
> Why would you run your terraform so infrequently that you'd forget about a botched statefile move?
* Why on earth is it surprising to you that theres a dev environment sitting somewhere that could be unused for some time?
> I also suspect that while the author's probably heard of a CI/CD pipeline, they're running their IaC from their local machine
* It is perfectly okay to develop and test changes on a local dev environment before putting it on a CI/CD pipeline. OP mentions that there is an Azure Devops pipeline in place.
> this guy just seems like an absolute cowboy who doesn't really know what he's doing.
* In dev yes. He clearly does know what he is doing.
This of course would never happen with AWS and cloudformation changesets. OP is sharing gotchas that arise when using Azure and Terraform, this is useful.
1) It is so weird to me that every cloud provider deletes backups when you delete the SQL instance. We take offsite backups that are decoupled from this process. Fortunately someone made a shell script to do this that has worked quite well: https://github.com/ovotech/cloud_sql_backup (you can then just copy the Cloud Storage buckets to S3, so when you Google account gets banned you still have your database).
2) I hate to be "that guy", but I'm starting to wonder about gitops. I really, really like having all infrastructure changes recorded in machine-readable format with history. Pry it out of my cold, dead hands. But you also lose a ton of important tools, like the diff and the sanity check before you deploy. You could have a CI rule that does the diffs, but then CI has to touch production which is not ideal. You could have CI do a dry run of HEAD and a dry run against your PR, and diff those two, but honestly no CI systems really let you check out multiple branches and use them as an input to the script, so have fun hacking that up. You end up with workarounds for workarounds and the net result is that your auditable infrastructure comes at the cost of taking a human out of the loop (and slows down experimentation). I don't think the model is quite right quite yet.
Doesn't seem hard? It would be a one-liner "git worktree" pre-build step in Jenkins, for example.
Just apply the changes against staging. Test everything is kosher. Then merge into master to apply to prod.
Not actually true for Google Cloud CloudSQL (at least MySQL). You can delete the instance and you don't see any backups any more in the Google Cloud Console, but they are actually there and you can restore from them. You need to know the instance name though.
I think the backups will be kept for 30 days. This is the same time you can't reuse the instance name.
Google could improve the UI about this a lot and display also deleted instances.
So while I think it is OK for the author to be a little humbled by his mistake, I would actually place the blame on the misdesigned environments. We should try to expect human mistakes and make them avoidable (by noticing them early in the test environment) or prevent them alltogether (by automatically testing and only allowing the change to production, but that is far harder to setup).
This is a nice aspirational goal, one that you should definitely aim for, but I have never seen a situation in a non-trivial app where the test and prod environment are even close to identical.
All large apps I've worked on end up with external dependencies that themselves don't have a test system and have real world side effects (like buying things), so straight off the bat, all of them need to be mocked out.
Then there's no way we'd get budget for a 1000+ node cluster on our test system, so naturally there are entire classes of high load / large network bugs that are less likely to crop up.
Finally the test systems are often not allowed to handle prod data for regulatory reasons. Thus, everything needs to be either synthetically generated or stripped versions of prod data, which unsurprisingly sometimes behaves rather differently.
Even if you think your test environment is a pretty good replica of prod, chances are it's probably not as good as you think. So you should always have processes in place to be able to slowly rollout deployments and halt/rollback if there are issues. Everything should always be backed up. And importantly, this should not all be controlled by a single deploy script with global permissions - anything that deletes backups should need explicit multi-step approval.
You have to invest big time in automation to build up and tear down any decent sized environment to save the spend there, and that itself costs money.
One of the smoothest run environments I've ever worked with had ten silos. A failed update would not impact more than 10% of the users, and we could quickly redirect them to the remaining 90% of the platform without a material performance impact.
And I wonder how many platforms can be made to work like that.
(we use gitlab-ci's built-in review process for approval).
At the end of the day, approval's still a human job, and humans make mistakes. Right? :D
Terraform is an incredibly powerful tool, and you can make some monumentally huge mistakes with it.
That said, Terraform breaks so often that if we did it all automated, we'd have a million more Git commits from trying to fix broken apply's.
- plan against staging
- get a PR approval
- apply against staging
- plan against prod
- apply against prod
- merge
Being forced to plan (and get someone up review said plan) before applying makes it far, far less likely you’ll do the level of damage described in this blog post.
The best strategy is to have a repository for your modules only so you can specify the version[0] you want to use, and separate environments by folders.
[0]: https://www.terraform.io/docs/language/modules/sources.html#...
re: SQL Server backups, I assume this was VMs running SQL Server rather than Azure SQL DB (the managed one) ? If it's the latter, then I think the backups will be retained even if the RG is deleted.
AWS RDS has a feature where if you delete a DB instance it prompts you to take a final snapshot of the data. I haven't used Terraform in just over a year now (have been using aws-cdk), but as far as I remember, Terraform deletes of RDS instances would require you to add a 'force' argument/option to override the prompt for the final snapshot.
Something like this seems like a no-brainer for Azure SQL when deleting instances.
Also, with TF its always best to plan, write the plan output to a file, and then run apply against that plan (after having studied it!)
First time I contacted them I was not hoping for much. But after a few contacts I realize they must be one of the best support orgs in the world.
They can quickly help out with anything from very trivial questions to helping out with very difficult faults, and do this without having to be harassed about escalating my tickets. Outstanding.
Spent over a month trying to resolve the issue but I must've been talking with tier 1 support because it was obvious the support technician wasn't very technical (asking very basic questions that I had already answered, couldn't answer any of my questions without getting back to me days later). Combined with the quirks and performance issues (30+ mins to setup VPN gateway) it was a very poor experience and I wouldn't choose to work with Azure in the future.
It can seem onerous at times, but ultimately it's a far smoother process than legacy manual or scripted deployments and is very repeatable and visible. It's also a good thing to have lots of eyes on critical infrastructure changes anyway.
Anyone know if az cli or powershell module prompts users before it deletes a resource group?
What it supposed to do is when was about to remove resource group, it would still see that other resources are relying to it and not to remove it.
AWS has done a great job at making it difficult-to-impossible to tear everything down in a single command. There are rate limits on resource deletion and no single place to remove items in bulk if not managed by Cloud Formation. I spent the better part of a couple days trying to clean up an account from all of the random supporting services that get created over the years. If I'm wrong about this, please let me know!
It sounds like the user running Terraform had access rights to multiple environments. This sort of thing was inevitable.
I also infrequently backup single sources off cloud to prevent ransomware payoffs. It has the secondary feature of providing a recovery source if things go off a cliff in the production environment.
Glad I'm not the only one that saw this as the right choice.
You are giving a relatively new tool full access to your datacenter. Here be dragons. This mistake is obviously something that would never happen when using the official UI, at least not without descriptive explanations and warnings in the process.
Terraform's "-detailed-exitcode" will help you tailor the pipeline if there's nothing to do (0 - plan success, no changes, 1 - error, 2 - plan success, changes).
I'm also of the opinion that prod/non-prod subscription and resource groups pairs should be "vended" to you by another department (i.e. you're a user not an owner). This prevents you accidentally destroying the entire hierarchy.
I think you meant stateful.
We don't do this because preproduction is being used by the business actively. It is treated like production environment. But it depends on companies.
Terraform will never delete something you don't tell it to.
Not necessarily. With some providers (e.g. Azure), Terraform will fail to recognize automated behind-the-scenes changes and try to revert them, causing serious breakage. This is why the "ignore_changes" meta-argument exists. See https://itnext.io/how-and-when-to-ignore-lifecycle-changes-i...
what the hell is Azure doing pretending to delete things?
How long do Azure keep your data after you think you've deleted it?
Is there a way to ensure data stored in Azure is actually destroyed when you ask for it to be destroyed?
Update: I'd forgotten that even filesystems often do a soft delete. https://lwn.net/Articles/462437/
Part of the reason is performance, I would assume, since a large delete could be hard on the overall system, and could slow Azure down for you + other customers. Also because some customers accidentally delete critical resources sometimes. (Or forgot to copy down any important config options from the resource before deleting it. I have some experience with that mistake.)
To solve his problem, after switching to the data source type from resource he needed to manually run `terraform state rm` to get rid of the resource from the state.
From Terraform's "An Overview of Our Recommended Workflow"[1]:
"The best approach is to use one workspace for each environment of a given infrastructure component. Or in other words, Terraform configurations * environments = workspaces."
So 1 workspace != 1 environment but workspaces are indeed intended to handle multiple environments.I've personally struggled with managing multiple environments using Terraform so I'm interested in what the best practices are.
[1]: https://www.terraform.io/docs/cloud/guides/recommended-pract...
"Blow Us All Away"