Terraform should have remained stateless
bejarano.io
bejarano.io
4. The previous Terraform configuration
This is effectively stored by state.We need this because if a resource is removed from the new config then Terraform needs to be able to delete the existing resource from the world. If we don’t have the state then Terraform must either:
1. Not delete it from the world
2. Or risk deleting something not managed by Terraform
If everything were managed by Terraform then perhaps we would not need state, but this is not realistic in my view.Could you give an example of a situation / transition that cannot be captured correctly by managing the state using these types of tags?
Cloud providers do not provide enough metadata to enable that mapping across all resources.
Do you have an example to back this up?
The article explicitly mentions OctoDNS as a stateless configuration management system for DNS as a good solution.
I guess it lets an attacker know that you're using Terraform, which might help them target their attacks.
Yes, if that's the case, then TXT records could easily be unsuitable. Depends exactly what metadata needs to be attached to your DNS records.
It's true that nothing extra is needed for simple/standard records, but once you start doing GeoDNS, failover, health check, etc. it's required.
In all cases thus far we've been able to find a way to store/indicate whatever we need.
(maintainer of octoDNS)
terraform is more than just cloud providers https://registry.terraform.io/browse/providers
IAM Users are taggable, but to get the tags on a given user, you must request them one user at a time from a known list of users. The "List all users" call doesn't return their tags. Obviously this is less of an issue for the TF state use case, but does add to the API call overhead for any tag-based approach.
Cloud providers having bad APIs is definitely the default state.
But I do not believe leveraging tags and/or metadata is the right approach - configuration for these resources could potentially be large (e.g. GKE resources) and most providers will have a size constraint on their metadata and tag values. Creating a metadata/tag key for each configuration key would also get messy, but solves the value size problem.
Why wouldn’t it be possible to not store the previous state at all? Terraform’s job is to reconcile what exists with what is declared - we should be able to rely on the provider’s APIs to understand what exists, perform the diff, and reconcile the changes.
Let's say your TF config declares a database Foo. Your AWS account has databases Foo and Bar. How does TF know whether it's responsible for database Bar and whether to delete it?
Where would it store it's history to make the diff against it?
Some issues with that:
- Fetching the whole state would be hugely impractical due to the number of API calls
- The risk of losing state information by a resource being deleted outside of Terraform is greater
- Again, not all resources have metadata that could be used to store state
This isn't losing state information though! That is the state. If state information were kept outside this it would now be wrong which means terraform *would do the wrong thing.
With the information being stored outside the resource, we know that it was deleted and the metadata about it.
Except for the Cloud's and API's that don't provide them but we stil need configuration managed. This is the world we live in, and in that world Terraform is the solution to the problems that we seek. In an ideal world Terraform would not be needed, but we don't live in that world.
Tags do not solve the deletion issue. It would require two step deployments, for example by adding delete = true to the config, applying, then removing the resource from the config entirely and applying again. But I don't think that's too bad tbh.
And also hoping that service's API has all the tools needed to find your scattered state within a reasonable amount of time in order to diff any changes you make in your declarations
Is this tool private or available for us to try out?
The stateless approach has worked really well for us.
So, fine for you in your personal small environment, but maybe not something suitable for Terraform in general.
If I have 3k machines running in an autoscaling group I think it would be ridiculous to have to hit each one of those with API calls to try and infer which were or weren't part of my state. Building a simple high availability VPC is about 72 resources just by itself.
I don't think people advocating for "the cloud is the state" realize how big even trivial environments can get- let alone complex ones.
State gives us a common schema and playing field to significantly simply the generation of dependency graphs and show drift. I imagine that even without a 'statefile', you would end up having to generate a similar graph in memory anyway.
If you turn the configuration into state, you don't need state anymore I guess?
Couldn't you do something like that?
Perhaps Puppet would have been the better analogy (i.e. `ensure => absent`).
While Terraform providers can be built to do a lot of implicit work, in practice they don't. In my experience, you still have to specify all the pieces.
The only place I can think of where auto-deletion came in handy was for infrastructure CI. However, that CI pipeline doesn't catch many critical issues, and it's very expensive to operate. So I can't say it's been a total win.
The whole point of TF is that it has state and doesn't require workarounds for these scenarios. Yes, you have to maintain state, but the state problems usually come from buggy providers, not Terraform itself. For example, try and use the GitHub provider to create a repo with GitHub Pages enabled from the gh-pages branch. It won't work because the authors didn't respect a fundamental rule of writing a provider: one state-changing API call = one resource. If you don't respect that, you have to do state handling yourself and you're almost certain to have state-related bugs.
Terraform would be much better without state. Not 10x better, but 2x.
When you change the name of the resource that will, obviously, break the connection, that's a big no-no when writing Terraform code.
One problem, as mentioned in a sibling comment, is if you don't have state then how do you know a resource has been deleted? One other option might be a system that compares the "desired state" from the previous version of the config. That might be an interesting approach - keeping a history of config. But that is, of course, state :)
I use TF a bit. I want to like it, but I do spend inordinate amounts of time faffing around with state files to make them match reality, and I'm not even doing particularly complex stuff. I'm told other tooling (eg: Bicep for Azure) dispenses with state entirely.
Regarding the state, it took me a very long time to come up with good organization for my code, but it works once you've gotten used to it. I really only need to mess with state files when there is a bug in the provider like the aforementioned GitHub issue.
Bicep/ARM still has issues with resource deletion, though. The default deployment behavior is to ignore resources that aren't described in a deployment template, so if you remove a VM from your template and redeploy, the VM will keep running until you manually delete it. There are a couple ways around this issue, but they all rely on having state external to the resource itself.
Disclaimer: I work on Bicep/ARM and think it's pretty great, but it's not perfect.
Orphan resources only exist as a concept if you manage entire estates in a single configurations, which is simplistic to the extreme.
If I use local_file to emit a certificate authority for an EKS cluster, does that mean every other file on my file system is suddenly orphaned? What if the networking team responsible for VPCs, routing and transit search for orphaned resources in my AWS account - should that include all my instances?
If you think Terraform would be better without state, I’d encourage you to put your money where your mouth is and build it.
What if terraform was able to perform a complete audit of your environment for each provider and say "this is what you've got, mate", then give you options for each resource:
1. Accept (turns it into terraform code)
2. Reject (removes the resource)
Taken further, this idea effectively means that manually creating something like EKS could be a way to auto-compose Terraform code and store it in git.
Pros:
* You have the options of writing code, using the console or using a script or binary to perform a complex setup
* You have 100% code coverage once Terraform compared code with reality and forced you to choose
Cons:
* Auto-generated code is shitty.
* Slooooooow without caching (which is only one of the functions of state)
To make these concepts work without a remote state you pretty much have to run your infrastructure tool as a service, and implement things like:
1. Polling of providers on an interval
2. Alerting when something changes
3. Eventually consistent cache that can be forcibly updated in the UI or during sensitive operations
4. "Infrastructure Composer" UI, like a vastly simplified IaaS console that only shows deployed infrastructure and lets you group infrastructure by service, and let you define code-level modules
5. Show discovered dependencies and ordering, and allow you to explicitly set them (these should be stored as tags on the actual infrastructure components where possible)
6. Use some hot AI to group code-level resources, compare their configurations and generate config data to apply to resources/modules using for_each loops, while respecting service boundaries and offering advice on when to refactor.
Edit: and actually, without a state there is no concept of a remote state. All codebases talk to the same "service" so circling a bunch of infrastructure and saving it to a specific git project would be easier than connecting to a bunch of remote states and reading specific resource days from them. Even if you're code repo doesn't declare a resource, the omniscient IAC service "knows all"
We keep coming up with these layers of abstractions around stuff that at the bottom is essentially already stateful but just not in a useful way. Doing it right from the ground up might help. Of course, left to the usual suspects, this would just turn into another Frankenstein blackhole of devops time. The usual suspects are billion dollar corporations that thrive on layers of complexity.
This is how we do our initial terraform for new projects. We build up what is missing by hand, then export the terraform to save time and effort in building our own internal templates. We throw away most of it, but it is very useful for not having to look up how every aspect is named and what options are needed.
0: https://docs.public.oneportal.content.oci.oraclecloud.com/en...
I believe, Terraform needs a mode that removes unmanaged resources, but that's quite the big ask and would result in yet another API / SDK change that providers would need to implement.
Measured cloud setup would have VMs A + X + Y
New cloud setup would specify VMs B + X + Y
You can easily identity X and Y as their names go unchanged, A and B would have similar config/metadata, instead of assuming A would be added and B would be removed you can ask the user if a rename happened.
This is not a new problem to solve by the way, this is how database migration tools handle column renames.
In most places I've been, Terraform is scattered, each run managing its own corner of the infrastructure. In this case, each Terraform run would delete everybody else's infrastructure.
Also, how do you know when to re-create things like random strings or numbers? Or null resources? Presently it's a mix of keepers (requires state) and taint (also requires state).
For example, if you changed an ID in your stateless Terraform, you'd have to insert some kind of code to destroy (or rename, if possible) the old resource. Or modify the Terraform DSL to include that kind of information, I suppose, and keep a historical record in the code perhaps. Then there's the question of what happens if someone modifies the physical resource out from under you -- could end up creating brand new resources rather than tracking the existing ones.
Also, it's nice to know that your Terraform instance is what created a thing -- if you ran a stateless `terraform destroy`, it's possible you could be deleting resources that someone else created that happened to match what your Terraform code defined. More of an edge case, I admit, but at scale these things have a way of happening...
That said, resources that don't have "physical" IDs work similarly to the stateless model by necessity. For example, VPC route table rules: [1, see the "Import" section].
Refactoring Terraform code is super annoying because of state, though, I'll give you that.
[0]: https://docs.ansible.com/ansible/latest/collections/amazon/a... [1]: https://registry.terraform.io/providers/hashicorp/aws/latest...
Do you have any advice or best practices to properly deploy and manage instances using Terraform? Also, what led you to use both Ansible and Terraform?
The first thing to realize about using Terraform is that _you_ can extend it with your own providers written in a language like Go and _you_ can cobble together your own modules written in the TF configuration language to orchestrate multiple providers or do repeatable work. There are really solid open-source modules for AWS operations that smooth out the kinks in the AWS API.
Second, use state and check it into an S3/R2 bucket. Keep your TF scripts in Git, and check them in too after changes. Make a checklist of what steps you take each time you modify a resource (first in the script, then in the state/live).
Third, learn the command line tools used to fix horked state deployments. It'll happen from time to time, and there are GOOD tools that already exist to fix issues. Also, the state is JSON, and you _can_ edit it by hand if you need to get something to work.
Remote cloud provider configuration is a complex problem space akin to programming-at-a-distance. It's hard because APIs are trash, APIs go down or flap in the middle of action, and APIs are slow, so debugging is tricky. Early on, a full tear down and rebuild policy helps, but quickly the slow APIs/slow cloud operations make you more reluctant to start from scratch.
Oh, and databases require a completely different management approach.
That said, I still endorse Terraform over Pulumi or the AWS CDK.
Ansible, Chef, Puppet and co existed long before Terraform. If the stateless way would work better, these tools take over the cloud infrastructure space. But they didn't. To me it seems like a good sign that the their approaches didn't fit to infra.
Additionally, one of the biggest terraform selling points in the early days was the change plan feature. The ability to see the entire change that is about to happen as a result of a config change. I don't think it's easy to implement such thing in a stateless system.
Is it possible to create a stateless config for a subset of the resources/providers with a better functionality? Absolutely! Octodns seems like a good example. Another great example would be any of the serverless frameworks out there that do a much better job than Terraform at managing the lifecycle of the functions. But can this approach be applied to every provider and resource?
The tools are not comparable in what they do.
These are in use but far cry away from popularity of tools such as Terraform, Cloudformation and others.
Ansible, Puppet, etc. don’t have intermediate stores of the hosts’ configuration, but then again they are used for different things.
Terraform defines and end-state and uses dependencies between resources to reach that state. Removing a resource from a tf file will remove it from the infra.
Ansible and Puppet are not used anymore because of the shift to cloud native, not because they are stateless.
TF's success, in my opinion, can be boiled down to:
a) Having a state - allowed them to provide declarative config with plan functionality
b) Choosing Go - removed the burden of installing or having to invest in self contained setup like they did with Vagrant & ruby.
I'm not a TF zealot. I hate my life every time I need to use it. And yet I know that there are no better alternatives out there. I'm sure there are some nice things in Pulumi or cdk-tf but I doubt switching to any worth the investment with existing TF project.
I would be glad to try and play with a tf-like project that works without state - "talk is cheap, show me the code".
Something that kubectl is not capable of. The lastAppliedConfig annotation does not help for purging, because once the manifest has been deleted on disk, there is no way of knowing what to delete from the server. The unusable apply --purge flag is the best example of this issue
I think the state mainly exist to know what has been created in the past but since been deleted from manifests and therefore needs to be purged. The caching/performance argument is rather weak, because Terraform refreshes by default anyway before any operation.
Beautiful summary.
For resources with flexible tags, one could easily imagine tags like Kubernetes's:
terraform.io/name
terraform.io/instance
However, for tag-less resources you have no choice but to store state to map real-world IDs with what is in the config.I wish Terraform "tried harder" to avoid state when it can be avoided. Perhaps it could introduce some soft state, where deleted resources are refreshed by looking at tags and not state.
I do everything with Terraform so I'm not super familiar with either of them. But teams are free to choose their poison.
when building https://carvel.dev/kapp (which i think of as "optimized terraform" for k8s) the goal was absolutely to take advantage of those k8s features. we ended up providing two capabilities: direct label (more advanced) and "app name" (more user friendly). from impl standpoint, difference is how much state is maintained.
"kapp deploy -a label:x=y -f ..." allows user to specify label that is applied to all deployed resources and is also used for querying k8s to determine whats out there under given label. invocation is completely stateless since burden of keeping/providing state (in this case the label x=y) is shifted to the user. downside of course is that all apis within k8s need to be iterated over. (side note, fun features like "kapp delete -a label:!x" are free thanks to k8s querying).
"kapp deploy -a my-app -f ..." gives user ability to associate name with uniquely auto-generated label. this case is more stateful than previous but again only label needs to be saved (we use ConfigMap to store that label). if this state is lost, one has to only recover generated label.
imho k8s api structure enables focused tools like kapp to be much much simpler than more generic tool like terraform. as much as i'd like for terraform to keep less state, i totally appreciate its needs to support lowest common denominator feature set.
common discussion topics:
* whats the lowest common denominator for apis that need to be supported
* how much state to store client side vs server side (in the api itself e.g. tags or in "assistive service" e.g. s3 api)
* is it enough to just store resource identifiers vs whole resource content (e.g. can resource content be retrieved at a later point; if content is stored, is it sensitive)
* how easy is it to recover from complete state loss
That deletion and modification problems happens in Ansible and other provisioning tools that relies in idempotency, and that is one of the things that makes Terraform different from them. A stateless Terraform is useless, it's better use other provisioning tools.
My hard line opinion is that if something NEEDS state management to exist and update, it's a pet, treat it like a pet. Don't mix pets with the rest of your automated machinery except to the minimum extent required, when absolutely necessary.
We had to rewrite an Ansible role because a patch level version upgrade on the recommended community package started destroying security groups... Another time we had to roll back our deployment Ansible image and update 30 repos because a minor level version bump in another Ansible recommended package suddenly required log groups to have an expiration value set in AWS or the module fell over with an null reference on the AWS lookup/comparison... So we couldn't use the latest version to fix it.
I've almost never had this happen with any AWS tool... Sure, there's drift possibilities, but those are controllable by mainly not letting humans do things, and not having multiple cooks in the kitchen changing things in automation, which are good ideas for terraform and Ansible to...
I also encourage modular designs, any of which (except cdn/db, and dns related stuff, in my use cases) can be torn down and re-built with only the brief outage nonexistence causes.. we only have done that once in 3 years, and we believe the issue was actually on AWS's internal side anyways.
I've spun up over 260,000 vms over 4-5 years with one cloud formation template and the basic SDK call, and we've never bothered to convert it to another tool because it's never broken.. we occasionally tweak it, use gp3 instead of gp2, etc, but it's never needed us to unexpectedly side track a sprint for 1-3 days
When I've used CDK, I actually strip the boot strap and avoid the functions that add state, and it makes great templates... But as I mentioned before, my use case might be different and less prone to issues.. I'm mostly using ec2 servers, lambda, and the typical supporting services (iam, cloud watch, s3, CloudFront, sqs, etc.)
This is also what you do with Terraform: that's what the Terraform Cloud product is about (or you can just build a CI pipeline in your tool of choice with a blog post's amount of work).
It also sounds like what you're saying is that you can just avoid all these problems by having everything automated from day one, but that's not the reality in any employer I've ever worked for. Unless you're starting a company today and happen to have an experienced infrastructure engineer on staff from day one you're not getting that world.
Terraform state also comes in handy as an audit trail. Through versions of your state file, you can see how your infrastructure changed in the past (You could use S3 versioning or you could have Terraform Cloud manage the state file automatically).
Terraform Cloud turns that concept into a relatively powerful feature. It becomes a compliance record: who changed what and how it changed, along with diff files.
As another commenter pointed out, it really seems like the alternative tool you're using (CloudFormation) has its own concept of state, but it's just hiding it from you. Pulumi also has a concept of state. I haven't really seen an infrastructure automation tool without that concept.
Ultimately, these tools need to have some way to track what resource the code is referencing. Call that metadata or state, you have to track something. And if you come up with something that somehow avoids it, I have to ask why? It seems like trying really hard to make a car without doors that also prevents rain from coming inside: why not just leave the doors on the car?
Certainly, it's hard to start it from day one. But it's not that difficult to move into it. If you have buy-in from your developer/operations team. We started with API and CLI calls for our beta version. Our next migration was to use partial Ansible control and we morph that into almost a monolith because of interconnected pieces that were typically required with this design. But it really didn't need to be a monolith, we just wanted to link things together and building a giant monolith was easier to make those references.
So we then split up into smaller Ansible playbooks, and we did lookups to create the linkage, which roughly broke the monolithic pattern and allowed us to do smaller deploys. But we still ran into breaking changes unexpectedly. So we decided to abandon the months of effort we put into Ansible and we started looking at terraform, because one of their salespeople promised our management team that it is cloud agnostic and we would only have to write once. After a week or two looking into that, we realized that we were basically just going to have to remake Ansible modules and we rarely weren't saving anything by migrating. Granted this was two and a half to three years ago. Things might have changed.
We then switch to SAM, and as we did that we extracted the Ansible side out of our deployment and we started redeploying brand new small SAM stacks, and started treating almost everything of our infrastructure as sheep instead of cattle, we completely redeployed our launch configurations our cluster are lambda functions, basically everything except the database, DNS, and the CDN with every deploy. This basically removed state is being an issue because the state is only needed for the first deploy, and we don't technically change the stacks afterwards since we simply replace them every time. For us, this also meant we could easily test and roll back if needed. Since we don't need to change the state of the stack back to its previous state, we simply changed the pointer to the previous stack, which typically was a DNS state change. But like I said, we focused on our states being only around DNS, etc. Which, almost isn't stateful, because our SAM deploys would insert zero weighted DNS records, and The state management is really just adjusting the weighted values.
I might suggest the Serverless framework (no serverless needed) which lets you write CFTs in most formats you prefer, use variables (including pulling config from S3), provides the canned scripts you otherwise add, and so on.
I wish MS would unbreak it's plugin to Serverless to support full ARM.
As others have written, there is state in the CFT/ARM/etc. services via named deploys. As you note, that allows for the deletion of removed things (including things for which you've specified another name). That's great.
The challenge with Terraform state use is that it is often too tetchy about what it doesn't know about and often reacts by tearing down and rebuilding assets, causing unavailability and data loss. It was even observed doing this for an entire stack due to a transitory network failure at PayPal. I loved --auto-approve because I don't want to reintroduce human failure and friction into the mix but Terraform's models and dependency on "community or when we get to it" means it just isn't safe the way the native provisioning services are. As you note, alignment of incentives and all.
Ansible is stateless because every operation is suppose to be idempotent. Unless your ansible is doing a HTTP PUT request to an API I suspect you're misusing the tool for something it's not meant to do.
State is a good thing with infrastructure and terraform got it right.
Terraform is not perfect, but looking at solutions we have available on the market -> its years ahead of competition.
You can still make changes to resources within stacks outside CF, and Cloudformation can still be very unaware of the drift.
Unless you're arguing semantics to which I'm not interested.
- The desired state in your .tf files
- The actual state in your provider (what you describe)
- The expected state in the state file
This essentially gives you drift detection and delta updates. A simple "terraform refresh" and "terraform plan" read your actual state and does a diff between desired state and actual state. If all you had was the real world state, then the absence of resources gives you zero information. The other way around: planning a change without comparing it against a stored state gives you no way to actually purge the real world state.You could technically argue that everyone should keep their .tf files in VCS/SCM and then have terraform first check the real world against the previous commit before creating a delta based on the current changes, but then you're just moving state to Git which is already a state backend...
The triangle this creates is why terraform generally is better than most other IaC systems which either don't have all three legs (and thus collapses them into a 2-dimensional also-ran tool) or they do but only for one special system (i.e. only AWS/GCP/Azure and no integration with anything else).
Next thing you know someone is coming to advocate against locking and hash comparison...
Edit: the best 'simple' explanation I could come up with is: you can't remove or update what you don't know shouldn't exist anymore. And you can't realistically 'download' the configuration of an entire cloud to 'check' all the tags for state information.
If you mean that they should have tried to keep using tags, not every resource and cloud provider supporting tags pretty much ends the possibility of that.
If you mean they should have gone completely stateless, I would say that state is great for detecting configuration drift. If Terraform tries to create a new resource, is it being created for the first time, or was it manually removed? You can only tell with state. Sure, you can get this information in a number of other ways, but one of the major strengths of Terraform is as an easy drift detection tool.
I get the desire for state not to exist, but in reality it is essential for what Terraform does.
I think whether that's desirable or not is a matter of preference. The nice thing about Terraform is that you don't have to go all in. I don't really want to have to tell it to ignore every resource that I want to manually manage, and I don't want it wasting the extra API calls on trying to find every resource that's not in my configuration every time I want to deploy.
Now, if they provided a separate tool for detecting manually created resources, that would be awesome, but it wouldn't fit into the typical flow of showing drift regularly just before deploying in my opinion. That's more about detecting whether you're about to overwrite something that was manually done, or if you're about to make changes assuming your infrastructure is as you last left it. I don't need to take inventory of every single resource I have every time before deploying.
It could do drift detection for existence rather easily if that was relaxed, but the rule follows the principle of least surprise.
And Pulumi is using Terraform providers under the hood.
Writing infra in imperative style can cause ALOT of unseen issues when developers start adding IFology or Design Patterns..
Pulumi abstractions are updated by a single corpo team.
Ill be EXTREMELY surprised if they wont laaaggg alot more behind when they start supporting more and more platforms.
Its simply a matter of amount of ppl working on the tool.
It’s promise is sustainably same-day parity with platform APIs. It is generated code, so it may not be semantically pleasing, but it should just work. I haven’t spent much time with it to form a nuanced opinion but I do think it’s a novel and reasonable approach.
https://www.pulumi.com/blog/pulumiup-google-native-provider/
I love/hate Terraform. It’s better than any other tool I’ve used for what it does, but the abundance of subtly leaky abstractions is tedious. And then when you mess up your state occasionally, yea that’s super annoying too.
I'm also curious to know how many people in this thread consider themselves developers as opposed to devops/cloud engineers.
The serious problem with CDK is that it's just a wrapper for CloudFormation - so you have all the limitations of a CFn backend with the added complexity of a full general programming language and ecosystem.
While Terraform is far from perfect you can land in a team, look at the terraform repo and immediately know what's going on, you've got very commonly used patterns / templates that can very quickly spin up your platform.
With CDK almost anything you create is a pet, a special snowflake that almost immediately becomes technical debt.
You end up spending a lot of time on maintenance and upkeep, updating libraries for (really bad!) node security patches, you have to test those updates across whatever packages and repos you're using them in, maintain software build/test/secure/deploy CI pipelines.
I've seen CDK at reasonable scale - it's painful with more than a couple of repos, even worse when they're spread across teams and products. With Terraform at least you just update Terraform itself after checking for breaking changes.
I could keep on ranting but I think no matter what I have seen and say - many developers will take a default position to arguing that infra/platform code is always better written in their language than any DSL, wrapper or templating system, I've heard this time and time again and only really agree with it when it's with Lambda only workloads, Often people cite "a real language is so much more powerful" etc... but in reality 99% of the time you just don't need highly complex programming logic when creating reliable and scalable platform using something like Terraform. (Insert "any fool can create something complex" quote here).
The one and only place I think CDK is genuinely a half-decent tool for the job is with Lambda only deployments as long as you keep it pretty lean don't don't over cook it with complexity and keep in mind that it's just CFn in the background so it has the same limitations.
I'm in a "DevOps" / platform engineering / automation role.
</poorly written midnight rant> :)
I’m not as experienced as you, but we have been using CDK for a lambda only system and it has been working well for us. CDK does have some warts, for example you can’t update multiple global secondarily indexes on a dynamodb table despite cloudformation having the capability.
I think the big advantage of the CDK is that it is more approachable for software engineers who do not have a lot of infrastructure experience. We currently are a dev team who also manage our infra and we do not have dedicated infra staff or devops/sre staff. Writing python code in a declarative style is easier for us devs instead of diving into terraformation.
That did require making the initial decision that octoDNS would own the zone completely so that it could know it's safe to remove records that aren't configured, but later work with filters and processors does allow softening that requirement.
I've always had similar feelings about Terraform's state, in fact I started in on a prototype of an IaC system specifically to see if it was workable to avoid external state. The answer as far as that POC made it was yes. I was able to create VPCs, subnets, instances, to the point I stopped working on it. It was generally straightforward, but there was a hiccup for things that don't have a place to store metadata of some sort and the biggest issue was knowing when they were "owned" by the IaC.
I think some of the other issues that things would eventually run into would be around finding out what exists (again to be able to delete things.) The system would essentially have to list/iterate every object of every possible type in order to decide whether or not to delete them. Similar to octoDNS this could be simplified by making the assumption that anything that exists is managed, but that's not workable unless you're starting greenfield and it would still require calling every possible API to list things.
Anyway, I see why Terraform went the way it did, but I still wish it wasn't so. Thinking about it now makes me want to pick the POC back up...
(maintainer of octoDNS)
We try to make sure EVERYTHING is done through service identities (Azure). No secrets in deploy scripts or resources.
Terraform will happily generate and store passwords and API keys for resources even if we don't want to use them and save the keys to everything in a file on a "random" disk and suddenly we need to lock this file down.
With a statefile we have to worry that secrets gets accidentally stored -- I am sure it id possible to configure it away case for case, but that is a more fragile approach.
With tools without a statefile the whole problem just goes away.
For us this was enough reason to disqualify Terraform.
If you don't have any state, and you have an empty module, did you just create it, or did you just remove all the resources from it? The former requires no action, the latter requires API calls to delete something that I no longer have a record of.
More generally, do I have to completely enumerate the entire state of every service available to my AWS account to determine whether there's something that shouldn't be there vs. the contents of my Terraform modules?
So here's a solution that could satisfy stateless proponents: we don't keep state for anything that we deploy and instead we should store state for everything else so we will know what not to remove /s
But seriously, as things are right now, state is necessary, it could be possible to get rid of it, but that would require cloud providers to designate their infrastructure that way.
And if they do that, in practice (because components building the cloud service likely are stateful) we would end up with a system with declarative language which in reality just abstracts away the state from us. Kind of what AWS CloudFormation essentially is.
Use some tag unique to the project and tag all resources on creation. When deleting resources delete everything not in your stack that has the tag. You can find these resources using the provider's tagging APIs, e.g. https://aws.amazon.com/blogs/aws/new-aws-resource-tagging-ap...
what if it’s stateful infrastructure like a db, and contains important state? is it already backed up? would it be expensive to restore? what else aren’t we thinking of.
even in a stateful model, deletion of top level, especially stateful, infrastructure should ALWAYS happen later, and maybe not happen at all.
creation is easy. deletion is it depends.
for ec2 i’m not sure i’d want unique names per instance. typically i prefer groups of instances with the same name.
Once you've fully adopted, then you can flip a flag to "I own the whole world" mode, where the default is to delete anything not found in the canonical configuration.
The solution for renames, both before and after full adoption, is similar - if you want to rename X to Y, change the canonical definition to Y, but add a tag saying "when you look at the current state of reality, you might find a thing called X; that's the old name for this, so rename X instead of creating from scratch".
I think it would sell like hotcakes.
* A table/column schema, which you can automatically synchronize into any DB backend you want via plugins
* Read/write logic in various languages - not an ORM, but a struct that represents a single row and handles the boilerplate.
* Maybe some kinds of richer query/join logic? If you go too far this becomes another ActiveRecord, but I think there's a middle ground.
* Batch pipelines with standard semantics - sort of like a materialized view, but computed via your big-data pipeline of choice rather than in-engine. Imagine that you have table A with 30 fields, table B with 15 fields, and you want to generate a downstream table with all 45. I think it's feasible to have a composable, declarative syntax that lets you create the 45-column table plus the pipeline that populates it with about 3 lines of configuration. Hard to turn into a product, because so much of that pipeline will depend on org-specific tech stack choices, but at the limit the "data platform engineer" could be entirely automated out of 80% of their job (and therefore be able to focus on more interesting things).
[1] https://cloud.google.com/blog/products/bigquery/inside-capac...
current_tables {
TableA {
Column1[string]
Column2[bool]
}
}
removed_tables: ["TableOld", "AnotherOldTable", ...]
Depending on your ergonomic preferences, you could also accomplish that by keeping the old table configs and adding an "is_deleted" flag. And once you've done one deploy, you can delete all the old tombstoned configs.The whole idea of TF is to not have to declare an absent resource for it to be destroyed, because the declarative approach already have the desired state.
And this is only during incremental adoption, where you'd soft-delete resources in your config by switching them to tombstones instead of removing them entirely, and adding tombstones for legacy unmanaged resources you want to remove (which builds up a nice history of those removals).
Once that's done, switch to "omnipresent" mode, delete all the tombstones, and never worry about them or state again.
I’m practice this results in a mess. When using Puppet or Ansible this method is also required. And it often leads to lots of code duplication or forgotten entries. I’d rather mess manually with state once in a while then constantly having to manages tombstones.
That may work for a database, but for a cloud provider for example(which i dare to say, its the main usage), it's unlikely to work, because of the amount of requests needed to check all possible resources and ids or tags. Also those applies usually happens multiple times in a day issued by different users/ci automations, so there's not enough api quota to cover them all and each plan would take forever to finish depending on your organization size.
But we also need a way to fiddle with the dev environment, while keeping track of everything that happened during development phase and making sure that it can be applied to the production DB in a single command with little room for errors.
Having a git repo in sync with the DB and a full history of changes in commit log helps a lot.
"terraform delete" only deletes what is in your config. "terraform apply" only creates and/or updates what is in your config. If you remove a resource from your config after applying it, it's your responsibility to manually delete it.
You'll either have to:
- Move the state management elsewhere, and invoke different commands depends on what and how resources are changed. This will make automation difficult, and doesn't solve the problem.
- Make Terraform assume that everything it sees is under its management, deleting everything not defined in the current configuration. This will make Terraform hard to adopt in an environment with existing infrastructure.
Adding a tombstone for deletion, or a formerly-known-as tag for renames, is only "state" in the way that reserved tag numbers in protocol buffers are "state". It is a little annoying to have to do, and it creates clutter that you eventually have to go back and clean up, but neither of those is a dealbreaker, and in the meantime it solves the second-biggest problem with Terraform, which is the inscrutability of what it actually thinks it's doing when it comes up with a plan you don't expect. (The biggest problem is how ridiculously inexpressive HCL is as a language.)
Under the hood, the terraform file is just JSON, so sed, grep, jq etc can be used to manipulate it (as well as any other tools you'd care to write).
But anyway, agreed, I think the stateful status quo is the way to go.
If you know that TF changes are guaranteed to be deployed within X days of writing them (e.g., with something like Atlantis, or even a weekly deployment schedule), then you can put a date in the comment of when the tombstone was added, and clean it up either automatically after X days or occasionally in a semi-automated sweep.
Not really. In every other system you can remove tombstones after a while.
Two Users create exactly the same resource with the same tags.
Which one should be removed by Terraform?
Now either way lets ignore that.
You want to refresh infrastructure to know what to do. Without the state you have to go through EVERY API CALL on every service even those you did not create to be able to determine the whole state of the infrastructure which would be super super long action.
Without dependencies you would also have to maintain and build dependency tree EVERY TIME you would try to apply infra.
And a good answer might be "all providers don't have that capability" and/or "providers can't efficiently answer questions about that, such as 'find me all things with this configuration tag'.
In your example, those two users wouldn't have the same tags, because you'd arrange it so that they didn't - either by user or a resource grouping based on the configuration itself. This is the choice made by some other tooling, for better or worse.
Route53 records, ECR repositories, Cloudwatch Alarms, IAM user groups, EC2 Launch configuration
It’s a real journey and not for people under short-term time pressure, but of all the pain-in-the-ass things I’ve learned in computing over the years it’s paid me back well-above average.
I don't understand where the 99% comes from. Maybe I'm just using the wrong services, but it seems more like 50% to me, anecdotally. Maybe 90% if you include hacks like storing "tags" in other fields like the description or resource name.
If it's truly 99% for you, it seems like it wouldn't be too hard to make a tool that generates an ephemeral TF state file from the services on the fly. Win-win? Or you'd run into the next problem, collecting all the state all the time would probably be really slow in big projects.
Pros:
- it can manage resources where not all information can be queried
- reading state is almost always significantly faster
- it can identify deleted resources
Cons:
- state can go out of sync
- every resource type has to implement state well while tracking changing features, so it’s fragile
- fixing broken state can be a PITA
Without making these trade-offs, terraform would only be able to support a smaller subset of the providers it currently does and perform poorly.
Having said that I don't mind state at all! There is always a storage layer and a schema somewhere storing some kind of information.
But of course, you could just write the code to explicitly include that pre-existing resource, rather than have to run a command to tell a state file to explicitly include it.
The problem with Terraform isn't that it keeps its own state file. The problem is that Terraform is just really dumb, regardless of the state file. It doesn't know how to auto-import existing resources. It doesn't know how to overwrite existing resources. It doesn't know how to detect existing resources and incorporate them into its plan. It doesn't know how to delete resources that were unexpected. It's so stupid that it basically just gives up any time any complication happens.
Puppet and Ansible aren't that stupid. They both will make a best effort to deal with existing resources and work around problems. It's not the presence or absence of an external state database that make those tools "more advanced", it's simple logic.
You are offerring an oversimplified view of cloud deployments.
How do you propose terraform is supposed to discern between existing resources that belong to you and the ones that belong to another colleague/project? That is the first problem and is an important one as many organisations have shared accounts where multiple distinct deployments coexist.
The second problem concerns the nature of deployments: in event-driven architectures, a cloud component firing an event and the event processor are decoupled from each other. For example, an S3 bucket can fire an S3 object creation event that can be funnelled into an EventBridge, however the event processor consuming the event off the EventBridge is an entirely distinct entity and has no explicit dependency on the event source (i.e. the S3 bucket in this example). S3 buckets firing events can be later replaced with somethingn else, but the event processor will remain unaware of the change and – from the deployment perspective – it does not even need to be redeployed; terraform has no way of knowing such things.
terraform can't and is not supposed to know such things as its purpose is threefold:
1. Declarative approach to the resource management – express the intentions, or the separation between «what» and «how» parts. terraform code is about «what needs to be created» but not «how it needs to created». terraform providers take over and handle the «how» part [mostly] transparently.
2. Dependency graph management – dependency management is hard.
3. Dependency tracking – it is akin to make/Makefile with the terraform state persisted in a dedicated space (locally or in a S3 bucket) as opposed to make/Makefile that makes use of the local file system and Makefile targets.
> It doesn't know how to overwrite existing resources.What do you mean by overwriting existing resources? terraform can easily overwrite configuration of and metadata (such as tags) for the existing resources, but a complete overwrite of an existing resource at least in AWS – that is not even possible to the best of my knowledge.
All of them "belong to you" if you have IAM access to the resource. It's trivial to restrict access to resources in AWS if you shouldn't be touching something. The fact that so many orgs don't use IAM properly is not a reason to make a tool unnecessarily difficult to use.
So the question isn't if it "belongs to you", the question is whether you as the user want to do something with that particular resource that you already have permissions to modify. This could be accomplished at least a half dozen ways:
1. Is there a record in the state file of Terraform creating it? No? Then it's somebody else's resource
2. Ask the user
3. Command-line options
4. The code / DSL
5. Tags
6. Naming convention
But it does none of these things. It just craps out with a generic error and you're left to manually fuck around with the state halfway through it already started applying changes in production.> 1. Declarative approach to the resource management
So it's declarative. Who cares? In the real world, there are resources that Terraform didn't create that you still have to deal with. There are also resources that Terraform did create, but you lost the state file, or you moved some modules, or changed the Terraform version by mistake and can't reverse your state file version, or a million other things. "Declarative" is not some magical word that means "the real world no longer applies".
> 2. Dependency graph management – dependency management is hard.
It's not that hard. If you already have a DAG and algorithms that can manipulate objects in it, there's only like 2 or 3 more functions you need to manipulate dependencies in a graph. This is really basic CS stuff. Dependency tracking is the same thing.
> What do you mean by overwriting existing resources?
I mean if there's an AWS record that Terraform wants to create and you didn't import it, Terraform fails. This is stupid, because 1) it's not even checking if the resource already exists (it has the data providers to do this), and 2) it's making me do something manually that it could do by itself with a flag or a config or DSL entry or a prompt.
The UX is fucking ridiculous. It's like whoever wrote Terraform has never actually used it, or just doesn't give a shit about blowing up production, or getting anything done in a reasonable amount of time. Classic "humans exist to serve machines", "opinionated" tech hipster bullshit. Must be the same person who did Git's UX. And impossibly, people defend it, as if technology is supposed to make our lives harder.
Let's take the most basic example: auto-generated ids. Many resources in AWS, GCS, etc have auto generated ids (just use tags you say, but many don't have tags or tags are used as part of some other system). Now, when terraform creates that resource you have to modify the config to contain the id. But if you have any sense terraform runs as part of a CI system that lets others review your code before merging, deploy to staging, etc.
So now does the terraform process need to make an automatic git push? What if there's a conflict? Does it make a PR that has to be manually merged? All of this is much more complicated than just having one JSON file in S3.
I have actually managed resources with Ansible where you have this problem and it's worse. And this is just _one_ thing.
Is Terraform's state story perfect? No. There are definitely annoyances, and one thing I'd love to see is a way to declaratively handle imports, renames, etc. when you need to, but it's better than the alternative.
Terraform isn’t about cloud apis only. Terraform allows me managing keycloack realms, ssh keys, cloud resources, postgres databases, git repos, imap accounts, and so on.
The state is there to be treated as the source of truth. It gives an answer to „do I have what I want to have”. With that state, it is possible to cross reference various resource types without having to load the curent state on every run. I’m surprised that the author did not see that as a performance issue.
Imagine that you have to load all route53 state, all buckets, ec2 instances, iam roles, … on every execution, and imagine you have 200+ machines… That’s what we used to do with puppet and chef, no?
Turned out that always fetching the view of the world from an api is pretty expensive and quickly exhausts api rate limits.
It’s a pretty weak article without any effort to suggest how could it work without a state.
Previously we have used Ansible for the same thing, in a largeish environment, in production. It had the obvious benefit of declarative and stateless.
The other comments here seem to focus on how that works badly together with manual state changes made externally from the system. The answer is that it requires another way of working where the state is the git repo. The question is not what to do when someone spawns extra test nodes, but why they did not do it in a version controlled manner.
Perhaps as a reaction to the discipline required, something about Kubernetes attracted a lot of people used to manipulating state by manual interaction. Every installation I have seen has a web interface in use, whereas not as many have in the Ansible world. I fully expect this pendulum to swing back and become more declarative and version controlled again.
The more i use tf, the more i think it would be better to remove _all_ dynamic features and use a real language to generate tf configs as flat, static files.
Unfortunately, it is not stable yet... but we all used Terraform <1.0 for years, and I'd say the experience was never that bad.
We want infrastructure automation to be boring and just work.
The risk with general purpose programming languages is that people will always find a way to outsmart themselves. Yes, sure, you can use the testing tool chain of the language of your choosing. But it's not like we have figured out to write software without bugs, despite all the awesomeness of modern languages.
But to explain where I'm coming fromm, this is something I ran into recently:
this breaks and returns null if http is not set: lookup(each.value.http, "port", 8080)
So I need to do this to make it work: coalesce(lookup(each.value.http, "port"), 8080)
Not a huge deal, but just one of these many things over the years where I think don't reinvent the wheel.
I wish Terraform could automatically detect the change and convert it into code.
state or idempotency, pick one.
hello? how is this even a question? i have a hard time believing that idempotent infrastructure mutations are not the right move almost always.
shameless plug[1], i’ve been exploring an aws specific approach to infrastructure that is stateless and idempotent for exactly these reason.
slow, finicky, stateful deploys are about as awful as it gets. add a pinch of lowest common denominator among all providers, and that’s a tough pill to swallow.
there’s got to be another way, aws should be fun!
If you want stateless, then you can use Ansible and use their providers. Enjoy spawning new instances everytime you change your infrastructure, rather than having existing ones change.
Describe loops and other properties in Ansible and have it build out a complex Cloudformation template for it to deploy. Use an Ansible variable for your stack name so you update/delete an existing stack and your good to go. Could even break it down to environments so dev = smaller instances vs stage/prod etc.
As every application developer knows, duplication of state is a primary source of bugs. To combat this, React/Flux type of architectures became extremely popular where state flows in one direction only. They dictate that, no you can't just cheat a little and use jQuery to modify some element, it _will_ get bulldozed on the next render. And a lot of Terraform headaches do come from this analogous reconciliation of what really is (our cloud env), what we want (our TF code) and this intermediate state of what TF thinks the cloud state is.
So, by saying that you cannot have resources outside those defined by TF there is actually a massive simplification with far reaching consequences possible.
How I imagine the experience would be:
- You could say that your Dev env is a shitshow and always will be, of manually created resources and only partially TFed. But your Production and Staging envs are opted in to "strict mode". This means that if there is a conflict you do only have two options: import the offender or destroy it, and the critical mindset change is that this is a good thing and will save us a lot of tears later on.
- Caching is an orthogonal concern. Terraform mixes these two together to its detriment, but the nature of a cache is such that you can safely blow it away and perhaps the next reconciliation will be slow, but it will be accurate. I also don't believe it would actually be that slow, tools like Cloudcraft map the entire metadata of your account in seconds.
- I find the excuse that some resources don't support tags intellectually lazy. Of the top of my head thinking about it for a minute, could you tag a parent resource with the child metadata you need? E.g. individual DNS records don't have tags, OK tag the Zone with childA=value. Same thing with tag length limits, you can work around it, concatenate values or whatever. However, in a truly strict mode you wouldn't even need metadata in tags because the TF code describes the entire target environment.
I hope Terraform would entertain such a strict stateless mode. Unfortunately it will probably take another tool, because the problem is not so much technical as it is an entire mindset change.
But a generic “none” backend as described in the OP would simply be impossible. The support for diffing desired vs actual state must be implemented in the resource provider, which in turn need to be supported by the cloud API. Ubiquitous labels, namespaces and global query performance seem to be the primary blockers today, judging by most other comments here.
Interesting thought if someone would attempt to make such a provider. Also interesting to look at existing providers if they avoid state internally when possible or just use it because it is available.
Looking at the challenges with kubectl apply --purge, already having all enablers laid out, it would require big effort.
Define everything in like cdk (... I use ruby to generate the tf.json), generate the code, import everything you can without error, and apply the rest.
Performance will be _bad_ but that will completely eliminate state problems.
CF tries to create as many resources in parallel as possible based on the dependency graph it creates.
The only way around it was making each Parameter resource dependent on the other (using DependsOn) to force the resources to be created sequentially.
Before the CF vs TF holy wars began, this is an API limitation that you would hit regardless of your IAC.
cdktf is great, but I would also rather do away with the state. I’ve gotten into too many problems that were only resolved by deleting everything and starting over.
I use this today to keep very “wide” modules synced
But, it is undoubtedly a pain in the arse in practice - invariably someone else is doing a change on a branch but has already applied it to the development infra using the shared state bucket, which then pollutes it for everyone else as you don't have those changes so your applys will now want to undo them.
As a first timer with this stuff, I think a key lesson learned is to balkanise your terraform quite heavily - we have a tfstate per environment but that is nowhere near granular enough, slicing it into smaller pieces would obviate many of those 'pollution' problems.
Just another layer of abstraction...
I wish either of them had a more graceful plugin story, though.
I've tried terraformer[1], but I can't tell how well maintained that is (it failed to get my creds, I had to modify the code to fix it, and then it crashed with an obscure error).
Anyone have a good approach?
The author compares Terraform to Ansible and Puppet, but these are not analogous tools. If you compare to Pulumi, you'll notice it also has state.
Keeping track of what is live and what is expected to be live is much of the utility of Terraform.
If the author is getting some Terraform state pain (e.g., from drift caused by a team making changes to infrastructure and forgetting to represent those changes in code), they should try using Terraform Cloud or building a CI pipeline for their infrastructure.
That needs some manual fixes i.e import state, killing resources etc but not all devs has access rights or knowledge to do this.
Not true since Terraform 1.1 introduced moved blocks : https://www.terraform.io/language/modules/develop/refactorin...
Wouldn't take much to hack something together to test this out, either... parse the TF for resources, lookup what is used for their IDs, run TF import with discovered IDs from the service provider and then your local state is up to date, run your plan / apply and blow away the state when you finish. But this is super gross IMO :)
then you can KNOW that no other random infrastructure should exist in an account.
terraform definitely does help coordinate the bunk beds of room mates. wouldn’t want them to accidentally discard each other’s pillows as they move in and out of the shared space.
separate billing is a nice bonus.
As soon as you get to multiple teams managing (say) AWS infra, you can't infer that a resource present in infra but not in tf file means a resource should be deleted.
The actual infrastructure and its configuration is the state. It does a diff to see what needs changing.
So, there is definitely still state, it's just stored centrally.
[0]: https://docs.microsoft.com/en-us/azure/azure-resource-manage...
What are tags, if not state?
If the resource is prefixed with "tf-" but missing in the terraform config, delete it.
1. Everything not declared is unmanaged. Deleting is done by a "deleted" attribute of some sort. The resource is declared as deleted.
2. Everything unmanaged is declared as unmanaged. The resource is declared as unmanaged and TF would ignore it.
Either of those would work without extra state.
Imagine you go down the ansible route and you write idempotent ansible code than you could argue: “See. I don’t need state. My code is idempotent. I just use this playbook to apply”
Now think of deleting resources.
You could have a delete playbook maybe. Then how would choose whether to create or delete stuff? Maybe a colleague gives you a ticket: “please delete”
Now there are various scenarios:
You do your thing (on your laptop) and tell everybody else in your team to not run the CreatePlaybook as you have to run the DeletePlaybook first for this one machine. Afterwards you delete the actual machine/resources from your ansible repository and tell the team: “please use the newest main branch”.
So: Here is your equivalent to terraform state, the coordination effort on your side since you are the only person who currently “knows” what’s going on with deletion/applying.
So your next idea is: “no problem, I make a Pipelines which runs the playbook on your behalf”. And the pipeline will update maybe the Git repo accordingly in some way - after the run (since ansible needs to know in the run, what to delete).
Everybody can see the pipeline . The pipeline will ensure you sequentially apply the playbooks to coordinate with your colleagues.
Next problem: how does the pipeline know which playbook and resources to run on?
You create a selection box for your hosts and for the playbook yaml to trigger the pipeline.
=> there is your state. Your ticket information are transferred to your manual labor to fill in the correct items in your selection box. To reestablish want went on you now have to look in the ansible code and the pipeline Paramus and pipeline logs.
More examples are: you only want to update some stuff with your ansible playbook. Therefore you might introduce tags on the resources and the playbook knows how to handle those tags. The extreme case might be:
TTag:state:present, tag:state:absent.
Then you can run a single playbook which can call the deletion and installation playbook for you and everybody is happy that you have everything visible in Git.
Problem here: your 2 step process to decommission things from git. First a commit which sets state:absent. Then run the pipeline and then another git commit to delete the code.
So what I am saying is: You can do all of this with ansible for sure. But you will have state somewhere:
In a Ticket, a Pipeline log, in git, in a wiki
I am not saying absible should not be used. It makes sense to configure things. (I personally would wrap a terraform hull around my absinble code and call it, just to have terraform handle the locking for my playbooks)
But just watch out for the hidden state in your workflows and better make it explicit. This is why people love GitOps for traceability.
You can all do this by hand and document your process in a Wiki for your colleagues so that they known what the “tag:state:absent” means for them. Or you can rely on somebody who has done this for you already and maintains documentation and what not.
What I could never wrap my head around was why the heck the tools had to expose so much complexity. I want the description of my infrastructure to be a single, or set of text files, written preferably in JSON or YAML like everything else, and I want to run a command that behaves in the same way as GNU make.
That's CloudFormation, there is a reason that people migrated away from it.
Almost all Terraform providers are ported to Pulumi or have Pulumi-native alternatives.
It's always seemed to me that one of the virtues of Terraform was that the language was just declarative, and rich enough to express what you needed without becoming a Turing tarpit.