Terraform is not the golden hammer
hub.qovery.com
hub.qovery.com
Spend enough time playing with it and understanding it, you'll end up like me thinking about all the shit you configure left and right such as hooking up Stripe's secret keys, Google Analytics and the webmaster console, and just about everything else we configure via web interfaces, and you'll think:
Why can't we use Terraform for this as well? Manage these SaaS products the same way we manage the rest of our cloud, test and audit changes, automatically roll secrets and update anything that needs updating the moment you change a setting.
Ah well. Not enough APIs out there. And it's difficult to write and maintain terraform plugins for these throwaway cases especially if they are going to use private APIs. Anyone know if Pulumi plugins are easier to write?
But like you say - now I've done that, I want to do it for every UI that I'm forced to log in to!
We've gotten away without that because the APIs we use don't contain any secrets. (The auth token for making the requests is just an environment variable.)
“What do you mean run an installer and update these files…”
If there’s ever a gap in what Terraform offers you can pretty easily fill it with Ansible.
I find it hard to figure out how to use GitOps with Ansible? How do you make a PullRequest which indicates that something should get deleted? You still would have to keep around an ansible playbook for the stuff you want to delete.
Infrastructure, all the monitoring as well as all the on-call rotation configurations (and anything else that is in that loop) should all be in code, and all changes should be reviewed the same way as application code does. If it doesn't, you can't really trust you're gonna be alerted properly when things start breaking.
I wish I could use it for personal things too, I'd rather have my bank account settings, my government tax information, yada yada in a personal terraform repository for example. Change of address? Commit a a change, check if the plan is good and apply to change it everywhere. Though having lots of experience with Terraform I can only imagine what the equivalent of trying to delete an S3 bucket that still has data in it is for a bank account.
Why can't everything have nice public APIs! And, while I'm at it, some sort of all-encompassing ticketing system, hell even if it were Jira. 'Pothole', assign local council. Blocked on '2021 roadworks funding increase', backlogged. Assigned to councillor. Won't fix. Ok - maybe I'm not making it sound great, but at least you could see some reasoning, and what the blockers are. Follow the chain to work out that 'communal lobby needs repainting' hasn't been acted kn by building management company because, ultimately, of global supply chain disruption and the contractor's supplier's supplier can't get any paint ingredients.
These are not the same as "real" pulumi providers that run across all supported programming languages, but I think they are a good enough fit for the cases you mention.
This is incorrect, and makes me wonder how the author has used terraform. Terraform will certainly detect differences between managed and current state for the resources it manages whenever you do a plan/apply.
The major challenge is that terraform _can only reconcile resources or configuration values it knows about_, and that depends very much on how a particular cloud vendor or terraform provider has modelled resources. I believe the Helm provider is one example where it (at least in the past) haven't had a good way to reconcile state.
Perhaps they are working with certain poorly implemented or buggy providers, and not realizing those providers were doing something different than terraform's properly working behavior that most providers implemented.
Bugs happen, but the first step is agreeing on intended behavior so we know what's a bug!
When working with AWS it will always reapply the terraform configuration.
A great use of this is for account hardening. You can run it daily to make sure it's configured correctly.
In terms of API use though I suppose that'd be quite expensive to plan - listing every possible AWS resource (in every region!) for example.
> For some resources like RDS or EKS, it won't check if the resource already exists or not. So if it's missing, nothing is going to happen as it's marked are deployed in the tfstate file
is also wrong.
It seems like the root of the problem is that the author wants to use Terraform to manage their AWS state, but also wants to use the web console to directly change things, so Terraform gets out of sync. Terraform has a command to handle this - https://www.terraform.io/docs/cli/commands/refresh.html
The key differences between the controller approach and the IaC approach are, I think, lots of little processes continuously reconciling state for all resources of a given type (many small loops that touch all resources of a given type on the entire cluster) versus a one-off process that tries to touch just the resources it cares about and if it fails it just gives up.
One thing Kubernetes definitely improves upon Terraform is that Kubernetes uses a YAML “assembly language” for its infra as code, but that YAML could be generated by a real programming language. Terraform expects you to write HCL, which is an accidentally re-invented programming language (every IaC tool provider thought static configs like YAML would suffice as a human interface, but as they gradually realized the need for more dynamism, Terraform and others would bolt on one dynamic feature after another until they had a slow, unfamiliar, and counterintuitive programming language). Terraform has a CDK that allows writing in other languages, but I’m skeptical that it liberated you from Terraform’s model of the world (e.g., if I rename a variable in CDK, does it try to destroy and recreate the underlying resource as with Terraform?). I’m also concerned that rather than allowing us to generate YAML in the obvious way, it will require bizarre inheritance patterns like the AWS CDK. I would be curious to hear from folks who have used the CDK.
I constantly see soo many people step on the same rake, it's incredible. Tools like Tilt let you use python, it's a much more sensible approach.
<https://crossplane.io> just graduated to CNCF Incubation and each of the cloud providers are working on K8s controllers and code generators (like Amazon Controllers for Kubernetes, Google Config Connector, and the Azure service operator).
I prefer CUE to JSON for TF and many other tools now
> Then we transitioned to an actor-based model where each resource was almost an actor, and there was a message-passing interface between them.
> This allowed the system to be highly concurrent the way Terraform is today, but also confusing for users to deal with and very difficult to build a programming model around, because the ordering of execution was so random and everything was happening concurrently.
https://www.hashicorp.com/resources/terraform-fireside-chat-...
They may still be right. Kubernetes' approach may seem more attractive but terraform is far more pragmatic in its design.
When you write a provider, you spend half your time converting data structures from the HCL submitted to your provider into the JSON your target service inevitably expects, then you spend half your time converting the JSON your target service inevitably returns into HCL for Terraform to consume, and then you spend another half of your time fixing bugs and polishing.
It's okay when you're building simple providers, but anything reasonably complex becomes unwieldy. I had a go at building some providers for AWS services that were not supported by Terraform or CloudFormation... and I just retreated to cheesy Lambda custom resources for CloudFormation.
The typical example is to start from there:
https://docs.aws.amazon.com/AWSEC2/latest/APIReference/API_V...
and arrive here:
https://registry.terraform.io/providers/hashicorp/aws/latest...
If you review it carefully, it is apparent how much coding effort and many moving parts were used to perform a transformation which seems disproportionately primitive.
An example is the fact that `for_each` is not supported on providers [1], an issue with 230 likes which has not been solved since January 2019. This had me resort to a Python script which generates a `.tf.json` file, definitely not ideal. Infrastructure as code sounds great, but in practice it's closer to "infrastructure as a non-standard markup language".
I think Terraform is quite possibly the best tool available, but there are clear flaws with both the model and the implementation. I think if I were to make a Terraform v2 I would make `plan` completely pure. This would avoid the provider issues, make validation and testing in CI easier and a whole bunch of other benefits. Of course there are downsides. For example EC2 instance IDs are random so you can't just include them in your pure plan. You would need some type of placeholder that is used for evaluation. This does cause some issues as it limits the operations that you can do with that value (so you can't pick the instance size based on the random instance ID) but overall I don't think it would be a major issue if the final substitution was handled well by the framework.
But this wasn’t just Terraform! The entire industry did this too. CloudFormation began as simple JSON, but over time they allowed you to encode the abstract syntax tree of a shitty programming language in your YAML, and CloudFormation would interpret it. However stupid that may sound, in the Kubernetes world, we have Helm which lets you generate YAML with text templates which is honestly the dumbest idea in the world (imagine a compiler that generates syntactically invalid machine code if the input program has an extra white space character).
Of course in all of these cases the answer is staring us in the face: use a static language (YAML, JSON, etc) to describe the desired state, and use a higher level language (like Python or Starlark or Dhall or etc) to generate that static desired state description. The only thing Terraform (or any IaC tool) should care about is the YAML description. That it is generated from Starlark or TypeScript is just an implementation detail.
Instead of that, though, we get CDKs which are so close, but admittedly I haven’t used them in anger yet.
It is however a great improvement over the previous ways of doing things, and probably the best out of the current similar alternatives out there (you might mention Pulumi as a strong contender, especially for AWS glue writing).
And - though as per the disclaimer, I may be biased - until a better tool comes up, I'd advise looking for specialized IaC CI/CD tools to ease your path with Terraform, like Spacelift[0].
It can help you with orchestrating dependencies among multiple state files; take care of scheduling regular drift detection/reconciliation without going into your way and locking your state; gives you a policy system for making sure preventable mistakes don't happen (i.e. recreating a resource you definitely never want to recreate); manages your credentials depending on whether you just want to run a plan, or apply your changes, and much more.
I can't imagine doing Terraform again without a tool like it.
Disclaimer: Software Engineer at Spacelift. If you want me to expand on the "and much more" part, you can find a demo-scheduling link in my bio!
[0]: https://spacelift.io
> I run once again the “terraform apply” command. But for some reason, Cloudflare API doesn't answer and I got completely stuck there without the possibility to update with Terraform this field because of linked dependencies.
You should be able to circumvent this with a `-target`.
That being said, I know exactly what they're talking about with helm. IME the helm provider was/is a complete mess and gets inconsistent state a lot. Helm specifically I would also keep out of TF until that is fixed, if ever. I haven't had that happen with other providers, though. Perhaps OP was just really unlucky ending up with the odd half-broken AWS module.
If your core competency isn't dependent on your cloud platform tf is a great tool. But using cloud APIs directly was great for us.
Compare to kubectl. Where you can write plugins in bash/shell and mark with execute bit, put it in somewhere in your $PATH as kubectl-blabla and use it as "kubectl blabla".
Also, just as you can write extensions to kubectl, you can write your own provider in Terraform if it does not exists. See https://registry.terraform.io/modules/waveaccounting/chatbot...
Also, Chatbot does not have a public API, that's why, it's only configured via Cloudformation. So the expectation is not fair either.
I've seen Cloudformation getting features years later. i.e
2021 - https://aws.amazon.com/about-aws/whats-new/2021/05/amazon-dy... 2015 - https://aws.amazon.com/about-aws/whats-new/2015/07/amazon-dy...
If you can configure something via CloudFormation you can integrate it via Terraform et al also, since they have resources representing CloudFormation stacks.
However, if AWS have not published metadata for a given service to be used across their various SDKs, it’s hard to take that service particularly seriously, so I’m not sure I’d bother with this.
This is fine, I've done it extensively myself for some of the bleeding-edge cloud stuff, but the importance of things like tracking state, managing hierarchical resource dependencies, or retry/back-off logic shouldn't be tossed aside simply because there are gaps in what's available in the Terraform providers. Especially where change management is important (basically any enterprise company).
I'd caution others reading this against abandoning something altogether and writing bespoke IaC tooling simply because the stable approach doesn't cover every (bleeding) edge case.
You'll spend a lot of time reinventing the wheel, and while it's fine for certain situations (like when you only care about desired state, not known state, for instance), you'll move faster (and likely safer) by sticking with tools like Terraform for the bulk of your infra, and augmenting here there with cloud APIs/SDKs when needed.
This makes your own effort for customisation minimal, keeps your knowledge portable and because your added features can be separated in to different files and the provider API is stable you can also easily backport/fast-forward new changes.
With fewer LOC I'm not sure what you mean, the provider code is pretty small, smaller than custom ansible, salt or puppet modules. Smaller than CDK and Pulumni as well. Sure, you'll have to write Go, but that's about the only hurdle.
Everyone doing a round of NIH for internal tooling is ultimately not making the tide rise.
Edit: don't get me wrong here, writing internal tools to do a job the right way for the right needs isn't "invalid" or something like that, but people often dismiss the rest of the lifecycle of knowledge and maintenance when making something completely custom.
I have issues with all the providers that make no sense like application configuration providers or all the flavors of kubectl providers..
Those are often very low quality and have various issues dedicated solutions don’t have.
An example could be helm and helm_provider. The former just works, with the latter Im constantly running into weird bugs that break terraform state..
In the sense that whenever there’s a new API or service available in any of public clouds and their official SDKs there will always be a delay before this new service/feature/API will become available in terraform.
First time I encountered it with GKE private clusters 3 or 4 years ago. Now it is AWS Keyspaces.
The second biggest reason is whenever you have a requirement for a hybrid or multicloud then well you are left with rigidity if HCL. It is probably doable but for what sake?
Solution: get a real language, write a STATELESS configuration management(IaC) system for your own needs and maintain it. The majority of public and private cloud providers ship SDKs in most popular languages that will help you build your own software solution and reduce your dependence on a third party which I would put under progressing operational risk category. Yaml/json/cue/toml for end user configs would suffice.
Example: for one of my previous projects were built a tool for a hybrid AWS-openstack setup, and were managing a dozen of busy environments.
Even makefiles are pretty straightforward, though you really want operations to only trigger when checksums differ -- timestamps result in a lot of redundant operations. As long as everything is idempotent, it's pretty straightforward.
My experience has been the exact opposite. Usually Terraform offers support for cloud services long before the vendor provides an SDK or supports it with their own offering (e.g. Cloudformation). There are still dozens of AWS services, for example that have no CF support offered by AWS.
tfstate files can be painful to manage, we had a lot of trouble with them at Capital One but mostly because:
1. People would modify state outside of TF which you should avoid.
2. People didn't architect their apps well which led to long lived infra. TF works best with cattle like infra.
terraform import feels very much like an afterthought which is why projects like terraformer exist.
For the cases where the native integration with the cloud provider does not provide the exact parameters you look for, it provides the alternative to integrate with terraform in the background, while making it transparrent to you.
I'm thinking of designing a series of tools to replace Terraform. The idea would be to break down how modern cloud environments are managed into a couple concepts, and then make a variety of tools that work within those concepts together, so that it's easy to expand and modify the way you use them for your use case. This would enable things like tailoring the use of the tools to a particular deployment strategy, or adding custom business logic, or replacing individual functionality, without being tied to one tool, language, etc.
Huh, is that true?
I'm just getting started with terraform but I assumed that was the idea of terraform (where it didn't happen would be a bug), and I think I had seen it happening for the few basic resources I have started out with (S3, cloudfront).
If the state doesn't match the actual configuration of S3, terraform notices, and the plan is to make it so. No? Am I confused and it hasn't been doing this?
Or is this is inconsistent, true of some resources and not of others? That seems surprising. What's the idea?
All that a plan does is evaluate what is going to change in the current Terraform state by performing a dry-run of the Terraform code that you have supplied.
If you would actually like to make changes to the Terraform state based on what the Terraform code evaluated then you run a Terraform apply - which will, for the resources deployed via Terraform, update the configurations themselves and update the Terraform state by using the Terraform code as the instruction set.
You can actually see this in action with plan and apply as the output will show you +,-, and ~ where ~ is settings that are going to change but are not new configurations or configuration to be removed.
Edit: Learned from some other comments that Terraform has a 'refresh' command that will take deploy+n time configurations done outside of Terraform and sync those configurations with the state. This might be what you ideally are looking for after deployments?
I thought I had seen terraform correcting it (to match what terraform thinks it should be) in some cases.
OP seems to suggest that in some cases it does and in other cases it doens't. I am surprised if that is inconsistent and unpredictable, and would have expected terraform to (modulo bugs) either always or never do that. And am wondering what terraform's intent is with that.
'terraform refresh' may be what you are looking for. This will update the state to match current configurations that may have been done outside of Terraform.
> You shouldn't typically need to use this command, because Terraform automatically performs the same refreshing actions as a part of creating a plan in both the terraform plan and terraform apply commands.
However, snom380 says this is what terraform is intended to do, and does with properly implemented providers, which makes sense to me.
I'm not sure if you are talking about the same things. We -- me, snom380, and the original quote I made from OP -- are talking about what happens when external actually existing live resources have diverged from terraform's state. I understand what you said that under a "perfect" situation, this would not happen. But it does sometimes for various reasons, what the original quote from OP is talking about and what I'm talking about is what happens when it does. I think maybe you are talking about something else.
I can only assume that the author has used some provider that hasn't implemented this properly (helm, I believe, is one example), or that they've run into one of the cases where terraform treats configuration/attachments to a resource as a separate resource (e.g. IAM role vs attached/inline IAM policies).
Terraform and yaml as well are so verbose and you have no clue whats going on from the local side of things.
How do debug your terraform markup locally. My guess is you can't.
That is one of the core features of Terraform? Detecting and fixing drift is useful.
I think people get into trouble with Terraform when they try to use it to do more than provision infrastructure. Things that should probably be part of build scripts, CI/CD pipelines or config management. Terraform isn’t good at those things; but it is very good at provisioning cloud infra in a cloud-agnostic fashion.
Heck, I never get how terraform can be cloud-agnostic... If everybody thinks having a same language (HCL) is equalivent to cloud agnostic, YAML exists...
It is literrally impossible to create a simple VM in 2 different cloud providers without defining them twice with their own specific parameters.
If you use AWS provider, resources start with aws_, if it's GCP it starts with gcp_ and so on. It is not possible to have a "resource vm { name = ... provider = aws }"
> It is not possible to have a "resource vm { name = ... provider = aws }"
It actually is via modules. It’s a lot of work with basically no benefit though, so in practice people don’t do this.
Pulumi is better at specifically this kind of thing however, since you can implement a common interface which can be specialised for each available cloud.
Indeed, terraform is great for exactly the opposite: For making sure both the initial infrastructure and subsequent modifications to it are first proposed, discussed, reviewed in code and applied afterwards instead of anyone with privileges tinkering with settings on an "as needed" basis in whatever console thereby ending up with an infrastructure where you have no idea who turned on the frobnicator, why it was set to 11, and what might be the consequences of changing the setting.
There’s a plug for managing Helm related resources with the author’s own SaaS product at the end, so I’ll file this under “half hearted hit job”.
What we really need is automatic reconciliation. I.e ask the provider what they have and then diff against that.
Or periodically auto-importing.
Are there any good solutions to auto-importing?
1. Inconsistent behaviour between providers. e.g. if a resource has been destroyed since the last `terraform apply`, then some resources/providers would recreate the resource, others wouldn't. (Similarly, there's not a guarantee that the state after running `terraform apply` matches up with what's there, if the provider is happy with its state file).
2. The dependencies of already-applied resources can block `terraform apply` if the upstream API for these resources suffers an outage.
3. If a `terraform apply` applies some resources before failing, this can result in an inconsistent state. Either the resources need to be deleted, or imported.
I'm not familiar with Pulumi; what aspects of these would Pulumi help with?
I very highly recommend investigating it more and trying it a bit. Terraform isn't mere cloud deployment.
As a small project you can start by deploying an EC2, RDS and some cloudflare records to go with them, all linked together with terraform. This will give you an initial idea of its capabilities.
Please don't read it as an attack, not the intent. Amazing that a HN regular can "learn about Terraform" only in the latter part of 2021!
I assume there is some selective reading and most articles referring to Terraform have a cloud or hashicorp reference in the title. If you don't care about either, you don't read the Terraform things on HN.
We saw a huge improvement in our build times after we started using AWS CDK directly.
Terraform can create cloudformation stacks as well, you just have to write the resources for it. It doesn't really make sense to do that. I also don't know that it's … "faster" in any way; cfn is really slow.
Terraform’s AWS provider calls the APIs directly, whereas CDK generates Cloudformation, an abstraction on top of the AWS APIs. For me, using Terraform was significantly faster than applying the same stack via CDK.
Or do you mean you’re able to iterate faster writing CDK vs TF?