"Cloud" configuration may have been Terraform's proverbial toe in water, but the truly untapped potential lies in the other providers. Anything that can be packaged as a Terraform provider exposing resource abstractions can be easily managed using convenient HCL syntax. This, IMHO, is the unfortunately buried lede.
Plus, the different philosophies of mutable/immutable infrastructure that the different capabilities/limitations in each tool encourage.
Sorry, but I have yet to meet a developer that used Terraform and willingly wants to keep using it after seeing it fail. When your tool cannot keep track of the resources it created or when it gets into situations where it doesn’t allow you to do certain things (like deleting all resources created) and you have to relearn how the underlying cloud works, it’s time to move on.
Do yourselves a favor and use Cloudformation (or your favorite cloud’s equivalent and just move on with your life)
Also, the terraform language, HCL? It's, I guess there's no better way to put this: not good.
Am I misunderstanding the complexity of what Terraform is trying to do? To me, it looks like a bunch of tiny API clients tied together with a topological sort --- in other words, it's just another species of "make". It feels like it carries a whole lot more complexity than that concept warrants, but just enough simplicity to make it not a serious programming language. It's, to me, one of those frustrating uncanny valley systems.
I am prepared to be totally wrong about this.
It's a good idea and not a bad design, but the user experience is pretty bad and nearly all the operational aspect is an afterthought. It was not built to be sold as a product, hence it kind of sucks as a product, but it's fine as a free tool. The fact that it's the best free tool we have for this task speaks volumes about how most companies are deathly afraid to work as a community to build better solutions.
I do think you're selling Terraform short, though. Sure, the core is the toposort-create-things. But it also stores the state of its created things and (crucially) has the ability to diff the actual state of resources against what it thinks they ought to be.
Being able to inspect existing resources and diff state is also what lets it import existing resources, so that they can be Terraform-managed going forward.
Terraform is also capable of determining if its planned changes can be performed in-place or if they require resources to be destroyed and re-created. That boils down to a boolean flag on a field, ultimately, but it's still something a dead-simple make clone probably wouldn't do well.
I've only really used Terraform seriously for AWS, so I'm not sure about the other providers, but the Terraform AWS Provider has an enormous amount of work behind it. Basically every resource API has schema validation written in the AWS provider, and depending on the resource there are often eventual-consistency issues handled by the provider. See for example [0].
In contrast, AWS CloudFormation: can't import existing resources; isn't always sure whether an update will require replacement or not; and, as of ~6 months ago, can detect configuration drift, but not correct it (!). Of course, CloudFormation wins in other areas...
[0]: https://github.com/terraform-providers/terraform-provider-aw...
Also CF is the “native” language of AWS. There are plenty of getting started examples from AWS where they give you the template. Also Elastic Beanstalk extensibility is built on top of CF.
Not to mention Codestar that will set up environments for you for common use cases and exports templates and the lambda environment lets you configure everything from the console, test it and then you can export the CF definition.
I've never had Terraform "forget" resources or disallow me from deleting them (unless I requested that). Maybe those are bugs that've since been fixed?
It sounds like we've had very different experiences with these tools.
Like one sibling comment mentioned, getting support from AWS is nice. You can buy Terraform support, too, but knowing HashiCorp it'd probably cost more than most peoples' AWS bills in their entirety. (This is me being a little cheeky and unfair—HashiCorp handles community support via GitHub Issues for Terraform really well.)
CloudFormation's built straight into AWS, so there's no need to set up state file storage or locking or worry about a state file at all, really. This has its own set of drawbacks, but it's nice for getting started, and in theory it makes CloudFormation more robust out-of-the-box.
CloudFormation StackSets are really nice if you have identical resources that need to be placed in many different regions and accounts. (Example: we use it to place GuardDuty and Config, both regional services, in every region.) With Terraform, this means copy-and-paste, as far as I know, though maybe Terraform 0.12 sets up the groundwork to make this better?
The Service Catalog is basically a way to let technical end-users manage products via CloudFormation templates. Something could be built to do the same for Terraform, but I don't think there is anything like that right now.
CloudFormation does (attempt to) roll back to a known-good state if an update fails. Terraform just stops in the middle of what it was doing. I have mixed feelings about that.
I find YAML to have better tooling/editor support than HCL, though I actually prefer HCL.
There are existing CloudFormation wrappers (e.g. Troposphere) that can give you a full programming language on top of CloudFormation. To my knowledge, there isn't anything similar for Terraform.
Some things that have changed with 0.12:
CloudFormation has had `AWS::NoValue` for as long as I can remember. Terraform <=0.11 had special values (like empty string, number zero, etc.) that were special-cased as no value. Terraform 0.12 now has `null` properly.
Terraform <=0.11's ternary operators were maddening because both sides were evaluated, which led to errors, unlike CloudFormation's !If. In Terraform 0.12, the ternary only evaluates one side.
if by "correct it" you mean update the current template to match what's there, that'd be good, as AFAIK there's no easy way to do this at the moment. if by "correct it" you mean revert or change resources, no thanks. that sounds like a production accident waiting to happen, and you can un-drift (?) resources manually already.
Auditability and controls are one of the many facets of IaC. We require code changes to be approved by another developer, and similarly we require infrastructure changes to be approved by another developer. Regularly working outside the IaC tool would be in violation of that policy.
It's in this approval step that the change-set (CFN) or plan (Terraform) should be carefully reviewed by a human. If someone's made manual changes, reversion of them should appear here, and those should be unusual and eyebrow-raising. At that point, it's either fix the IaC definition of the infrastructure, manually un-drift it as you say, or do some workaround to ignore specific changes.
(To reiterate, IMO no one should ever run CFN/Terraform unattended on prod infrastructure, and there should always be a step to review the change-set/plan.)
I'll also say that the sword cuts both ways when it comes to prod outages and manual changes. Not so long ago, I ran into a prod-impacting issue when turning on multi-AZ for an RDS instance. In the other regions that had multi-AZ enabled, someone had manually added an extra parameter to the RDS parameter group, one that was required for certain app functionality to work. No one ever added it back to CloudFormation and that knowledge was eventually lost. When we enabled multi-AZ in a different region, we expected no problems at all, but instead we ended up with a whole section of app functionality breaking.
(This would've been before drift detection was a thing in CloudFormation, but actually I don't think RDS parameter groups are supported in CloudFormation's drift detection right now anyway. [0])
[0] https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
there is also a way to do this for cloudformation. look up cloudformer.
That said, I don't think CloudFormer does the same thing. I haven't used it before, so please correct me if I'm wrong, but to me it looks like it takes existing resources and generates a CloudFormation template out of them. But then you're still expected to upload that template and create a new stack, with all brand-new CloudFormation managed resources—is that right?
So for example, CloudFormer for an RDS instance would probably be a no-go.
In the Terraform case, after you've imported resources, there's no need to re-create them. There's the question of if what's in the template will match what the resources actually are, but there are also tools to generate Terraform templates straight off resources.
https://www.hashicorp.com/blog/hashicorp-terraform-0-12-prev...
Speaking of real code (not YAML or HCL) as infrastructure, anyone have experience with Pulumi?
It actually uses terraform's API's under the hood which is comforting in a way because you know it's building on a solid foundation.
I would not go back to writing HCL after experiencing Pulumi if I can avoid it, using a "real" programming language just feels a whole lot more natural and allows for much more powerful abstractions.
Users can pick and chose what modules, how many and set some params which are then validated.
The components are then rendered as json, fully declarative, no counts, no fancy TF hacks.
Data sources and a couple of home grown providers take care of whatever needs to be dynamic.
This way a lot of foot-guns are removed for consumers, and it simplifies the whole thing. At the expense of writing some code... worth it though, in a more enterprisey setting.
A bunch of oldschools sysadmins who "don't code"? Terraform is ridgid and on-rails enough that it probably helps sort of keep things sensible compared to just using boto. Almost like it was a framework, specifically defined to do that sort of thing.
It does sort of suck though, but what sucks less?
Edit: My solution is to stick as much as possible into k8s, but obviously that comes with its own warts, and to be fair to terraform, a lot of terraforms warts are just the underlying API warts leaking through.
If you are saying a nonsensical argument like "not IaC because it can't rollback" I can say CloudFormation is not IaC because it can't even import a local file to grab some values. Or run external commands
it was a jab at Terraform’s behavior
I am very curious though. Would you mind explaining this from a high level (to someone who knows cloud technology about as well as your dog)?
> Provision and Manage any Infrastructure Use infrastructure as code to consistently provision any cloud, infrastructure, and service.
Hence it's meaningless to me.
Given something like "I want three auto-scaling groups, each containing a minimum of three instances of type m4.xlarge, in the AWS us-west-2 region. And I want an S3 bucket that has permissions set up so that only the code running on those instances can read and write to the bucket. And I want a load balancer between all of them. And I want the instances to run Ubuntu 18.04 and to install these 6 dependencies on startup. And I want a large pepperoni pizza[1]." Terraform will read your credentials and make it happen.
[1] https://github.com/ndmckinley/terraform-provider-dominos
They both have their drawbacks, Terraform obviously suffers from not being a first-class service to AWS (since their service CloudFormation is a direct competitor). It is also possible to accidentally discard your Cloud "state" file, that keeps track of every instance that already exists and its current state (so TF can do a "diff" on what is there vs. what you're trying to apply), which definitely causes some headaches. In my experience the benefits of the design decisions from the tool far far outweigh any of the cons.
On the contrary, I have never had a good experience with CloudFormation. The workflow is long/slow, some of the AWS Best Practices are hilariously bad (looking at you "Paste this entire Python file as a text string into a yaml file), and more than a few times have I gotten into a state where CF-generated instances cannot be deleted and require the intervention of an AWS Rep.