Terraform 0.12
hashicorp.com
hashicorp.com
it was a jab at Terraform’s behavior
"Cloud" configuration may have been Terraform's proverbial toe in water, but the truly untapped potential lies in the other providers. Anything that can be packaged as a Terraform provider exposing resource abstractions can be easily managed using convenient HCL syntax. This, IMHO, is the unfortunately buried lede.
Plus, the different philosophies of mutable/immutable infrastructure that the different capabilities/limitations in each tool encourage.
Sorry, but I have yet to meet a developer that used Terraform and willingly wants to keep using it after seeing it fail. When your tool cannot keep track of the resources it created or when it gets into situations where it doesn’t allow you to do certain things (like deleting all resources created) and you have to relearn how the underlying cloud works, it’s time to move on.
Do yourselves a favor and use Cloudformation (or your favorite cloud’s equivalent and just move on with your life)
Also, the terraform language, HCL? It's, I guess there's no better way to put this: not good.
Am I misunderstanding the complexity of what Terraform is trying to do? To me, it looks like a bunch of tiny API clients tied together with a topological sort --- in other words, it's just another species of "make". It feels like it carries a whole lot more complexity than that concept warrants, but just enough simplicity to make it not a serious programming language. It's, to me, one of those frustrating uncanny valley systems.
I am prepared to be totally wrong about this.
It's a good idea and not a bad design, but the user experience is pretty bad and nearly all the operational aspect is an afterthought. It was not built to be sold as a product, hence it kind of sucks as a product, but it's fine as a free tool. The fact that it's the best free tool we have for this task speaks volumes about how most companies are deathly afraid to work as a community to build better solutions.
I do think you're selling Terraform short, though. Sure, the core is the toposort-create-things. But it also stores the state of its created things and (crucially) has the ability to diff the actual state of resources against what it thinks they ought to be.
Being able to inspect existing resources and diff state is also what lets it import existing resources, so that they can be Terraform-managed going forward.
Terraform is also capable of determining if its planned changes can be performed in-place or if they require resources to be destroyed and re-created. That boils down to a boolean flag on a field, ultimately, but it's still something a dead-simple make clone probably wouldn't do well.
I've only really used Terraform seriously for AWS, so I'm not sure about the other providers, but the Terraform AWS Provider has an enormous amount of work behind it. Basically every resource API has schema validation written in the AWS provider, and depending on the resource there are often eventual-consistency issues handled by the provider. See for example [0].
In contrast, AWS CloudFormation: can't import existing resources; isn't always sure whether an update will require replacement or not; and, as of ~6 months ago, can detect configuration drift, but not correct it (!). Of course, CloudFormation wins in other areas...
[0]: https://github.com/terraform-providers/terraform-provider-aw...
Also CF is the “native” language of AWS. There are plenty of getting started examples from AWS where they give you the template. Also Elastic Beanstalk extensibility is built on top of CF.
Not to mention Codestar that will set up environments for you for common use cases and exports templates and the lambda environment lets you configure everything from the console, test it and then you can export the CF definition.
I've never had Terraform "forget" resources or disallow me from deleting them (unless I requested that). Maybe those are bugs that've since been fixed?
It sounds like we've had very different experiences with these tools.
Like one sibling comment mentioned, getting support from AWS is nice. You can buy Terraform support, too, but knowing HashiCorp it'd probably cost more than most peoples' AWS bills in their entirety. (This is me being a little cheeky and unfair—HashiCorp handles community support via GitHub Issues for Terraform really well.)
CloudFormation's built straight into AWS, so there's no need to set up state file storage or locking or worry about a state file at all, really. This has its own set of drawbacks, but it's nice for getting started, and in theory it makes CloudFormation more robust out-of-the-box.
CloudFormation StackSets are really nice if you have identical resources that need to be placed in many different regions and accounts. (Example: we use it to place GuardDuty and Config, both regional services, in every region.) With Terraform, this means copy-and-paste, as far as I know, though maybe Terraform 0.12 sets up the groundwork to make this better?
The Service Catalog is basically a way to let technical end-users manage products via CloudFormation templates. Something could be built to do the same for Terraform, but I don't think there is anything like that right now.
CloudFormation does (attempt to) roll back to a known-good state if an update fails. Terraform just stops in the middle of what it was doing. I have mixed feelings about that.
I find YAML to have better tooling/editor support than HCL, though I actually prefer HCL.
There are existing CloudFormation wrappers (e.g. Troposphere) that can give you a full programming language on top of CloudFormation. To my knowledge, there isn't anything similar for Terraform.
Some things that have changed with 0.12:
CloudFormation has had `AWS::NoValue` for as long as I can remember. Terraform <=0.11 had special values (like empty string, number zero, etc.) that were special-cased as no value. Terraform 0.12 now has `null` properly.
Terraform <=0.11's ternary operators were maddening because both sides were evaluated, which led to errors, unlike CloudFormation's !If. In Terraform 0.12, the ternary only evaluates one side.
if by "correct it" you mean update the current template to match what's there, that'd be good, as AFAIK there's no easy way to do this at the moment. if by "correct it" you mean revert or change resources, no thanks. that sounds like a production accident waiting to happen, and you can un-drift (?) resources manually already.
Auditability and controls are one of the many facets of IaC. We require code changes to be approved by another developer, and similarly we require infrastructure changes to be approved by another developer. Regularly working outside the IaC tool would be in violation of that policy.
It's in this approval step that the change-set (CFN) or plan (Terraform) should be carefully reviewed by a human. If someone's made manual changes, reversion of them should appear here, and those should be unusual and eyebrow-raising. At that point, it's either fix the IaC definition of the infrastructure, manually un-drift it as you say, or do some workaround to ignore specific changes.
(To reiterate, IMO no one should ever run CFN/Terraform unattended on prod infrastructure, and there should always be a step to review the change-set/plan.)
I'll also say that the sword cuts both ways when it comes to prod outages and manual changes. Not so long ago, I ran into a prod-impacting issue when turning on multi-AZ for an RDS instance. In the other regions that had multi-AZ enabled, someone had manually added an extra parameter to the RDS parameter group, one that was required for certain app functionality to work. No one ever added it back to CloudFormation and that knowledge was eventually lost. When we enabled multi-AZ in a different region, we expected no problems at all, but instead we ended up with a whole section of app functionality breaking.
(This would've been before drift detection was a thing in CloudFormation, but actually I don't think RDS parameter groups are supported in CloudFormation's drift detection right now anyway. [0])
[0] https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
there is also a way to do this for cloudformation. look up cloudformer.
That said, I don't think CloudFormer does the same thing. I haven't used it before, so please correct me if I'm wrong, but to me it looks like it takes existing resources and generates a CloudFormation template out of them. But then you're still expected to upload that template and create a new stack, with all brand-new CloudFormation managed resources—is that right?
So for example, CloudFormer for an RDS instance would probably be a no-go.
In the Terraform case, after you've imported resources, there's no need to re-create them. There's the question of if what's in the template will match what the resources actually are, but there are also tools to generate Terraform templates straight off resources.
https://www.hashicorp.com/blog/hashicorp-terraform-0-12-prev...
Speaking of real code (not YAML or HCL) as infrastructure, anyone have experience with Pulumi?
It actually uses terraform's API's under the hood which is comforting in a way because you know it's building on a solid foundation.
I would not go back to writing HCL after experiencing Pulumi if I can avoid it, using a "real" programming language just feels a whole lot more natural and allows for much more powerful abstractions.
Users can pick and chose what modules, how many and set some params which are then validated.
The components are then rendered as json, fully declarative, no counts, no fancy TF hacks.
Data sources and a couple of home grown providers take care of whatever needs to be dynamic.
This way a lot of foot-guns are removed for consumers, and it simplifies the whole thing. At the expense of writing some code... worth it though, in a more enterprisey setting.
A bunch of oldschools sysadmins who "don't code"? Terraform is ridgid and on-rails enough that it probably helps sort of keep things sensible compared to just using boto. Almost like it was a framework, specifically defined to do that sort of thing.
It does sort of suck though, but what sucks less?
Edit: My solution is to stick as much as possible into k8s, but obviously that comes with its own warts, and to be fair to terraform, a lot of terraforms warts are just the underlying API warts leaking through.
If you are saying a nonsensical argument like "not IaC because it can't rollback" I can say CloudFormation is not IaC because it can't even import a local file to grab some values. Or run external commands
I am very curious though. Would you mind explaining this from a high level (to someone who knows cloud technology about as well as your dog)?
> Provision and Manage any Infrastructure Use infrastructure as code to consistently provision any cloud, infrastructure, and service.
Hence it's meaningless to me.
Given something like "I want three auto-scaling groups, each containing a minimum of three instances of type m4.xlarge, in the AWS us-west-2 region. And I want an S3 bucket that has permissions set up so that only the code running on those instances can read and write to the bucket. And I want a load balancer between all of them. And I want the instances to run Ubuntu 18.04 and to install these 6 dependencies on startup. And I want a large pepperoni pizza[1]." Terraform will read your credentials and make it happen.
[1] https://github.com/ndmckinley/terraform-provider-dominos
They both have their drawbacks, Terraform obviously suffers from not being a first-class service to AWS (since their service CloudFormation is a direct competitor). It is also possible to accidentally discard your Cloud "state" file, that keeps track of every instance that already exists and its current state (so TF can do a "diff" on what is there vs. what you're trying to apply), which definitely causes some headaches. In my experience the benefits of the design decisions from the tool far far outweigh any of the cons.
On the contrary, I have never had a good experience with CloudFormation. The workflow is long/slow, some of the AWS Best Practices are hilariously bad (looking at you "Paste this entire Python file as a text string into a yaml file), and more than a few times have I gotten into a state where CF-generated instances cannot be deleted and require the intervention of an AWS Rep.
Also, I'll go out on a limb and say that I dislike the flexibility of iteration allowed in HCL 2. I know that people overwhelmingly asked for it, but my opinion is that it demonstrates a fundamental misunderstanding of how the v. 11 and earlier system was designed, and just how powerful completely declarative code can be.
Additionally you have to rewrite everything for each cloud provider, so it just expands the work required, all with no real IDE integration. I’d just write against the cloud provider APIs directly or look at Pulumi when they get interactive debugging.
Personally I’m excited that loop operations now exist in 0.12
But thanks for the reference to Pulumi, had not heard of this and it looks very interesting.
The lesson here is that if you plan on making a language, even a DSL, you want to be sure you're really up for it since it's a lot of work.
Then, I needed to launch infra in GCP and I messed around with terraform unsuccessfully for a few days before writing about 10 lines of gcloud CLI commands into a makefile.
Now I just check a makefile into my project and just break things up into little shell scripts.
Solved so many problems and headaches.
If there wasn't a chip on my shoulder telling me I had to use what the next person would expect else as a contractor I risk being seen as unprofessional I'd do it in a heartbeat.
I've never used Terraform but CloudFormation just seems to suck, the documentation is poor and relatively few people are sharing their stack files. I've lost count of the number of times I've hit an error only to find Google hasn't heard of it.
It's a bit disconcerting when one of the BIG aspects of the Terraform 0.12 release is some support for preliminary types.
Pulumi just uses Typescript.
I'm not entirely sure why Terraform is going in the direction of reinventing the wheel when it can leverage stuff out there.
I haven't found a solution for it, but I have lots of resources that are almost identical, except for a few arguments, but right now it seems like my options are to either write my own Ruby script to output .hcl files, or live with two resources that are almost identical, and live with the mistakes that could entail by future modifications.
I haven't had the time to look into HCL2 yet, but maybe it is solved.
resource "aws_instance" "nomadclients_tick" {
ami = "${data.aws_ami.nomadclient_tick.image_id}"
instance_type = "t2.micro"
count = 3
iam_instance_profile = "${aws_iam_instance_profile.consul-join.name}"
subnet_id = "${element(aws_subnet.consul.*.id, count.index)}"
vpc_security_group_ids = [
"${aws_security_group.rule01.id}",
"${aws_security_group.rule02.id}",
"${aws_security_group.rule03.id}",
"${aws_security_group.rule04.id}",
]
}
resource "aws_instance" "nomadclients_tock" {
ami = "${data.aws_ami.nomadclient_tock.image_id}"
instance_type = "t2.micro"
count = 3
iam_instance_profile = "${aws_iam_instance_profile.consul-join.name}"
subnet_id = "${element(aws_subnet.consul.*.id, count.index)}"
vpc_security_group_ids = [
"${aws_security_group.rule01.id}",
"${aws_security_group.rule02.id}",
"${aws_security_group.rule03.id}",
"${aws_security_group.rule04.id}",
]
}
The examples I have found feels very difficult. resource "aws_instance" "nomadclients" {
ami = "${count.index > 2 ? data.aws_ami.nomadclient_tick.image_id : data.aws_ami.nomadclient_tock.image_id}"
instance_type = "t2.micro"
count = 6
iam_instance_profile = "${aws_iam_instance_profile.consul-join.name}"
subnet_id = "${element(aws_subnet.consul.*.id, count.index)}"
vpc_security_group_ids = [
"${aws_security_group.rule01.id}",
"${aws_security_group.rule02.id}",
"${aws_security_group.rule03.id}",
"${aws_security_group.rule04.id}",
]
}
The downside to this approach (in TF < 0.12), however, will become apparent when you want to modify the number of instances in one or both pools. This arrises from the way hcl v1 manages the counter index with each resource in the state; That is each resource is linked to a specific index and when that index is altered in someway, terraform will attempt to adjust the resources at each respective index accordingly.Because of these headaches, we decided to abandon hcl for these use cases and ended up writing our own preprocessor that takes a JSON configuration and generates individual JSON-based HCL (ie. .json.tf) that terraform can then use. We can leverage proper templating to generate many variants of a single resource in a higher level language, while still using terraform to manage the infrastructure and it only adds one additional step to the plan + apply process.
locals {
nomad_ami = [
"${data.aws_ami.nomadclient_tick.image_id}",
"${data.aws_ami.nomadclient_tock.image_id}",
]
}
resource "aws_instance" "nomadclients" {
count = 6
ami = "${element(local.nomad_ami, count.index)}"
...
}[0] https://www.terraform.io/docs/modules/index.html
[1] https://www.terraform.io/docs/modules/sources.html#local-pat...
[2] https://www.terraform.io/docs/configuration/variables.html
I have two use cases in mind: Use case 1: need to have a reproducible way to generate terraform folders for multiple almost-identical deployments (this can be solved via modules) Use case 2: need way to "promote" a deployment from staging to production (load a .tf file, change a few basic params then save again to a different folder --- doesn't really work with templates since I want to load an existing .tf file)
One thing that might work is to use the combination of two tools: - https://github.com/virtuald/pyhcl = reads .tf, can export .tf.json - https://github.com/kvz/json2hcl = convert .tf.json to .tf but feels hacky...
Any other recommendations for .tf parsing and generation? (preferably in python or scriptable via python)
main.tf.json
Terraform 0.x before 0.12 was heavily limited due to the syntax, you just worked around so many limitations. We are actively ignoring those limitations and we started working without our own templates using jinja that just generate vanilla terraform. No counts and no if/else needed.
Maybe with .12 we can move some of these back to plain terraform.
Upgrade so far is not smooth though, there are a lot of pains with the type system vs the plain old “I’ll figure it out for you”.
ALL of this being said, I really appreciate Hashicorp’s work on this, we could not imagine our life without terraform.
I switched from Terraform to Pulumi for a personal project recently and haven't looked back (no affiliation). Writing it with Typescript means you get excellent IDE support (using VSCode here) and access to the enormous JS ecosystem. I've also found myself creating a number of useful abstractions - that I would never have bothered to with HCL - like IAM helper functions, eg:
const limitedReadAccessPolicy = createPolicy(
"product-table-read-access",
allow(
["dynamodb:Get*", "dynamodb:Query"],
[table.arn, interpolate`${table.arn}/index/*`]
)
);
vs resource "aws_iam_policy" "product_read_access" {
name = "madu_${var.env_name}_product_read"
policy = <<EOF
{
"Version": "2012-10-17",
"Statement": [
{
"Action": [
"dynamodb:Get*",
"dynamodb:Query"
],
"Effect": "Allow",
"Resource": [
"${aws_dynamodb_table.products.arn}",
"${aws_dynamodb_table.products.arn}/index/*"
]
}
]
}
EOF
}[1] https://app.terraform.io/signup?utm_source=banner&utm_campai...
Having your infrastructure in an inventory-like codebase makes it clearer to reason about. It can't go too crazy like most programs end up like
It always looks at things using provider specific resource, while IMHO it should just expose a bunch of predefined resource types (see rOCCI specs e.g.) and then allow you to attach a specific provider to it.
IMHO the biggest win as a user would be having not to have an implementation for every provider over and over. Do we really need to have a consul module for Azure, AWS, Tudeluuu and god knows who? No.
That being said, Terraform in the long run still is the most reliable tool in that space.
The whole situation about state management is... lacking. Experience says the one thing no client ever wants in the cloud but always on prem is state.
Hm, can you elaborate? S3 state seems perfectly serviceable, and I don't immediately see why I would want to operate on-prem resources just to maintain state.
Currently Terraform Enterprise is the only approach to solving that and its a good one but... only if you're one of the big guns and even those sometimes think twice because the infrastructure state requires much more security than that offers.
State really shouldn't have anything secret in it in the first place, so 3rd party having access to it shouldn't matter.
For light usage, I find managing Terraform's state to be a significant hurdle. You basically have no choice but to set up secure remote storage unless you want to check passwords into source control. In contrast `kubectl apply` is so easy to use since it's stateless. It just creates or updates any resources provided, and it even supports --prune if you want the set of configuration to be treated as comprehensive.
It seems like the main things that Terraform adds:
1. The ability to work with providers that require you to store their generated IDs to reference later. With kubernetes, the kind and name of the resource is enough to identify it; it does get assigned a UID, but you don't have to include that in the configuration since keys that are excluded are left as-is.
2. The ability to work with multiple different providers. I'm not sure how often you do have a single terraform project(is that the term?) with more than one provider, but I guess using the same set of tools, even if the configuration is provider-dependent, is nice.
Is that accurate? Does Terraform offer any other advantages?
If you were building a configuration mechanism for your system from scratch to allow your users to configure it as code, would you make .a Terraform provider over a command line tool that can apply [--prune] that same configuration?
> I find managing Terraform's state to be a significant hurdle Have a look at https://app.terraform.io/signup/account, it's the free version of Terraform Enterprise and it makes managing your state very easy.
>If you were building a configuration mechanism for your system from scratch to allow your users to configure it as code, would you make .a Terraform provider over a command line tool that can apply [--prune] that same configuration?
I would make a Terraform provider, you get plans, modules, the possibility to integrate with other providers easily and creating a new Terraform provider is actually very easy!
Infrastructure as code becomes progressively more important as you add resources until you're at the point where terraform does the job of a whole team of administrators who would be spending all day clicking through UIs.
Personally, I think staging/production is the greatest advantage. I can experiment and break things on staging, then when I apply the plan to production I can be confident I’m getting the same infrastructure.
Infrastructure-as-code brings many programming amenities. You can write comments. You can easily see what a resource depends on and you can grep the repository to see what depends on it. You can version control it, which gives you a record of who changed what and possibly why. You can do code reviews.
It does take more time then clicking through a UI, but I can’t imagine operating all but the smallest infrastructure without Terraform or similar.
It's easy to misuse ansible though. People try to consume it like puppet and use it for whole-system-declarative-state, which is usually not worth the squeeze in ansible.
Treat it as glorified bash-scripting with smarter data handling and yaml-driven config files, and you'll have success.
I also made the mistake of `terraform plan`ning and updating my code as I went along. Just use `terraform validate`. Otherwise you're going to inadvertently promote the statefile before you're done dealing with all the issues (and you don't want that because it prevents you from aborting your upgrade and switching back to 0.11 till you're all ready). Not a real problem because the statefile is versioned but an annoyance nonetheless.
The type issues were mostly easy to deal with, except for where things that were assignments now being blocks. For instance, look at this:
master_authorized_networks_config {
cidr_blocks {
cidr_block = "10.0.0.0/8"
display_name = "Example /8"
}
cidr_blocks {
cidr_block = "10.1.2.3/32"
display_name = "Example /32"
}
}
That looked like: master_authorized_networks_config = {
cidr_blocks = [
{
cidr_block = "10.0.0.0/8"
display_name = "Example /8"
},
{
cidr_block = "10.1.2.3/32"
display_name = "Example /32"
},
]
}
before and `terraform 0.12upgrade` isn't about to help you navigate this. Especially if you previously assigned from a variable. In that case, it's going to make this monstrosity of a `for_each` over that thing. Jesus Christ.Still, I'm thrilled for the new stuff with the more type-safety. Not going to complain. If this is the price, then I'll pay. I just wish they'd done more to help the upgrade, but it's an 0.x release so fine.
As you supposed to download this, use it’s language and syntax which is all it’s own thing, to define you services and then export that to a YAML setup that AWS CloudFormation (for example) is expecting?
I assume there are reasons I wouldn’t just define it myself in YAML directly?
Now, other than the portability of AWS to Azure to GCP, why do I want this?
Because with CloudFormation I can pull down and put up my “stacks” in layers and this isn’t necessarily the same as API calls (although on the backend maybe it is).
Terraform supports each of those providers, but the resources are specific to each provider.
You cannot use the same Terraform configuration on AWS and Azure.
Nor would you want to, per se. As far as I know, this is by design.
The state of the infrastructure is stored, ideally, in the cloud. You apply your code during which terraform identifies which changes need to be made by comparing the current state with your local changes and executes those changes as you watch.
The end result is well-defined, testable, repeatable cloud infrastructure provisioning.
Infrastructure as code. Store it in git. Profit.
Ok. But so is CloudFormation YAML files, mostly.
So is the advantage of Teraform that it works beyond AWS?
I never leave the AWS ecosystem at all.
I’m not challenging what you do, just trying to figure out what makes sense for me.
Regarding IaC "rollback" capabilities - I don't think it really exists in a way that makes it reasonable for people to depend on. The path forward is to stand up new infrastructure, canary test, etc - then steer via load-balancers the traffic to your new nodes, and then finally destroy the old ones. I loathe trying to maintain a fleet of nodes that can drift into various states of dis-repair. I love the idea of blowing everything away and having a 100% clean and predictable environment again. It makes me happy.
</end rant>
Hope that helps - I really have fallen in love with Terraform after having to cobble down N number of CLI tools for various cloud providers, hypervisor providers, etc. Terraform at least gives an easy to use language and abstracts me above recursive API call logic that I would otherwise need to write for basic things.
Hmm, I’ll have to look into it some more. I really like the idea of layering my application into dynamo set up, lambdas, polices, etc. and being able to update a specific layer, pull it down and put the new one up.
Last thing I want is more headache.
Atlantis (https://github.com/runatlantis/atlantis) is also a really cool project for managing teams working on terraform projects.
I imagine using git to store state would be mightily inconvenient.
Has anyone tried the approach of writing some go code that imports terraform, rather than using the terraform CLI? This would give you the full power of golang to set up your resources.
But HCL is giving me Puppet DSL PTSD.
Folks at Hashicorp should just embed JS engine and let users write definitions using a real language (JS).
It's totally doable using library such-as Otto.
On the other hand, I'm usong React-Native, lol
Looks like a minor update to me...
I mean, it has been around for like 5+ years...
I am even more astounded that people are happy to use a product that by its own definition is not stable.
Same happen in ruby a lot. You find a Gem that claims to follow semantic versioning and it is still in 0.y.z after years of being used in production which flies in the face of https://semver.org/#how-do-i-know-when-to-release-100
However I know of no common versioning scheme where 0.y.z is considered production ready.
OpenSource projects have another pace compared to commercial solutions
https://learn.hashicorp.com/terraform/aws/eks-intro
So, I'd say Kubernetes itself does more heavy lifting but not all of it?
https://github.com/JasonCarter80/contoso_aws_k8s/tree/master...
Ansible seems to focus on managing infrastructure, Terraform shines when creating infrastructure.
it attempts to and encourages declarative configuration, but is very hard to keep that way. it is difficult and requires determination to make Terraform do something in a non-declarative fashion.
the end result is that when I look at a Terraform configuration I can very easily tell what is going on, because the end result is exactly what I read. where with Ansible, it very much depends on understanding the current state of the server you are about to run this playbook on, and you just have to cross your fingers and hope for the best.
Terraform is a tool designed to overcome the problems that "best practices" will supposedly prevent you from introducing in your Ansible playbook.
don't get me wrong. Ansible is a wonderful tool for certain applications (configuration management). I would just never use it to spin up infrastructure again.