Terraform 0.7 released
hashicorp.com
hashicorp.com
By the way, contributing to the project was very straightforward and the participants involved are super nice. I recommend you send them a patch if you have an itch!
[1]: https://www.terraform.io/docs/providers/do/r/volume.html
"Terraform is a tool for building, changing, and versioning infrastructure safely and efficiently. Terraform can manage existing and popular service providers as well as custom in-house solutions."
I'm using Terraform now, and I like it much better. It has its own language for defining resources, but I find it very readable and easy to write by hand.
My goal is to have my "infrastructure as code" so it can be tested in a sandbox and changes can be reviewed as pull requests. Going forward Terraform is my choice.
I've seen a lot of people start writing their own CF management systems, in Python or whatever, and end up with a lot of infrastructure defined in the logic of the homegrown management tool - like a badly written PHP page where presentation and logic are all mixed up. You see that once, you never want to go there again.
Terraform does not cover all AWS features, but the exceptions are few and tend to be lesser used features. The file format is human readable (this is subjective to some extent), and can be composed and modularized easily.
https://opencredo.com/terraform-infrastructure-design-patter...
https://github.com/hashicorp/best-practices
https://atlas.hashicorp.com/help/intro/use-cases/multiple-en...
https://www.terraform.io/intro/getting-started/modules.html
https://www.terraform.io/docs/state/remote/index.html
I'm not a big fan of buzzwords and soundbites, but Terraform comes pretty close to fulfilling the ideal of "software-defined infrastructure" in a way that is accessible, easy to use, easy to expand, and just makes sense - in the way a good programming language just makes sense.
We'll continue to use Terraform at my day job, but I don't recommend it to colleagues nor clients if they use only AWS resources. Simplicity is the ultimate sophistication, and adding yet another abstraction does not contribute towards that philosophy.
It's really unfortunate you had these experiences and I can't back up the statement when I say that this isn't normal and every release certainly does improve stability in a big way. Terraform is a very, very heavily tested (unit + acceptance) project and we fight very hard against regressions with multiple tests per bug in many cases at different levels.
I'd encourage you to continue giving future versions a shot and if you experience anything like this again, let us know and we'll react to that quickly relative to other types of issues.
I try to pick boring technologies that are battle tested through years of tech sector use; there's no glory for being on the bleeding edge, with only pain when things go south. So goes the ops struggle.
MIGBY, we have also been bitten in production in much the same way. If you are using only AWS resources there is no need to use terraform. Using it at all seems like a questionable decision to me since it's still firmly in "early adopters only" status.
If you're AWS-only then CloudFormation is the way to go. And IMO it will continue to be for a long time, until terraform is much more battle-tested and -hardened. Unless of course you're operating in an environment where you can afford to be exposed to the cutting edge of infrastructure-as-code. I definitely am not.
Personally, I use CloudFormation for my own stuff, because aside from reliability concerns, giving up on cfn-init to not write JSON (when I don't now, I use cfer, a Ruby DSL) is totally no bueno for me. Terraform offers no direct equivalent, though I've had HashiCorp sales people try to push Consul chewing-gum solutions that I'd have to manage much more directly than CFN metadata/cfn-init have been for me. Maybe in a couple years it'll be somewhere where I feel safer recommending it.
They also allow you to manually get out of the UPDATE_ROLLBACK_FAILED state, which was usually where you needed to contact AWS.
Cloudformation also supports Lifecycle events. They are a bit complicated to set up right, but they allow you to tell CloudFormation how to wait and how to proceed in response to resource signals (for instance, the deployment can wait or roll back if new instances in an ASG fail to be deployed). Terraform has none of that. It has "create before destroy", and "depends on", but that's about it AFAIK. If CloudFormation fails it will at least attempt to roll back to a sane state (and usually succeeds). Terraform will give up and leave you to clean up the mess.
Apart from that, Terraform has an advantage because of the modularity, ability to easily bring resources in, reference other resources in remote states etc.
Even if your project only expects to use AWS can you say that for certain of all your projects?
Also, HCL is way more readable than JSON. You can actually do a code review on a terraform plan.
And terraform as a CLI tool is pretty easy to use and expand. On top of this, you can write your own providers fairly easily and it's a more 'standard' tool as in, you don't get vendor lock in.
On top of all this, terraform is an open source Go project: you can easily make your own tools and import it's internal packages to hack things to your needs (caveat: internal packages =P).
CloudFormation is also nice in that it has access to a few API features which Terraform doesn't - specifically wait-conditions, automatically rolling-back updates when there's a problem, and a lot of the magic around UpdatePolicy/rolling deploys (although that doesn't work the way one would expect - a story for another day).
Having said that, Troposphere is pleasantly readable and has nice docs. The cross-platform integrations (for example, CloudFlare) add to the "just works" feeling. Some things are a pain to do in Terraform due to its strictness around being declarative and not having conditionals (e.g. having production and development environments that are similar, but not exactly the same), but there are some well-known hacks around them.
It also takes a very strong position around failed updates: it'll stop mid-change, tell you something's broken, and have you fix it. In the same situation, CloudFormation would roll back the changes to the state you were in before you started. Which one of those two failure modes you prefer is up to you. There are also some issues around having passwords and the like in your statefiles when you use RDS. You can avoid this if you go whole-hog into the HashiCorp stack (TF + Vault + Consul + etc.), but it's bugged me.
Lastly, If you're going to build something big and complex in TF, I suggest you do some reading of Charity Majors' rants on TF (https://charity.wtf/tag/terraform/). She has probably already run into the problem you're going to have.
So which would I use? Depends. If I want to create the same environment in a bunch of places (which also includes blue/green deploys), TF. If I want to do more complex things than that, CloudFormation.
working with json can be a pain, but there's lots of great tools out there that output cfn templates from either dsls or nicer formats
In the end, you choose your toolset and I have made a deliberate decision to stay as close to the provider's bare metal solution and build from there. I can build my own tiny "abstraction" a couple hundred lines of Python generating CF template and parsing the template into readable human format.
# you don't need to state the list in here variable "test" { type = "list" default = ["test1", "test2"] }
module "test_module" { test = "${var.test}" }
# in module test_module variable "test" { type = "list" }
# and now you can use it as ${var.test}
Hopefully it helps
- Very steep learning curve. It uses its own configuration language (HCL) which means you have new syntax to learn on top of its overwhelming API.
- It's stateful, so you have a state file that you have to store and keep synchronized in your team. There's capabilities for storing the state remotely but it's very much not ideal. You seemingly can't generate that state file from scratch based on your current provider's state.
- The templates are atrociously ugly and the string interpolation is limited, has extremely rough edges.
- It's very much an 0.x product. I keep hitting bugs. EC2 machines don't destroy properly if they have mounted drives. S3 buckets can't be force-destroyed if they have versioning enabled. Some normalization issues here and there which cause constant changes to show up.
- The custom format means the files aren't easily parsed and/or generated unless you're using hashicorp's own hcl go library. This sucks, to say the least. TOML with some conventions and jinja2 templating would have done the job just fine and would have been a ton more readable imho.
It's still good software and design. I would definitely pick it up for new project, but I wouldn't bother migrating existing ones just yet unless you know you'll need it. I use it alongside Ansible and the two pair quite nicely.
I still wish they'd gone with TOML :) They have good reasoning as to why not JSON/YAML, but those are the issues TOML actually solves and it's just so much more readable imho.
After all, it means that you have to start from scratch. If you any existing infrastructure — instances, load balancers, DNS, whatever — then anything you declare in Terraform will conflict, even if it's declared exactly identically.
A workaround is to manually reverse-engineer your current state into a state file, or use a third-party tool like Terraforming [1], but it still breaks the second anyone (either accidentally or intentionally for whatever reasons) bypasses Terraform and touches the world directly, so Terraform simply isn't resilient by design. Surely one of the primary uses of a tool like this must be that if something is modified behind your back, you are secure in the knowledge that you can whip out Terraform and it will force the world back into the right shape in a few seconds?
That being said, 0.7 has an import command now, so they recognize the problem. It's extremely limited and is not a complete solution yet, however.
I never understood the purpose of the state file. A tool like Terraform, from my perspective, should be stateless. It ought to compare its target to the world, and then attempt to converge the world to match the target. You could have local state as an option, in particular to detect interference, but I don't see why it's needed at all times.
My current solution is to use Salt (not Salt Cloud, just the Salt "boto" states, using local masterless mode), but it's pretty terrible. I would really want a tool like Terraform, but without the flawed state management.
One of the major features of the version released today is the ability to import existing resources in to Terraform state.
That being said, 0.7 has an import command now, so they recognize the problem.
It's extremely limited and is not a complete solution yet, however.> ... it still breaks the second anyone ... bypasses Terraform and touches the world directly, so Terraform simply isn't resilient by design.
This shouldn't be true. If you create a new resource that was never under management by Terraform, then yes, Terraform will ignore it. This is by design so that you can use Terraform to manage infrastructure and use other processes to manage other infrastructure that perhaps isn't quite migrated yet OR doesn't fit Terraform's model for whatever reason.
However, if you change a resource under management or even remove it, Terraform will notice this. It is very resilient to detecting drift.
> I never understood the purpose of the state file.
The purpose is to map your resources to what exists in the world. There must be some way to make a mapping there to be able to do diffing, drift detection, etc.
Some cloud platforms provide interesting ways to work around this (AWS tags are often abused for this) and we actually experimented with that approach early on. Unfortunately, not everything supports tags and there was no way for us to separate what was Terraform managed and what wasn't.
One of the features we actually have planned for Terraform in the future is the ability for it to tell you what IS NOT under Terraform management (by comparing local state to global state). For companies that have everything under management, this will be a great check to find any rogue resources. For companies working on adopting Terraform, this will help find things that still need to be migrated/imported.
So, hopefully our state improvements in 0.7 and in the short term will help sway you. They're definitely a top concern. But I also hope I explained my way through some of the design here!
Thanks for the feedback and we hope to see you in the community soon!
I greatly prefer the BOSH model: "give me a manifest and I will make the world look like it".
Documents are easily versionable. Commands aren't.
Documents narrow scope and can be more easily made idempotent as an automatic guarantee. Commands require us to remember to check the state ourselves, manually, before doing something.
That said, Terraform has a significantly different low-level model. BOSH looks at the size of the problem and imposes non-negotiable primitives to scope that complexity. Terraform is more ambitious. I suspect Terraform will be easier to adopt incrementally, but harder to manage on a per-unit-of-state basis.
[0] Loosely related: http://chester.id.au/2012/06/27/a-not-sobrief-aside-on-reign...
This is how Terraform works if you want a one-time setup. You can just throw the state away after that if you want. But Terraform is supposed to be used (and is very very often used) to do ongoing minimal changes to infrastructure, often dozens of times per day.
In this scenario, Terraform's configurations are completely declarative: it is a state of the world you want to reach, it isn't what to do next.
However, we require the state file in order to find the resources we own and then refresh the state of the world so we can do a diff.
We fixed the bug. But somehow, that previous build wiped the state file. When we pushed the fix, terraform freaked out and started deleting all our instances, including S3 redirects that weren't managed by it which we explicitly told it to ignore.
That sucked.
No harm done, we're still in beta to catch bugs exactly like those, but our initial enthusiasm with terraform has been waning very quickly. Like I said, I love the design, but it's just not there yet.
I've seen BOSH used this way.
You just have to accept that editing a file and checking it into version control is a good thing.
After a while, infrastructure tends to stabilise into a predictable format for the given distributed system under management.
I think Terraform would definitely be easier for a fast-developing system because of its bias towards interactivity. But in the long run most systems cease to be fast-developing.
At Pivotal we're seeing more and more customers going all-in on BOSH. Not just for Cloud Foundry, but for pretty much every stateful service they can lay a hand on. They like that it's based entirely on versionable, auditable files.
> However, we require the state file in order to find the resources we own and then refresh the state of the world so we can do a diff.
BOSH will inspect the world when you ask it to, or you can ask it to do so continuously. As a rule, most operators prefer (like Terraform users) to do it manually.
It's great that you're working on fixing that, but it looks like it will be some time before it's quite there, given the seemingly slow pace of development.
It's nice that Terraform can support untaggable resources, but I don't understand why that's the general case and not a special case. I can't remember an AWS resource that isn't nameable or taggable.
ElasticIP is one among many. For a complete list: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Using_Tag...
This is how BOSH has been doing it for years for IaaS setup and other systems like Salt/Puppet/Chef/cfengine for servers. Compare-and-repair is a fairly well-accepted, trustworthy model.
Disclaimer: I work for Pivotal, we donate most of the engineering on BOSH.
[0] https://bosh.io/
- Often beats AWS to support on new features. AWS is traditionally not very good at coordinating service releases between IAM, CloudFormation, and whatever.
- Modules better than nested stacks
- Much better insight into what's going to change. This still seems to be the case with CloudFormation's change sets.. Particularly with nested stacks, which were sent to earth to torment us.
- Nicer configuration than CloudFormation
Woes:
- Juggling multiple state files. This seems to be necessary if you are going to stamp out infrastructure across accounts and regions; this is double true if you want to roll changes gradually across regions/accounts. Prepare to write wrappers.
- Creating modules is the only wait to organize tf files into folders. This can really rub you the wrong way because modules are isolated from the root and other modules' variables. You have to start passing in and out resource references :| There was an import proposal made 7 days ago that looks very nice.
- Not much in the way of rollbacks..
- Potentially less flexible than CloudFormation; CloudFormation offers wait conditions, conditionals, rollbacks, update cancellations(with rollback), etc
I feel strongly about two things:
1) I screwed up at the beginning so hard (there were _no_ best practices to be found so it was a bit wild west) that I've come to the point where I need to use something else, or rewrite everything.
2) Terraform is best when you are in a silo by yourself, or you're using something to do the state applies for you (Atlas, or DIY through Jenkins and friends).
It does a lot right that other tools don't do (stringing together resources is a cinch), but it also behaves in strange things and the statefile is both a huge positive, and a huge negative.
If you need to move fast it's not a bad tool to choose at all.
The BIG caveat here is as follows -- If you have a mixed-cloud setup and use AWS, GCE, Azure, DigitalOcean, etc.. at the same time; It's probably hands down the best tool for managing everything.
Notes:
- I only use AWS. I use some of the other TF providers mixed in with it. (Consul, Datadog)
- I started from scratch (no AWS resources -> 100% Terraform), the general opinion seems to be this is the better place to start from.
- I use a majority of Hashicorp's stuff and am quite partial to liking their products. Terraform is something I have the most opinions on. I have no problems strongly suggesting Packer, Consul, Vagrant, and Vault to anyone.
- One of the main reasons I'm considering alternatives is that, at my current position, I'm having trouble getting folks to wrap their heads around Terraform.
+++ Use jetbrains free IDE (IntelliJ, PyCharm...) and install the plugins for HCL support.
You'll get syntax highlighting + auto completion for variable and infrastructure names + visual report on lines with errors and quickfixes.
Now HCL just because the easiest language in the world. (Seriously, it's not even a language, it takes 1 hour to master the syntax, top.)
+++ Use terraform to create all the infrastructure backbone. (VPC, subnets, routing tables, DHCP options, etc... )
Terraform is the single best tool in the world to do that. It's also the only one currently in existence that is actually able to do that. (ansible and the likes are missing critical modules).
+++ Do NOT use terraform to manage instances nor security groups.
These things change wayyy too often and you'll struggle with states.
Use one of the ansible/salt to create and manage them. They refresh their inventory to keep with external changes (at least ansible does) and they allow to run scripts on the resources. (e.g. tag all the EBS volumes of the instance, after the instance is up).
Bonus: For every instance "name" you create, create and put it in a security group with the same "name". Then you can manage security groups by name ;)
+++ Terraform can import existing infrastructure since v0.7 (this release :D)
Just sayin'.
+++ The state files is annoying to share and keep updated
It used to be put in git. (it's still possible).
It can be put in S3 (or similar options). The sharing is faster and it makes less state issues.
+++ Because of the state files, terraform is good for "static things" and terrible for "fast moving things".
That makes terraform appropriate for the VPC/subnet/... provisioning (the first point in this list )
But it's terrible for instances and security groups (another point in the list )
Known your strengths and weaknesses, choose the right tool for the job ;)
Wound up taking a day and writing a ruby script that checked out an STS token, and exec'ed a new bash shell with the correct environment variables. the $PS1 has the account name in there so it's not too hard to know what shell points where. Also helpful that we set up a "control account" that handles consolidated billing and user login.
We've had some pretty good success so far, however I can definitely see how that feature would be good baked-in somehow.
Having this built into TF would be nice, but there are enough tools out there that don't support AWS role-jumping that I suspect we'll end up using that wrapper for a long time.
[0] https://github.com/aws/aws-sdk-go/issues/472#issuecomment-23...
Kudos to the TF team/community and Mitchell