- Very steep learning curve. It uses its own configuration language (HCL) which means you have new syntax to learn on top of its overwhelming API.
- It's stateful, so you have a state file that you have to store and keep synchronized in your team. There's capabilities for storing the state remotely but it's very much not ideal. You seemingly can't generate that state file from scratch based on your current provider's state.
- The templates are atrociously ugly and the string interpolation is limited, has extremely rough edges.
- It's very much an 0.x product. I keep hitting bugs. EC2 machines don't destroy properly if they have mounted drives. S3 buckets can't be force-destroyed if they have versioning enabled. Some normalization issues here and there which cause constant changes to show up.
- The custom format means the files aren't easily parsed and/or generated unless you're using hashicorp's own hcl go library. This sucks, to say the least. TOML with some conventions and jinja2 templating would have done the job just fine and would have been a ton more readable imho.
It's still good software and design. I would definitely pick it up for new project, but I wouldn't bother migrating existing ones just yet unless you know you'll need it. I use it alongside Ansible and the two pair quite nicely.
I still wish they'd gone with TOML :) They have good reasoning as to why not JSON/YAML, but those are the issues TOML actually solves and it's just so much more readable imho.
After all, it means that you have to start from scratch. If you any existing infrastructure — instances, load balancers, DNS, whatever — then anything you declare in Terraform will conflict, even if it's declared exactly identically.
A workaround is to manually reverse-engineer your current state into a state file, or use a third-party tool like Terraforming [1], but it still breaks the second anyone (either accidentally or intentionally for whatever reasons) bypasses Terraform and touches the world directly, so Terraform simply isn't resilient by design. Surely one of the primary uses of a tool like this must be that if something is modified behind your back, you are secure in the knowledge that you can whip out Terraform and it will force the world back into the right shape in a few seconds?
That being said, 0.7 has an import command now, so they recognize the problem. It's extremely limited and is not a complete solution yet, however.
I never understood the purpose of the state file. A tool like Terraform, from my perspective, should be stateless. It ought to compare its target to the world, and then attempt to converge the world to match the target. You could have local state as an option, in particular to detect interference, but I don't see why it's needed at all times.
My current solution is to use Salt (not Salt Cloud, just the Salt "boto" states, using local masterless mode), but it's pretty terrible. I would really want a tool like Terraform, but without the flawed state management.
One of the major features of the version released today is the ability to import existing resources in to Terraform state.
That being said, 0.7 has an import command now, so they recognize the problem.
It's extremely limited and is not a complete solution yet, however.> ... it still breaks the second anyone ... bypasses Terraform and touches the world directly, so Terraform simply isn't resilient by design.
This shouldn't be true. If you create a new resource that was never under management by Terraform, then yes, Terraform will ignore it. This is by design so that you can use Terraform to manage infrastructure and use other processes to manage other infrastructure that perhaps isn't quite migrated yet OR doesn't fit Terraform's model for whatever reason.
However, if you change a resource under management or even remove it, Terraform will notice this. It is very resilient to detecting drift.
> I never understood the purpose of the state file.
The purpose is to map your resources to what exists in the world. There must be some way to make a mapping there to be able to do diffing, drift detection, etc.
Some cloud platforms provide interesting ways to work around this (AWS tags are often abused for this) and we actually experimented with that approach early on. Unfortunately, not everything supports tags and there was no way for us to separate what was Terraform managed and what wasn't.
One of the features we actually have planned for Terraform in the future is the ability for it to tell you what IS NOT under Terraform management (by comparing local state to global state). For companies that have everything under management, this will be a great check to find any rogue resources. For companies working on adopting Terraform, this will help find things that still need to be migrated/imported.
So, hopefully our state improvements in 0.7 and in the short term will help sway you. They're definitely a top concern. But I also hope I explained my way through some of the design here!
Thanks for the feedback and we hope to see you in the community soon!
I greatly prefer the BOSH model: "give me a manifest and I will make the world look like it".
Documents are easily versionable. Commands aren't.
Documents narrow scope and can be more easily made idempotent as an automatic guarantee. Commands require us to remember to check the state ourselves, manually, before doing something.
That said, Terraform has a significantly different low-level model. BOSH looks at the size of the problem and imposes non-negotiable primitives to scope that complexity. Terraform is more ambitious. I suspect Terraform will be easier to adopt incrementally, but harder to manage on a per-unit-of-state basis.
[0] Loosely related: http://chester.id.au/2012/06/27/a-not-sobrief-aside-on-reign...
This is how Terraform works if you want a one-time setup. You can just throw the state away after that if you want. But Terraform is supposed to be used (and is very very often used) to do ongoing minimal changes to infrastructure, often dozens of times per day.
In this scenario, Terraform's configurations are completely declarative: it is a state of the world you want to reach, it isn't what to do next.
However, we require the state file in order to find the resources we own and then refresh the state of the world so we can do a diff.
We fixed the bug. But somehow, that previous build wiped the state file. When we pushed the fix, terraform freaked out and started deleting all our instances, including S3 redirects that weren't managed by it which we explicitly told it to ignore.
That sucked.
No harm done, we're still in beta to catch bugs exactly like those, but our initial enthusiasm with terraform has been waning very quickly. Like I said, I love the design, but it's just not there yet.
I've seen BOSH used this way.
You just have to accept that editing a file and checking it into version control is a good thing.
After a while, infrastructure tends to stabilise into a predictable format for the given distributed system under management.
I think Terraform would definitely be easier for a fast-developing system because of its bias towards interactivity. But in the long run most systems cease to be fast-developing.
At Pivotal we're seeing more and more customers going all-in on BOSH. Not just for Cloud Foundry, but for pretty much every stateful service they can lay a hand on. They like that it's based entirely on versionable, auditable files.
> However, we require the state file in order to find the resources we own and then refresh the state of the world so we can do a diff.
BOSH will inspect the world when you ask it to, or you can ask it to do so continuously. As a rule, most operators prefer (like Terraform users) to do it manually.
It's great that you're working on fixing that, but it looks like it will be some time before it's quite there, given the seemingly slow pace of development.
It's nice that Terraform can support untaggable resources, but I don't understand why that's the general case and not a special case. I can't remember an AWS resource that isn't nameable or taggable.
ElasticIP is one among many. For a complete list: http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Using_Tag...
This is how BOSH has been doing it for years for IaaS setup and other systems like Salt/Puppet/Chef/cfengine for servers. Compare-and-repair is a fairly well-accepted, trustworthy model.
Disclaimer: I work for Pivotal, we donate most of the engineering on BOSH.
[0] https://bosh.io/
I feel strongly about two things:
1) I screwed up at the beginning so hard (there were _no_ best practices to be found so it was a bit wild west) that I've come to the point where I need to use something else, or rewrite everything.
2) Terraform is best when you are in a silo by yourself, or you're using something to do the state applies for you (Atlas, or DIY through Jenkins and friends).
It does a lot right that other tools don't do (stringing together resources is a cinch), but it also behaves in strange things and the statefile is both a huge positive, and a huge negative.
If you need to move fast it's not a bad tool to choose at all.
The BIG caveat here is as follows -- If you have a mixed-cloud setup and use AWS, GCE, Azure, DigitalOcean, etc.. at the same time; It's probably hands down the best tool for managing everything.
Notes:
- I only use AWS. I use some of the other TF providers mixed in with it. (Consul, Datadog)
- I started from scratch (no AWS resources -> 100% Terraform), the general opinion seems to be this is the better place to start from.
- I use a majority of Hashicorp's stuff and am quite partial to liking their products. Terraform is something I have the most opinions on. I have no problems strongly suggesting Packer, Consul, Vagrant, and Vault to anyone.
- One of the main reasons I'm considering alternatives is that, at my current position, I'm having trouble getting folks to wrap their heads around Terraform.
- Often beats AWS to support on new features. AWS is traditionally not very good at coordinating service releases between IAM, CloudFormation, and whatever.
- Modules better than nested stacks
- Much better insight into what's going to change. This still seems to be the case with CloudFormation's change sets.. Particularly with nested stacks, which were sent to earth to torment us.
- Nicer configuration than CloudFormation
Woes:
- Juggling multiple state files. This seems to be necessary if you are going to stamp out infrastructure across accounts and regions; this is double true if you want to roll changes gradually across regions/accounts. Prepare to write wrappers.
- Creating modules is the only wait to organize tf files into folders. This can really rub you the wrong way because modules are isolated from the root and other modules' variables. You have to start passing in and out resource references :| There was an import proposal made 7 days ago that looks very nice.
- Not much in the way of rollbacks..
- Potentially less flexible than CloudFormation; CloudFormation offers wait conditions, conditionals, rollbacks, update cancellations(with rollback), etc
+++ Use jetbrains free IDE (IntelliJ, PyCharm...) and install the plugins for HCL support.
You'll get syntax highlighting + auto completion for variable and infrastructure names + visual report on lines with errors and quickfixes.
Now HCL just because the easiest language in the world. (Seriously, it's not even a language, it takes 1 hour to master the syntax, top.)
+++ Use terraform to create all the infrastructure backbone. (VPC, subnets, routing tables, DHCP options, etc... )
Terraform is the single best tool in the world to do that. It's also the only one currently in existence that is actually able to do that. (ansible and the likes are missing critical modules).
+++ Do NOT use terraform to manage instances nor security groups.
These things change wayyy too often and you'll struggle with states.
Use one of the ansible/salt to create and manage them. They refresh their inventory to keep with external changes (at least ansible does) and they allow to run scripts on the resources. (e.g. tag all the EBS volumes of the instance, after the instance is up).
Bonus: For every instance "name" you create, create and put it in a security group with the same "name". Then you can manage security groups by name ;)
+++ Terraform can import existing infrastructure since v0.7 (this release :D)
Just sayin'.
+++ The state files is annoying to share and keep updated
It used to be put in git. (it's still possible).
It can be put in S3 (or similar options). The sharing is faster and it makes less state issues.
+++ Because of the state files, terraform is good for "static things" and terrible for "fast moving things".
That makes terraform appropriate for the VPC/subnet/... provisioning (the first point in this list )
But it's terrible for instances and security groups (another point in the list )
Known your strengths and weaknesses, choose the right tool for the job ;)