Terraform 0.9
hashicorp.com
hashicorp.com
Every place I've seen Terraform used invariably runs into these shortcomings and the workarounds range from using Erb and Jinja templates to generate Terraform templates to just ditching it entirely and using home-grown solutions. Terraform should have been a library in an actual programming language instead of a gimped external DSL that re-invents the wheel.
Compilers are like, useful.
0: https://www.hashicorp.com/blog/terraform-0-8/#conditional
As for DSL vs. programming languages, the halting problem provides a pretty compelling argument in favor of limiting your system if you want to have "verifiable convergence". Personally I'm a huge fan of HCL too. It fixes a lot of what's bad about YAML and TOML and has a built in pretty printer.
I pretty much agree with you, they should of provided 2 entrypoints IMHO, they can make there DSL and all that and make it (eventually) turing complete and at the same time they should just provide composeable objects in <some> language as well.
I've never quite figured out who the target audience is for these kind of DSL(s), is it people that can't program, or people that can't program well/at all (without adult supervision)?
They all suck in equal measure, but each in a different fashion. You can only try to pick the one tool that sucks the least for your particular use-case and hope that two years down the line you made the right choice.
As far as I'm concerned Chef is the right way to do configuration management. Everyone else with their YAML DSLs is doing it wrong. Expressing dependencies between resources does not require specialized external DSLs and sometimes you need to do imperative things to get to the end which is either impossible or just horribly convoluted with any tool that tries to hide complexity behind a markup/serialization language.
Several human centuries of effort have been spent on improving these things (syntax highlighting, auto-complete, stack traces, error messages, interactive debuggers, etc.) and then all these DSL authors decide to just chuck it all out. Boggles the mind.
More than a few times Terraform has wrapped an error message from AWS in the hopes of being helpful while hiding all the details of the actual error message and no hints how to rectify things in the DSL.
e.g.
https://plugins.jetbrains.com/plugin/7808-hcl-language-suppo...
The things that Terraform solves for, namely keeping environment state and infrastructure relationships, are really hard to do with other CM tools without diving deep i.e. writing python or ruby
You might consider those things advanced. I consider it the bare minimum to get proper work done.
It's markup. It is meant to be easily readable by devs and non-devs alike while being reasonably capable.
The beauty of these tools is that they are written and be expanded upon with popular languages. If I wanted to make Terraform deploy Dell servers, all I have to do is write a plug-in with Go. What's more is that I don't have to teach a TF user anything new since its usage pattern will be the same as any other TF plug-in.
You must not have seen many TF deployments. I've seen several and each one is a unique snowflake. Some use Jinja, some use Makesfiles, some use Erb, some just use heroic copy-pasting. There are no conventions or standards whatsoever. I don't know what you mean by readable either. When was the last time you navigated 20 Terraform modules while mentally filling in all the variables and thought to yourself "Ya, this is totally readable".
Who are the non-devs reading infrastructure code? I think the target market for Terriform (probably) understands the basics of programming languages like loops, conditions, and variables.
I see your point about plugins/expansion. I think infrastructure hasn't gotten past the "write a plain file as an API" stage, so writing a library to generate those files feels like hacking on top of an incomplete platform.
If I could use puppet to manage openstack instances at work, I'd use it over Terraform in a heartbeat because the foundation is much more solid, but unfortunately such a provider does not exist.
I use terraform and want to like it, but every bug and weird behaviour and nonsensical limitation is making it harder and harder. Furthermore, these aren't really fixable implementation issues. The terraform language needs a complete redesign.
You don't need explicit loops and branching to do the kind of declarative programming terraform needs, but being able to use basic datastructures besides strings would be very useful.
Instead of if statements, pattern matching constructs would probably suffice for avoiding some of the concerns with avoiding loops and branching (limits and predictability upon the DAG and resulting state machine being generated as I understand it).
In any case, I've found plenty of power by using custom data providers that use whatever logic I want and feed that to Terraform providers and modules to instantiate. I think the level of maturity with the Terraform community is still in the early stages and emphasis will shift to data provider based constructs for orchestration.
I think part of your issue with Terraform is that you want it to do everything, when you should just use it to get the foundation up and then use a provisioner with more power to do the finer detail work to fit your needs. Also, lets remember it's not even 1.0 software at this point. We started playing around with it around 0.6 and things have come along way since then. Maybe it's just not the tool for your needs.
In fact if Terraform was a library it would not do everything nor do I want it to do everything. At the end of the day Terraform traverses a dependency graph and generates a sequence of commands to run. The traversal is basically a topological sort and the command generation is a bunch of API calls. None of those things require a specialized DSL.
Given a list:
[ A, B, ... N ]
Create a list of objects: [
{ foo: const, bar: A },
{ foo: const, bar: B },
{ foo: const, bar: C },
...,
{ foo: const, bar: N }
]
You say, why, this complicated programming task is not a job for terraform. Then why are some elements in place of the dsl functions, and some missing?I have a feeling that the creators of may cloud tools don't use the cloud, or only to have pet infrastructures instead of pet servers. To set up something automated, complicated, scalable, reproducible, you have to result to do it yourself, while these tools are supposed to do it for you.
I feel like ops people are writing these tools for ops people, but instead of good old bash, now in Go. What is the problem with this?
* They reinvent well known patterns, but call them differently.
* They disregard the common knowledge of software development tradecraft.
* They don't get these things right for the first (or second time), just redo their previous workflows, and we are still where we were 10 years ago, just with different tools.
I personally prefer Ansible, but it has its weak points as well, and its development seems to have slowed since RedHat acquired it. The main advantages of Ansible are extensibility and reuseability. I can write custom modules, and have a descriptive DSL for stuff not thought of for the upstream devs. I can also simply reuse parts, which is a very weak point in terraform.
Still we use terraform, for it has its own merits, but IMHO if you want to do something, either do it, or don't do it, but don't do a half-assed "solution" which promises much, and fails to deliver. This is my feeling with terraform, where I could not even create a simple mapping of values, if those are not strings! (I know, "ops" people love strings. Only those pesky "devs" love structured data)
I am expecting the Renaissance you mentioned in devops shortly, just like how React transformed the scene.
In short, DSLs should be inside programming languages and we shouldn't be afraid to write code.
It's like you said; no loops, no proper variables etc.
Terraform itself though is fantastic and powerful. It's lacking some pretty important stuff still (for example, not having to refresh the entire infrastructure state when you change one single variable for one single tiny little service), but that just makes it an alpha. It's still better than the alternatives by a mile and then some.
But, yeah. Fuck HCL. The good news though is that I think this is "easily" fixable by someone who would really work on it: Because the underlying environment is 100% json, it's easy enough to generate; so you could implement a saner environment for specifying the state and feed that to terraform. That's enough for me to bet my infrastructure on it: knowing that today's issues are fixable, rather than the effort being doomed from the start.
If your looking to dive in, I wrote a short introduction blog post on getting started with Google Compute Engine.
https://blog.elasticbyte.net/getting-started-with-terraform-...
- The new 'State environment' feature should resolve the issues discussed here: https://charity.wtf/2016/03/30/terraform-vpc-and-why-you-wan...
- The new locking feature means we don't need to use https://github.com/gruntwork-io/terragrunt
It only took me about 30 minutes to setup a terraform file that could bring up and teardown an entire web stack with a single command. I was so shocked by how straightforward it was that I brought the stack up and down a few times just to make sure it was actually repeatable.
However, when things went wrong we ended up with cryptic error messages and often when `terraform up` failed, then `terraform destroy` would fail too, leaving us using AWS console to jump in and start clearing up the resources. Particularly painful because AWS has no awareness that your stack came from a Terraform config, so you have to navigate through the subresource menus destroying things one-by-one.
We ended up switching over to Cloudformation. The extra verbosity does suck, but the tooling for running/updating stacks feels far safer. We can review the changeset when we update the stack, view a realtime event log as resources are created, deleted or updated, and best of all we can always tear the stack down from a single point.
It's allowed us to do the following for our developers:
Full integration with the CI/CD process.
Open a branch, commit it to GitHub, add a special label and you will get a completely new setup, just to test your branch. New application name in the service discovery tool (Consul/Eureka), new RDS instances, new network, etc, etc. Once happy, the developers can merge into dev and the branched version of the infrastructure is destroyed.
If you want a great starting point for learning how to do things the right way https://github.com/hashicorp/best-practices/tree/master/terr... helped us out a lot.
I'm launching a venture that requires deploying and maintaining cloud infrastructure across the three major cloud providers (AWS, GCP, Azure) and am considering options. Most of our infrastructure is kubernetes but Fugue caught my eye for their AWS deployments. Curious if anyone has used both.
But in another place we decided to use troposphere[0] and I liked it more (at least for AWS-only shops). A lot more flexibility and you're relying on Cloudformation to do some of the heavy lifting around states. We basically ended up with a tool based on that library which would generate deployable CF stacks (and some other things).
Given the inputs you stated, I am full CF because I don't want to manage state files, and I want some more advanced things like Deletion policies and automatic tagging of resources with the stack info. We use these tags in a lot of lookup scripts when operating on the infrastructure.
Unfortunately when you inevitably make changes to your CF config, `terraform plan` doesn't help because it's likely you changed something inside the `template_body` and Terraform just knows the string has changed and doesn't tell you what part of the config changed. Parameterising as much as you can helps.
The primary advantage Cloudformation has for myself is a lack of need to manage the state that Terraform spends so much effort upon. In fact, Cloudformation does this fairly opaquely but was much more obvious when I tried to deploy a CF stack when S3 was down a while ago - being able to use an alternate state file location would have been handy then.
Some of the things I like about it are that it allows me to have multiple environments (Prod, Dev, QA, etc.) where each stack can be slightly different but still related. It also has global variables (ie. office IP address) that allow you to change something in one place and have it be updated everywhere. Finally, AWS resources are defined in Handlebars partials it's easy to add new ones or maintain existing ones as AWS adds more functionality.
It's still a work in progress and I updated it almost every week. If you like working with NodeJS I'm always looking for help. :)
I've got to say, I always end up falling in love with products with cool names, nice logos, great design and easy interfaces. Terraform and most of the Hashicorps products I've seen tick all of those boxes. I'm amazed at the quality of the tools.
However a question:
The documentation makes it really clear that Terraform is not a configuration management tool. Do I understand this correctly that it means you need another tool like Chef to run all base setup code, like install firewalls, set timezone, install other packages? I've been using Ansible to install a baseline of packages for my servers.
We have been doing "state environments" using different var files and different state files for a while now, while keeping the same terraform files.
For instance, say you're sharding across three machines. And now you want to up that to four. Each instance of the application is smart enough to do it provided it knows which shard it's in.
We deploy fully baked AMIs containing an application. Do you make separate AMIs per shard? Or do you supply configuration afterwards? Or do you template and generate your plan file each time?
I haven't found anything on this so I'm hoping someone can tell me.
What you could do is generate a different packer configuration for each of those types (guessing you use packer since you mentioned AMIs), which will create a different image for each machine type. You can scale each type individually.
Happy to help if you need specific advice.
Do people usually just build k AMIs if they have k partition groups that they want to serve?
For a sort of simplistic version of the problem, imagine the data was splittable into 6 partitions and you hashmod the partitions across 3 groups. So instances in group 1 will serve partitions 0,4; instances in group 2 will server partitions 1,5; and so on. Does one build an AMI for group 1, one for group 2, and one for group 3? Does one provide configuration informing which group a specific node should serve through other means? Or does one stick that in a Terraform definition?
Generally though, if there is some custom config data that must be passed on BOOTSTRAPPING the node, you can use templated values for user-data on cloud-init based systems, or the remote_exec provisioner. This can even include the modified shard name (with count). If you're talking about updates AFTER the machine has been in service, I do not believe this is 'officially' supported by Terraform but it can be done [0]. You may want to look into a proper CM tool w/ idempotency for that sort of thing (Chef, Salt, etc.)
0: http://stackoverflow.com/questions/37865979/terraform-how-to...
OT but what I really want is a service where I just click a button and they configure the server for me with automated security updates.
I'd first ask what makes you want to move out of Heroku, especially if you have no sysadmin experience at all.
Re OT: you should really check out https://www.engineyard.com/, which provides a layer of services on top of AWS. You can handle a Rails or Node application with little experience and have security updates (Disclaimer: I'm a happy customer, but again I use many different providers so I don't think I'm too biased).
See my other comment for reasons why I want to move away from Heroku. I can expand on the cycling bit though:
I keep my game's state in memory all the time and really just want to write to the database once in a while. The random cycling prevents this since it just terminates the app without any warning so any state is lost.
Not really.
What Terraform does is to allow you to put all your configuration in (versioned) code. So you can have all your existing configuration, and then change the relevant definitions to Digital Ocean or Linode and make the commit 'Move to DO/Linode'.
Exactly how its configured - security or otherwise - is still up to you though.
> OT but what I really want is a service where I just click a button and they configure the server for me with automated security updates.
Why are you moving from Heroku to Digital Ocean then? Heroku is much closer to, if not exactly that.
It basically comes down to the daily cycling which I have no control over and the low memory on the hobby dynos. My game is constantly reaching the limit and without rebuilding the whole architecture it leaves me with either upgrading to a much more expensive dyno or migrating away.
I also miss just being able to ssh into my machine and do stuff.
This is all anecdotal but at work we recently switched from marketing sites/blogs on internal infrastructure to farming them out to DO. The biggest concern with the person who manages our fleet is the security implications of handling multiple machines outside our network. In just a very minimal amount of Google fu I found a plethora of management services that do the work you're talking about on your VPS infrastructure. Some specialize in just DO but a large number are relatively cloud agnostic.
We have also just recently invested in using Laravel Forge for provisioning and it's relatively trivial to write a script that will run 'sudo apt-get update' but proper security usually requires a bit more than that. Package upgrades to major versions can sometimes cause quite a bit of havok. Having a service or team of individuals that can keep this brain power to handle your case properly is often more than worth the price to someone like me. Finding the right one is its own chore but from my limited search there's quite a lot out there to cover just how hands off you want to be.
Anyone actually know where this product would fall in the development pipeline? It sounds amazing.
Edit: Okay, got it, it's a configuration management application which only applies if you're manually handling your AWS containers.
Terraform allows you to manage VPC, network, routing table , firewalls and gateways in the cloud. There is no other software to do that (only manual click in the AWS website).
It can also create AMI, instances and auto scaling groups but there are many tools in that field that do it better.
Terraform is a tool that captures Infrastructure as Code. Flynn appears to be a PaaS, akin to something like Heroku with scheduling capabilities on top. Under the hood, there is some sort of infrastructure that Flynn is deploying to; it's providing an API on top of it to easily manage deployments and scaling. Terraform is the tool that defines the layout and composition of underlying infrastructure resources.
Another one of Terraform's strengths is that it supports multiple providers. Most of this thread mentions AWS, but Terraform also supports GCP, Azure, and DigitalOcean. But Terraform can manage anything with an API, so you can codify things like GitHub teams and permissions, DNS configuration, PagerDuty escalation policies, and so much more (https://www.terraform.io/docs/providers/index.html).
All of these definitions are captured in code. You can bring that code under management of source control to get peer review, pull requests, or permission models (signoffs) too.
If you are an individual developer happy with Heroku or Flynn, Terraform probably isn't the best choice for you, especially if you do not have a background in operations or you do not want to manage servers. However, most medium->large scale production applications cannot be supported by such technologies and demand a tool like Terraform to manage the complexity.
I hope that clears up the differences a bit.
Disclaimer: I work for HashiCorp, the company that makes Terraform.
https://medium.com/@zwischenzugs/dynamic-terraform-environme...
is there a 'proper' way to create dynamic environments?
The best time to look into something like Terraform is if you begin managing more than just AWS resources. Even so, you can choose to use CF for AWS and TF for non-AWS. Or, if you want to start taking advantage of the higher-level features Terraform is beginning to support.
At the end of the day though, I'm a big fan of working with what you're most expert with until circumstances require otherwise.
the biggest win for TF (IMO) is that it doesn't lock you into any provider. It is a lot easier to design TF configs to work in cross-cloud scenarios (well, "easy") than it is to port CF over to something else (which is AWS only.