Creating my personal cloud with HashiCorp
cgamesplay.com
cgamesplay.com
Often I hear from enterprises that Terraform is cloud agnostic, but that's often very wrong. Terraform modules are still specific to the cloud platform and a rewrite is required to port an app running on AWS to GCP.
If you use AWS, you're probably better off to use AWS Cloudformation and for GCP Google Cloud Deployment manager.
A business reason is often that the engineers are more familiar with Terraform already, but learning Cloudformation is really not that hard...And if you can't work out how to use a basic tool, you should not be running infrastructure anyway, because IT is a constant cycle of change.
I'm not saying Terraform is not good, but I just think a native solution of the platform is preferable over a 3rd party tool.
Comparing CloudFormation to Terraform is a little like comparing HTML and CSS to Javascript (though Terraform isn't nearly as nice to code in as Javascript -- and I'm not exactly a big fan of Javascript). You can cover most use cases with plain HTML and CSS but the moment you need to get a little more intelligent with your code you get stuck.
I think that's what AWS CDK[0] and Terraform's CDKTF[1] are trying to solve.
Given the context of your example, I'd liken Terraform to CSS and CloudFormation to HTML; CDK/TF to Javascript. It's not a great analogy, but Terraform as is right now is just close enough to a programming language to deceive you into treating it like one. But it really isn't and those issues become glaringly clear the more you use it.
I'm yet to try Hashicorps CDKTF but from what I've used of CDK it felt like the audience was a little different to those that would use TF. CDK feels more for orgs that would have the same team who write the application code (eg lambdas) also write the infra, a bit like Serverless (sls) is. Whereas Terraform tends to be more suited for orgs that like different teams managing infra from those managing application development. Generally speaking of course.
Ultimately all the above tools work fine for production systems so its often just a question of preference.
I'm definitely in the latter camp, which is something I find frustrating. I get that for a developer the syntax familiarity might make CDK easier, but for me as a non-developer the pain of groping around the terrible documentation and learning how classes are supposed to be used far outweighs the annoyance of fixing YAML indentation.
Ultimately I worry people are jumping on the "true IaC" bandwagon without acknowledging that if their infrastructure is supposed to be somewhat static and immutable, a declarative language might actually be better.
Also the API gateway configuration is generated based on registered endpoints within a larger application, but that's not being done directly in the CDK, but as a separate explicit step in the Makefile as a dependency of the CDK targets.
Disclaimer: I work for AWS, but not on anything related to this, other than using it.
My issue with anything that ultimately compiles to a JSON or YAML state/config is that you lose the dependency tree (or your dependency tree becomes rigidly defined by the way the transpiler converts your code into JSON). It causes so many problems on any large project that ultimately the only solution is to break the project up into smaller distinct projects within the same git repository. Which is basically the same solution to working with JSON/YAML CF stacks directly.
If someone can create a language (or SDK for $PROGRAMMING_LANG) that then worked with AWS APIs directly, (like Terraform), didn't just transpile back to JSON, and isn't as verbose as using boto3 directly -- well I think that product might stand a real chance of displacing Terraform.
I've got a lot of strong thoughts on how this could be done right so I did consider having an attempt at it myself using the shell scripting language I'd designed as a rough foundation. But having a full time job already, two young kids, and a growing in popularity open source project (namely my aforementioned $SHELL), I realised any Terraform competitor I did create would be doomed to either never being maintained, or the project would literally burn me out.
But I do think that it would be really interesting to build an IaC project, and you’re right that it probably would be doomed. :)
So what is stopping you from using "a proper programming language" to generate the json/yaml cloudformation template?
This is what you see in GCP docs from day one. On AWS, they brainwashed everyone in this corner of "static template with parameters", so that you can "reuse" a template to build your custom stack. It's great for "look what I can do, mom" (but I have no idea what it's doing) but nobody sane would ever trust a 1-km long yaml/json and deploy it. So if you anyway have to inspect it, why not make it easy to inspect? Split into modules, add docs, etc = code to run.
I have no idea how we switched from random scripts to "reusable" random scripts (ansible &co) to random static configuration and then the cherry on top: "reusable" random static configuration. Insane. Abstractions on top of abstractions.
CDK is on the right track, but even there it's a mess, again for the sake of hiding complexity: constructs and deployment. Where did One thing well and Keep it simple stupid go? :)
Nothing and there are already products out there that offer that. However I think the issue will always fall back to the problem that you're compiling from a functional or procedural language down to a dumb data interchange format. That can cause a variety of issues such as losing your carefully crafted order of execution.
> This is what you see in GCP docs from day one. On AWS, they brainwashed everyone in this corner of "static template with parameters", so that you can "reuse" a template to build your custom stack. It's great for "look what I can do, mom" (but I have no idea what it's doing) but nobody sane would ever trust a 1-km long yaml/json and deploy it. So if you anyway have to inspect it, why not make it easy to inspect? Split into modules, add docs, etc = code to run.
I don't think anyone was "brainwashed" by CloudFormation and the solution you describe is exactly the approach Terraform takes.
> I have no idea how we switched from random scripts to "reusable" random scripts (ansible &co) to random static configuration and then the cherry on top: "reusable" random static configuration. Insane. Abstractions on top of abstractions.
This I agree with. It's not just AWS though, you see YAML-based config all over the place from CI/CD pipelines to Kubernetes pods. And they all suffer from the same problems. It's easily my least favourite thing about the DevOps modernisation of what would have been random sysadmin shell scripts 10 years ago. Frankly I'm not convinced these YAML files are any more readable nor less brittle than the duct tape we wrote in #!/bin/sh before.
> CDK is on the right track, but even there it's a mess, again for the sake of hiding complexity: constructs and deployment. Where did One thing well and Keep it simple stupid go? :)
As I'd written elsewhere, I think CDK is aimed at a subtly different audience. CF, TF, etc are more focused around the sysadmin side of DevOps, whereas CDK are more focused around the application developers end of the DevOps tool chain. That's not to say sysadmin guys can't use CDK, but rather than CDK doesn't just aim to deploy infra, it embeds and deploys the serverless applications (like lambda code) as well. It's more akin to the "full stack" style of developers and not every team nor individual who's job it is to manage cloud infra is going to be an application developer. Particularly in larger organisations. So there definitely is a place for cloud infra stacks to be described in less sophisticated languages (even if it doesn't appeal to you and I personally).
Anecdotally: non-devops developers find it easier to edit Terraform than CloudFormation. They understand the resource and data blocks much easier than YAML.
Terraform bridges the gap between purely YAML CloudFormation and purely JavaScript API calls. (JavaScript used as an example there, it could be any language).
Terraform makes tradeoffs to satisfy its problem domain and be easier to use than CloudFormation. Plus Terraform often gets features faster than CloudFormation because cf seems to lag behind the AWS API.
I'd like to write some of my simpler Terraform stuff in AWS CDK, Pulumi, or something similar, to see if I can make use of more language features like inheritance or better decision logic. (currently I try to keep as much logic outside of Terraform as possible. It's supposed to be declarative and not have "real" coding language features shoved into it)
I think terraform was better before, but now CloudFormation for me is good enough, so I don't see the point of Terraform.
Cloudformation also has Drift detection now and I really dislike Terraforms state files.
So terraform is critical for our bootstrap procedure, for documenting our configuration.
Terraform offers far more constructs than CloudFormation and there is still a lot to be said for using the same language from describing your AWS infra as GCP, Github, any on prem infra, etc (even if the resources differ between providers requiring bespoke code for each). It's a bit like those who advocate node.js because it means the same developers can right frontend and backend code and in the same language.
If you like the tooling around CloudFormation in AWS then you're better off with serverless (`sls`) or even Amazon's own CDK (https://aws.amazon.com/cdk/) over YAML-based CloudFormation stacks in my opinion.
That's not to say I don't think Terraform doesn't have its warts: properties are non-guessable, IDE integration is pretty mediocre, and it's overly verbose in calling modules (so bad that sometimes the calling code has just as many lines as the module itself!) but I do still think it is the least worst tool available at the moment. And one can always use a 3rd party tools like Terragrunt if you want to fix some of the shortcomings of Terraform while still taking advantage of it's benefits (though personally I'm on the fence about whether the industry really needs yet another transpiler that compiles to code that needs to be transpiled....it's starting to feel like it's just abstractions all the way down....)
1. It's slow to return results
2. sometimes the only way you can generate a list of available properties is to type the first character, which is really annoying if you don't even know what the first character might be
3. it offers suggestions for properties that aren't even valid for that resource
4. there is no description against any of the properties so you're still left guessing (or having the docs open in another window) anyway
Each clouds SDK in the language the team is most familiar with is by far the best option.
State can be stored in git. Any version of my infrastructure is a git checkout away.
I use Go, and the documentation for the AWS SDK includes copy-paste examples
Try and checkout Terraform from 6 months ago and run it? Frequently I cannot even get someone’s tutorial example written a week prior to work without edits.
I checked out a year old commit in my infra repo; rebuilt an entire ECS stack deprecated 18 months ago.
Just be programmers. The cloud ops scene is just reselling the same delusions as Unix grey beards and Windows server admins. It’s about making hardware do the right thing, not hand wavy semantics.
Most of the people I work with just regurgitate memes. Very few actually test them for truth.
Creating resources with the SDK is easy—keeping them in the desired state is very hard.
Instead, we use a real language to generate a description of what we want (in YAML or HCL) and a reconciliation engine takes over from there.
I don't like the idea of having a Turing complete language to do that. I would prefer things to be declarative. Honestly I would just be happy with a better terraform without yaml and better tooling for terraform.
I tried to use pulumi and they didn't support something I needed, or maybe I was too dumb to understand how to do it, so I insta quit.
They're also pretty pushy in selling whatever they're selling.
I've worked in a couple organizations that had a lot of success managing things like cloud access security brokers, web application firewall appliances, ticketing systems, identity providers (with and without multi-IdP federation) and more.
Doing this without Terraform is incredibly manual and requires even more manual process to keep these systems in sync with your cloud. Having a common automation framework that can manage this is indispensable--and useful even if you're not using Terraform to manage your core infrastructure.
However, my biggest issue with Terraform is that the promise of dependency graphs is in practice broken. Providers will break once the underlying resources are not yet there. The hack to solve this is having several Terraform directories which you run after each other.
Still, I think multiple clouds is the way forward. You can negotiate prices down and use services are best suited for the projects. Especially with other players such as CloudFlare and Backblaze, which already have Terraform providers available.
Want to use Cloudflare with AWS and another 3rd party provider that offers services in an AWS region? Simple with Terraform.
However, since then I gave Terraform another shot and dang am I glad that I did. It is fast, easy for me to get started with. I would much rather waste effort and time on Terraform which is cloud agnostic than spend a bunch of time learning something specific to AWS
I would agree if you have your green-field approach and can commit to a single platform.
From my current experience I use AWS, Azure some other SaaS hosted products like ElasticSearch, Instana, Opsgenie, Kubernetes, databases, Grafana, Prometheus. And with all these products you have a bunch of people in the companry which specialize in their domain and can't know all the tools in all details but have to talk to each other.
So what makes terraform so special in my case is that you can streamline the interaction between multiple teams by focusing on defining well-known interfaces between those teams. The interfaces in the case of terraform would be:
variables (inputs)
outputs
Or you can have specialized teams, which will offer terraform modules for other teams to use.
So for me terraform does not have to be agnostic as this is not the point of it. The benefit of terraform is to streamline interactions (inputs,outputs) for your needs. Teams could automate their things with a python script for example, and use terraform just as a "hull" to offer a way to pass inputs,args to you python script and report back some outputs.
What a Dockerfile did, was to establish a well-known interface on how to define what a container is, by giving a standard-way on how to declare a "CMD,ENTRYPOINT,PORT,etc."
And terraform in that sense gives you a standard on how to define your inputs,outputs when you build,configure infrastructure,Saas, etc.
Also all the major clouds (even the minor ones) have pretty strong first class support for terraform.
I had a job interview where someone asked me what I would prefer Terraform or Cloudformation for AWS.
I said Cloudformation because it's managed by AWS who writes the actual software as well. And they kind of smugly said Terraform is better because it's cloud agnostic.
I was thinking...have you ever USED Terraform.
Even AWS has come up with a CDK to avoid building CF and use real prog langs.
I've never used it though.
I just personally think that if youre going to be a part of an ecosystem, things go much more smoothly when you stay in that ecosystem as much as possible.
Also Amazon is a massive company compared to hashicorp so I feel more comfortable about my infrastructures longevity with AWS tools.
Not saying hashicorp is going anywhere anytime soon and if it did the open-source community would probably take over, but it eliminates that tiny tiny risk.
I would probably use the AWS-CDK now instead of CF.
https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
I'm quite late to respond here, but just wanted to clarify one thing: Terraform is WORKFLOW agnostic, not TECHNOLOGY agnostic. This is a key part of our product philosophy that we make the 1st element of our Tao: https://www.hashicorp.com/tao-of-hashicorp
I've talked about this more with more references in this tweet: https://twitter.com/mitchellh/status/1078682765963350016
I don't think we've ever claimed cloud portability through "write once run anywhere;" that isn't our marketing or sales pitch and if we ever did make that claim please let me know and I'll poke some teams to correct it. Our pitch is always to just learn one workflow/tool and use it everywhere, but you explicitly WILL rewrite cloud-specific modules/code/etc.
With Terraform, the big win for folks is learning how to write and use Terraform, then knowing a fully supported official tool (to some extent) for hundreds of API-driveable systems. Instead of educating an engineer on CloudFormation, Azure ARM, etc. they learn ONE tool and ONE syntax and then adapt that to their cloud-specific knowledge.
More details are in the tweet I mentioned above, but I hope that helps. Fully respect you not being a fan of Terraform, I don't mind that, I just wanted to make sure for yourself and others reading that it is clear that we also don't believe Terraform is cloud agnostic in the sense you described.
Why not provide a cloud host tier for startups akin to CLoudflare Pages?
I also think sometimes why i don't like it very much and how i would make it different.
How the state is handled, including potential secrets in it, is just frustrating. Having root secrets for your whole setup exposed/unsecure is bad. The state is relativly fragile and cumbersome to clean up or fix. I also can't grasp that tf even needs a state and the cloud providers can't return the current state just fast enough. Only a lightweight cache would then be needed.
And probably due to implementation details, plans show sometimes changes when there would be no changes necessary.
For me its a good tradeoff to use terraform for setting up a k8s environment and then handling everything with ArgoCD.
Google Connector is a very great thought: you create a k8s resource and the cloud provider executes it for you on their cloud. No terraform needed anymore at all.
YAML is not a programming language, and any attempt to turn it into a programming language is fraught with pain, woe, and suffering — at least on the part of the poor suckers who are tricked into using it.
Give me a real programming language, with interfaces to the appropriate constructs necessary to do the job, via an SDK. In AWS, that means using CDK, not CloudFormation.
Feel free to argue over whatever aspects of TerraForm that you don’t like, but please, for the love of ${DEITY}, please do not recommend CloudFormation as your suggested alternative.
To bw fair, while CF can be written in YAML for less visual noise, its fundamentally JSON.
Also, its not a programming language, its a target-state description language.
> Give me a real programming language, with interfaces to the appropriate constructs necessary to do the job, via an SDK. In AWS, that means using CDK, not CloudFormation.
Technically,“CloudFormation via CDK”. You can't do “CDK, not CloudFormation” because CDK is just a tool for generating CloudFormation.
In the end I went with docker stack + machine + swarm + docker flow proxy and it's been pretty smooth sailing.
I don't need anything else installed on the machine, I can develop stacks locally with docker compose, deploy them to swarm and I can even add replication on multiple machines for services that need it. I manage secrets with stackoverflow's blackbox and the secrets live in the repository encrypted. I can easily run commands on the local or remote stack via docker machine. I cheated with cron jobs and I just synchronise /etc/cron on all hosts with a directory in my repo. The commands are mostly running commands inside containers.
I have one stack for redis, postgres, docker registry and multiple stacks for other applications. The flow to deploying stuff is basically: generate id from git hash, build docker image, push to self hosted registry, stack deploy with the git hash as version.
A few caveats I found:
- Dependencies with docker health checks will take the duration of the health checks to become available to services needing them. Either you skip health checks or setup things so that infra is running before your services.
- In order to use docker-machine on multiple machines / with different contributors you need to export the key used to setup the machine and export it (there is a npm package that works very well, but it would be nice to have it natively).
- You have to cleanup the docker registry manually every once and then (I had to write a cron job for that)
- I'm unsure about the future of docker, albeit things work pretty well as they are.
https://gitlab.com/stavros/harbormaster
All it does is manage Compose applications, with a sane directory structure. It's been working great, both for my personal use and for a few companies running production workloads on it.
I love that it's super simple and the workflow it has is fantastic, I just push to a repo and everything else happens automatically.
Can you go into a bit more detail about your setup?
Honestly, if I had stuck with Compose files then Harbormaster looks like it's a reasonable way to manage them. But I felt like everything I wanted to do was just a little bit more difficult when I was using them, especially after comparing to Nomad.
The biggest issue I've encountered on it is when you move out of AWS (and you can't apply EC2-based IAM) but still have S3-hosted artifacts. Specifically it cannot receive Vault secrets for the artifact credentials because the nomad templates get applied at a much later stage.
https://github.com/hashicorp/nomad/issues/3854
I've used an nginx-based S3 proxy in the past to get around this. Not ideal but it works.
What I particularly like is that I have a mixture of Docker containers, VMs, and LXD containers all centrally managed.
Overall I found nomad to be fairly intuitive, and the single binary/single job paradigm of the Hashicorp stack is very appealing.
This was a much-needed reminder that the target audience for my documentation is my future-self.
In my lab I'm actually running a two-node cluster, but that's 'even worse' and engenders the occasional mildly surprising failure states.
Anyway, I can highly recommend setting up Nomad (and some friends) on a single host. It's going to be much more robust & interesting than 'some janky shell scripts'. : )
So, that said, yes, server (whitebox) plus desktop (xeon, 32gb) - both running Debian, with Consul, Nomad, Traefik, promtail etc running on both, but no shared NFS between the two, so I've got constraints on most of the important stuff to run on the server (prometheus, cortex, loki, nodered). In practice, desktop is running 24/7, just it occasionally gets a reboot.
Having consul/nomad running as server-client on two machines is undeniably weird, and requires some careful consideration around bootstrap_expect= settings.
Almost all of this is around having a useful facsimile of my work environment, rather than (say) running an SSG public site.
https://www.nomadproject.io/docs/drivers/qemu
https://kubevirt.io/user-guide/virtual_machines/disks_and_vo...
Is terraform generally used to deploy workloads to nomad instead of writing tasks directly?
It can be, although it has some weird shortcomings. For example, if the job is already present in Nomad but not running ("dead"), I don't think you can use the terraform provider to start it again.
HashiCorp themselves suggest using terraform to provision the base nomad system (ACL, quotas, base system jobs), but perhaps not your actual applications:
> This can be used to initialize your cluster with system jobs, common services, and more. In day to day Nomad use it is common for developers to submit jobs to Nomad directly, such as for general app deployment. In addition to these apps, a Nomad cluster often runs core system services that are ideally setup during infrastructure creation. This resource is ideal for the latter type of job, but can be used to manage any job within Nomad.
https://registry.terraform.io/providers/hashicorp/nomad/late...
In my work helping companies with nomad, I've seen jobs run a few different ways:
* Write and submit HCL jobs directly to nomad
* Terraform templating and the nomad provider (as above)
* https://github.com/hashicorp/levant
* A DIY templating thing (e.g. python and jinja2 templates)
* A webapp that submits jobs as JSON directly to the nomad API, perhaps modifying it to match certain policies (kind of like k8s validating / mutating admission webhooks)
There's also a new thing similar to Helm: https://github.com/hashicorp/nomad-pack