I've seen a lot of people start writing their own CF management systems, in Python or whatever, and end up with a lot of infrastructure defined in the logic of the homegrown management tool - like a badly written PHP page where presentation and logic are all mixed up. You see that once, you never want to go there again.
Terraform does not cover all AWS features, but the exceptions are few and tend to be lesser used features. The file format is human readable (this is subjective to some extent), and can be composed and modularized easily.
https://opencredo.com/terraform-infrastructure-design-patter...
https://github.com/hashicorp/best-practices
https://atlas.hashicorp.com/help/intro/use-cases/multiple-en...
https://www.terraform.io/intro/getting-started/modules.html
https://www.terraform.io/docs/state/remote/index.html
I'm not a big fan of buzzwords and soundbites, but Terraform comes pretty close to fulfilling the ideal of "software-defined infrastructure" in a way that is accessible, easy to use, easy to expand, and just makes sense - in the way a good programming language just makes sense.
We'll continue to use Terraform at my day job, but I don't recommend it to colleagues nor clients if they use only AWS resources. Simplicity is the ultimate sophistication, and adding yet another abstraction does not contribute towards that philosophy.
It's really unfortunate you had these experiences and I can't back up the statement when I say that this isn't normal and every release certainly does improve stability in a big way. Terraform is a very, very heavily tested (unit + acceptance) project and we fight very hard against regressions with multiple tests per bug in many cases at different levels.
I'd encourage you to continue giving future versions a shot and if you experience anything like this again, let us know and we'll react to that quickly relative to other types of issues.
I try to pick boring technologies that are battle tested through years of tech sector use; there's no glory for being on the bleeding edge, with only pain when things go south. So goes the ops struggle.
MIGBY, we have also been bitten in production in much the same way. If you are using only AWS resources there is no need to use terraform. Using it at all seems like a questionable decision to me since it's still firmly in "early adopters only" status.
If you're AWS-only then CloudFormation is the way to go. And IMO it will continue to be for a long time, until terraform is much more battle-tested and -hardened. Unless of course you're operating in an environment where you can afford to be exposed to the cutting edge of infrastructure-as-code. I definitely am not.
Personally, I use CloudFormation for my own stuff, because aside from reliability concerns, giving up on cfn-init to not write JSON (when I don't now, I use cfer, a Ruby DSL) is totally no bueno for me. Terraform offers no direct equivalent, though I've had HashiCorp sales people try to push Consul chewing-gum solutions that I'd have to manage much more directly than CFN metadata/cfn-init have been for me. Maybe in a couple years it'll be somewhere where I feel safer recommending it.
They also allow you to manually get out of the UPDATE_ROLLBACK_FAILED state, which was usually where you needed to contact AWS.
Cloudformation also supports Lifecycle events. They are a bit complicated to set up right, but they allow you to tell CloudFormation how to wait and how to proceed in response to resource signals (for instance, the deployment can wait or roll back if new instances in an ASG fail to be deployed). Terraform has none of that. It has "create before destroy", and "depends on", but that's about it AFAIK. If CloudFormation fails it will at least attempt to roll back to a sane state (and usually succeeds). Terraform will give up and leave you to clean up the mess.
Apart from that, Terraform has an advantage because of the modularity, ability to easily bring resources in, reference other resources in remote states etc.
I'm using Terraform now, and I like it much better. It has its own language for defining resources, but I find it very readable and easy to write by hand.
My goal is to have my "infrastructure as code" so it can be tested in a sandbox and changes can be reviewed as pull requests. Going forward Terraform is my choice.
Also, HCL is way more readable than JSON. You can actually do a code review on a terraform plan.
And terraform as a CLI tool is pretty easy to use and expand. On top of this, you can write your own providers fairly easily and it's a more 'standard' tool as in, you don't get vendor lock in.
On top of all this, terraform is an open source Go project: you can easily make your own tools and import it's internal packages to hack things to your needs (caveat: internal packages =P).
CloudFormation is also nice in that it has access to a few API features which Terraform doesn't - specifically wait-conditions, automatically rolling-back updates when there's a problem, and a lot of the magic around UpdatePolicy/rolling deploys (although that doesn't work the way one would expect - a story for another day).
Having said that, Troposphere is pleasantly readable and has nice docs. The cross-platform integrations (for example, CloudFlare) add to the "just works" feeling. Some things are a pain to do in Terraform due to its strictness around being declarative and not having conditionals (e.g. having production and development environments that are similar, but not exactly the same), but there are some well-known hacks around them.
It also takes a very strong position around failed updates: it'll stop mid-change, tell you something's broken, and have you fix it. In the same situation, CloudFormation would roll back the changes to the state you were in before you started. Which one of those two failure modes you prefer is up to you. There are also some issues around having passwords and the like in your statefiles when you use RDS. You can avoid this if you go whole-hog into the HashiCorp stack (TF + Vault + Consul + etc.), but it's bugged me.
Lastly, If you're going to build something big and complex in TF, I suggest you do some reading of Charity Majors' rants on TF (https://charity.wtf/tag/terraform/). She has probably already run into the problem you're going to have.
So which would I use? Depends. If I want to create the same environment in a bunch of places (which also includes blue/green deploys), TF. If I want to do more complex things than that, CloudFormation.
In the end, you choose your toolset and I have made a deliberate decision to stay as close to the provider's bare metal solution and build from there. I can build my own tiny "abstraction" a couple hundred lines of Python generating CF template and parsing the template into readable human format.
working with json can be a pain, but there's lots of great tools out there that output cfn templates from either dsls or nicer formats
Even if your project only expects to use AWS can you say that for certain of all your projects?