For a small example, I _still_ get +1 notifications on this critical issue nearly every day: https://github.com/mitchellh/packer/issues/409 - no response from dev team whatsoever.
For a small example, I _still_ get +1 notifications on this critical issue nearly every day: https://github.com/mitchellh/packer/issues/409 - no response from dev team whatsoever.
My experience.
Windows guests: buggy
FreeBSD guests: buggy
Ansible provisioning: buggy
Windows support was incorporated from a community plugin. Now it's in limbo maintained neither by the original creators nor the (seemingly) lone dev from hashicorp assigned to vagrant.
They use DRM to protect their closed source proprietary plugins and it makes it a nightmare to build a version from source that works with said plugins. Building the installer is closed source for some reason.
After finally being able to run with modified source, I was able to monkey patch their paid VMWare plugin to add a capability so I could use private_network with static IP on Windows guests . I posted the monkey patch to github issues and it took two weeks for their paid plugin to be updated.
I wanted to love and contribute to vagrant, but the somewhat dishonest docs, unresponsive dev team, and DRM have turned me into a grudging user only because there are no alternatives.
Vagrant is definitely a wolf in sheepskin in terms of appearing open source and open to community collaboration.
(Having wrote a multihypervisor disk-cloner in Ruby.)
Consul is...fine, I guess? It's not particularly useful for me, but that's because I have a Zookeeper background and am comfortable with its primitives. I'd use it if I had a need for it. OTOH, I've had extremely bad experiences with Terraform, both in terms of trying to build tooling around it and in having it hose states when resources fail to create correctly (I've mentioned previously around here that multiple clients independently invented the term "terrafucked" for the results of their Terraform state) and I don't feel like I can trust the tool with something as critical as my cloud infrastructural resources. I might bite the bullet and use it if I were working on Google Cloud or Azure, but in AWS, CloudFormation (with tooling on top) is sufficient.
Disclaimer: I am the founder and CEO of Boxfuse
If I'm wrong, please let me know how I can write, as I would using Chef and Packer, a set of directives to install and prep Zookeeper, then discover other Zookeeper nodes in my cluster at runtime, while remaining aware of nodes that are replaced due to system failure. (If the explanation exists and involves "our proprietary, closed-source agent," or "our open-source agent that talks to a proprietary backend", you don't have an answer.)
As for the notion of "immutable" servers, this industry term means servers that aren't updated in place, not servers without a read-write file system or read-write memory.
In the case of service discovery with client-side load balancing you can easily integrate client libs for services like Eureka or Zookeeper directly in your JVM application, or you can ship an agent (like Consul for example) and run that. You have a minimal Linux x64 system after all.
And no you don't have to pay $100/month. The licensing is based on a freemium model and your first app is free forever. And at the end of the day all you need to do is make the decision whether those monetary costs outweigh the value you get. And if it doesn't that's fine too. It simply means Boxfuse isn't the right fit for you.
You explicitly asked for suggestions for alternatives. All I did was provide one in case it may prove useful for you.
I'll be honest: I cannot envision a company that should use Boxfuse. Not one. They're better off with something like a Racker template that takes a few parameters and feeds them into Packer than your spooky-action-at-a-distance stuff.
Packer has a lot of user-friendliness problems (destroying build artifacts on failure, as the OP noted, being just one of them), but it does constrain the universe sufficiently that I can just run Chef and get on with my day.
Everyone has a different experience and perception I suppose.
I agree we have a long way to go, but I disagree that our tooling is "very incomplete" OR "un-battle-tested".
Ignoring Vagrant as you did, Consul is used at multi-thousand node (per datacenter) scale by dozens of companies and a couple specific companies are using it at an even larger scale. And that's ignoring the thousands of hobbyist and smaller company usage at dozen to hundred node scale. I only don't mention specific names because I don't have explicit permission, but you'll just have to trust I don't intend to lie here.
Vault as another example: if you interacted with a financial institution, your transaction at some point likely hit a Vault cluster. Did you visit some websites today? One of the world's largest CDNs has fully deployed Vault for internal TLS cert management.
Those are just a couple examples.
Or, ignoring tech usage completely, we just had our first seven-figure quarter after only three quarters of sales (and that was dozens of deals, not just a handful). You just can't get those sorts of numbers without real world usage.
On completeness, I think the adoption speaks for itself. Tools don't get adopted at the scale they're being adopted without being complete enough to productively solve a real pain point. I believe we have a long way to go but what we have already in most of our tools is relatively complete by measure of being able to get real, meaningful, and productive work done.
Its unfortunate that we can't get to every open source issue and resolve every problem for everyone, but please try to understand that the issue inbound across our projects is massive in addition to trying to build enterprise solutions for customers and run a business. We'd love to hire hire hire to handle all the community inbound but that'd be irresponsible of us financially. Our teams are slowly growing and we're also promoting more and more community members to core committers who help out quite a bit, too. Packer has ups and downs since we don't have a full team around it at the moment but its on our radar of things to work on.
Ultimately, we can improve in every area and we'll strive to do so. In the process, we're motivated and encouraged by our community and also by the "serious work" we see our tools doing every day across various industries.
I'll follow up with you via email to see where we've fallen short for you. I find these criticisms educational and would like to see where we went wrong. I'm not doing that to hide anything from the public, but only because its hard to have meaningful back-and-forth discourse in a few nesting levels of comments. :)
To be sure, there is always room for improvement. HashiCorp does an incredible job at providing active support for the tools they create, and the active community that they've built up over the years is proof of their dedication to their ecosystem.
Keep doing what you're doing, Mitchell!
I've not had much luck with Terraform, I never really "got it" with Otto, it feels like Packer is unloved, and I'm also struggling to understand Nomad (vs. Kubernetes, Mesos, et al).
Hopefully constructive feedback -
There's some disparity between the level of "marketing polish" and "technical polish" with HashiCorp - almost all of the software had hype surrounding the announcement and they have slick websites but launch maturity has varied, and several products have felt like they lacked follow-on investment. Sharp edges haven't always been identified.
Keep pushing. :-)
You guys rock! Hashicorp's tools are fantastic and pleasure to learn and use.
We're a small company and we're going to officially purchase Atlas in few weeks - because it saves tons of time. Also, every developer in our company (rather small company, where people have great flexibility what tools to use) has been cheering and clapping when we demoed CI/CD flow in Atlas. It's so easy to deploy and run services.
Cheers! Roman
It's true, he doesn't have permission.
But at least one company in fintech is willing to admit appreciation of Hashicorp tooling.
Anyone curious come chat with Joel in Napa in a couple weeks.
https://www.hashiconf.com/talks/managing-vault-in-a-federate...
And if you're into distributed systems, talk to us seriously:
"From a values perspective, we're trying to understand the way the world works — that's what our business is — and so we're really interested in people that have a sort of deep curiosity, people that have the patience to understand deep and complex systems," Kreiter said. "Now, whether those are biological systems, or economic systems, or political systems, it doesn't really matter. Somebody who has an interest in and an ability to understand that deeply is interesting to us."
And that's why we like Mitchell.
Thanks for these tools.
I'm so glad you're killing a product. My complaint has always been that hashicorp was doing so many different products it was unable to actually deliver any of them at high quality.
Packer is a reliable and trusty piece of tech for us, but terraform (which has so much potential) is also so rough around the edges we've wrapped it in a mess of our own tooling to make it somewhat usable.
I hope you can find a way to focus on one or two things and do them well.
For example we use Packer daily and it's just fantastic. Yes, once you need to debug it, it's a pain in the ass. Also it makes some basic assumptions for you (for instance it turns off project wide ssh keys while provisioning [in GCE], so you have to manually add them anytime you want do something).
We also use Terraform daily and it works well. Couple of bugs now and then, but it does things so effectively that once it works, you save a ton of time.
Things like ZooKeeper have been around a lot longer and are definitely also in widespread usage.
I will continue using Consul (it's DNS interface for service discovery is awesome) and Vault, but most everything else seems like it's a bit too ambitious and ultimately lacking.
They mostly work, but it's never as seamless as they make it out to be.
I do appreciate the work that's put into these tools but I also believe anything can be criticised. That doesn't mean the author has to fix what everyone says is wrong but an attitude of "well it's free, either PR or stfu" doesn't help either.
Seems like a 'principle of least surprise' violation.
Is that something others have observed?
Personally, I prefer it to alternatives like Boto/Troposphere because it's fully declarative and not coupled to a single cloud.
To give you an idea of how many problems I had I currently have 4 different tfstate files from 4 days of testing. I had to go into AWS and manually delete resources because it couldn't recover from the errors it created.
One example: I was using the ECS option and changed the container source for a service. Seems easy enough and something that should work. Terraform wedged itself after applying the change so badly that I had to blow away the entire setup to get it reset to where it could even run `plan` without erroring.
Otto looked nice but it had fundamental, basic issues and it seems like nobody was actually working on it. I +1'd a bug with the PHP implementation where it didn't give you the option to change the web root and never got an update. This is something that every single decent PHP framework out there REQUIRES and wasn't supported. Otto PHP seemed like it was designed simply to work with Wordpress.
I definitely find it a lot easier to manage and reason about if I mostly avoid third-party Terraform modules. Out of probably a dozen different Terraform projects, I've never run into a situation which I needed to manually resolve. This includes both projects which I started myself and cases where I'm helping to improve/manage client deployments.
> I was using the ECS option and changed the container source for a service.
What do you mean by this? I've used ECS with Terraform extensively and never had a problem with updating the container/image which a service referenced.
That being said, I never used Otto. It definitely seems like they tried to bite off more than they could chew and I wasn't really interested in such high-level solutions.
Yes. I'm actually in the middle of trying that now.
I want to set up a vpc, a few web servers (1-10) with an autoscaling policy behind the vpc along with a bastion server and a cron server, a code deploy setup to work with autoscaling, cloudwatch logging and monitoring, a load balancer, an elasticache instance, and an rds instance. I've been working on this off an on for months. If you or anyone else can point me in a direction to simplify this I'd be grateful.
The core of the problem I had with terraform (outside of the ECS issue) is that there is one AWS service that gets soft deleted. I can't remember what it is right now but it really threw tf for a loop. So I'd setup the stack, do some testing, decide to shut everything down for the day with a `terraform destroy` and the next day i couldn't resume because tf thinks the resource exists but aws doesn't think it does.
I've also seen tfstate get weird after a slow Elasticache spin up or termination. If it takes over 10 minutes it times out. The main thing I don't like about Terraform is that they don't support conditionals, which can be annoying.
You can look at my github.com/RichardKnop/coreos-cluster as an example (that one sets up a CoreOS cluster but you can take just the VPC, RDS, security groups, subnets and NAT bastion from there. I also have couple more terraform repos on my GitHub that deploy AWS infrastructures like you described.
Also look at the GitHub of Government Digital Service (GDS), I think it's alphagov. They have a lot of nice terraform stuff there from their experimentation with different PaaS.
Best devops tools (along with ansible) I've ever worked with.
Use Packer for all AMI building, probably one of the easiest tools to get started with.
Use Consul for service discovery with consul-template and envconsul, all work nicely together.
Terraform. I've been using for the past few months, but it has its problems. Modules use can be very tricky to manage, with no real guidelines on how to use them or TF itself. The isolation can be more complex than say, Cloudformation, which has its faults too. With CFN you have a concept of a infra stack, which will share common resources, with supplementary stacks providing additional resources. Unless you are happy to have a load of main.tf files in sub directories, which seems to be the pattern, you end up with a giant load of TF as one "stack". The benefit is never hard coding things, which you are probably going to do with CFN, plus you have a much better scope of changing resources in TF, but if you don't need that, TF can offer little over CFN with a language to generate it, like CFNDSL or troposphere.
We also regularly break though our AWS API limits for TF because it touches so many things, if you have TF in a single git repo with multiple hooks, your TF can take a long time to actually plan. Recently, I did a plan on 0.7 to see what the impact was for upgrading, without realising that it would overwrite the state file to comply with 0.7, making it impossible to switch back to 0.6. I did roll back to a previous version of the state, but even that caused some issues that I had to manually resolve.
Finally, Atlas. I would say that it is far from complete. There are a lot of things that UI is missing. The API for atlas is also non existent, so you have to use the UI constantly unless you want to auto deploy, which is probably fine for a dev environment, but not production. If you have a plan apply that needs manual intervention, it will sit there while commits pile up above it, so you either let apply and let atlas hit your account with massive amounts of API calls, or you cancel all the plans currently being blocked. It obviously works very different with say S3 and using TF from a CI, but i'm not sure how much focus there is on Atlas vs other products. Atlas also doesn't let you set up things like organisational access tokens, for Atlas itself and for Github. Right now, if we removed someone from our Github or Atlas org, stuff will simply break.
- Packer for production workloads for > 2 years
- Consul for production workloads for > 1.5 years
- Terraform for building production (and all other) infrastructure for > 1 year
Terraform is probably my favorite you have to be a bit careful running it against live infrastructure but it is so nice to be able to figure out who did what to your machines rather than clicking having people click around in the AWS dashboard.
The concept is nice, and the concept of execution is nice too. The execution itself is a failure.
But, hey, they know how to make shiny websites.
consul and vault are very nice for devops.
TerraForm has been huge for us. Packer has been great. We plan on rolling out Vault fully and consul to our entire infrastructure.
The tools I listed above are essential along with Chef, ansible, and git are essential for our tools stack.
Your assessment is roughly correct, Ansible is definitely easier to get started with, and if you just want to manage a small number of nodes that you have SSH access to, it's a good choice. Chef has a larger ecosystem, and a different communication model (VMs under management communicate out to a Chef server over HTTPS, no need for inbound ports to be opened). The tradeoff is a slightly longer learning cycle. I'm biased towards Chef because it's what I know, but having spent some time with Ansible it's a good choice for some use cases.
Terraform seems to be a fully unique product. No CloudFormation isn't really the same thing, nor are Azure ARM Templates. Terraform has solved multiple problems for me that were not solvable any other way except for lots of manual orchestration done with bash scripts.
I haven't used Nomad yet, but I've heard good things and am happy Kubernetes/Mesos has some competition (no I don't consider Swarm Mode a real competitor yet).
- packer, that is a really nice and simple image building tool. It saved lot's of headache compared to tools like oz, because packer fails fast and saves debugging time.
- terraform, that is comparable to OpenStack Heat or AWS CloudFormations. Each have their own issues. But terraform being platform agnostic is an advantage for me.
- vagrant is not modern anymore, but it's ubiquitous.
One should probably think that about sending a PR for a feature request or a bug report if it really impacts them. An attempt, even if in the wrong direction, would probably push the priority of the underlying issue higher for the core devs working on the project.
If you can't do that, at least attempt to be constructive.
- Packer: Not my favorite, honestly. I can't argue that it's doing the job, but it feels inflexible and hard to integrate into a coherent workflow. It seems that Hashicorp Atlas can smooth this over in principle; we don't use it, because it didn't seem to fit with our use of Terraform at the time we got started, and we have a semi-home-grown alternative now.
- Terraform: We're using Terraform not only for low-level infrastructure stuff (VPCs, subnets, etc) but also for application deployment. I'd say our success with Terraform was due to a couple things. First: we picked up Terraform at a time when we were in the process of a total infrastructure rework in our org anyway, so we were effectively starting from scratch. Second: I spent a few months using Terraform for toy things and learning what it was good at, what it was less good at, and building a "pattern library" of techniques that had worked out. Once we started applying it to real problems, we just cherry-picked suitable patterns from that library and used them. I expect that Terraform is tougher for someone who already has significant infrastructure deployed and is trying to manage it with Terraform with few changes, since there are definitely approaches that are harder to model in Terraform than others.
- Consul: I really enjoy the simplicity of Consul. Getting a cluster up and running is pretty straightforward. Once you have it running, you get a highly-available, datacenter-aware key/value store and a service registry. We honestly don't use the service registry very much, but we have made extensive use of the consul-template utility in conjunction with Terraform's consul_key_prefix resource to have applications/services announce where their endpoints are for consumption by their clients.
We actually decided against using Vagrant because it was "more bulky" than our app developers were willing to tolerate. Instead we continued with our previous solution (running the apps direction on the users' laptops with a README in each app describing how to get it running) being optimistic that the new Docker for Mac and Docker for Windows would be awesome enough to get the good parts of Vagrant in a lighter package.
Vault showed up a bit late for our "architecture remix" so we solved our Vault-ish problems in other ways. I like its design in theory, and would probably give it a try if the opportunity arose.
Similar story with Nomad: too late for us, and we'd gone down an alternative path before it showed up. Can't really speak to it, since I only dabbled with it very briefly.
I'm sad but honestly not surprised to see Otto phased out. I was initially excited when it was announced last year but I could never really figure out how to get it to behave in the way I expected... I always felt like I was fighting it, and doing things in a way it didn't expect. I think there's room for the Hashicorp family of tools to "tessellate better", but Otto seemed like a very coarse, heavy solution -- essentially wrapping and templating the complex tools underneath -- where I was more hoping for the tools themselves to grow features to close the gaps.
This turned in to a bit of a rant, so I'll stop. :D
I advise anyone using Terraform in production to wrap it up in some sort of automation. Hashicorp would of course like you to use Atlas :D but you can get a long way with CI/automation tools like Jenkins, Rundeck, ...
We have a wrapper script which: - configures the remote state in a predictable way (setting up remote state properly is one of the more fiddly parts of Terraform usage) - takes a snapshot of the current state - runs "terraform plan" to produce a plan file - takes a snapshot of the current state, which has now been refreshed by Terraform - pauses here and waits for human approval of the plan - takes a snapshot of the current state one more time, even though it's usually just another copy of the last state we snapshotted - runs "terraform apply" to apply the plan created earlier - takes a snapshot of the final state
All that state-snapshotting is an insurance policy against Terraform getting itself confused. There are definitely some gotchas in this area[1] but honestly we've only actually made use of these zealous state snapshots on two separate occasions, and they were both on our pre-production staging environment (which we deploy to more carelessly, as a dry run for production) rather than our production environment.
I have thought about open sourcing that wrapper script but sadly it has some assumptions about our environment built into it (e.g. locking using a specific service in our world, so that two deploys can't run concurrently) and I've not had the time to scrub them out and generalize it.
[1] https://gist.github.com/apparentlymart/657885e730d1e5abc6ea
On the other hand, Terraform is much easier to get started with and much less opinionated.
Disclosure: I work for Pivotal, we donate the majority of engineering on BOSH.
I still think Terraform is great though.
Terraform is also popular.