Salt: Like Puppet, Except It Doesn’t Suck
blog.smartbear.com
blog.smartbear.com
I know that I should be doing this in a more robust way but whenever I try and read up on configuration management tools like Puppet and Chef, they're all described in comparative terms - Puppet does X better than Vagrant which does Y better than Chef, etc. I quickly lose patience and get back to digging myself into a deeper technical hole.
Is there a non-recursive explanation of what these tools are able to do and where someone like me should start?
Edit: Thanks for the helpful responses!
When you want to change the state, you edit the file - say, to add a package.
At any point you can set up an entirely new server from scratch using that same code.
You can't do that easily with bash scripts.
(you could if they worked like virtualenv requirements files but that only works for python's virtualenv...)
The most important thing about it is that you describe the state of the server. And puppet applies it for you. It's idempotent (you might know this term from REST). You can run puppet multiple times and the end result would be the same.
Normally (without config management) you would write scripts that do the configuration. If you ran them twice, things could go wrong, or they would overwrite the stuff from last time. Scripts are descriptions of what steps to take. Puppet manifests are descriptions of what state you want the server to be in.
1: Stop Service
2: Upgrade
3: Start Service
Right now I have just made three non-declarative statements about my tomcat server. I am describing a process that runs through a series of steps to reach a new conclusion.So puppet has these beautiful (sounding) properties. But declaratively describing the state of my servers isn't a real problem I have.
I need to deal with the way that my servers change. I need to upgrade services, including starting and stopping. Puppet absolutely fails at this. Any attempt to describe a process is antagonistic to puppet at a fundamental level.
Whenever people seek a solution to these problems they always get fed that line. That puppet is declarative and idempotent, which is fine. But we should be _very_ clear that puppet does a spectacularly bad job of solving a very ordinary problem that I have, and which I would expect almost everyone else working with large numbers of servers to have too.
If you're interested in trying out Chef - which seems like a good fit from your comment - read on.
For tomcat you could take a look at this community-provided cookbook, which contains resources that make it simple:
https://github.com/opscode-cookbooks/application_java
If you explicitly need to stop/start tomcat beforehand (I believe this resource does a restart after deploying your app) you can do that with the "service" resource built into Chef - though doing so may mean that you're restarting every time you converge the node, which is probably not what you're looking for.
You WANT to stop the service? Really? Stopping the service is an actual headline-level goal, and not a behind-the-scenes detail?
Personally I tell puppet "make sure that the latest version of X is running", because that's what I actually want, and I don't care how it does it...
Maybe puppet has hooks to integrate with the rest of your setup, I don't know - either way, caring about how things are restarted isn't that crazy imo.
There are numerous dependencies to that task - one of which may be stopping the service before upgrading. There are dependencies to stopping the service - maybe telling the load-balancers or your monitoring system. When the service has been upgraded, you can trigger notifications that other items within your config can depend upon (for example, when an upgrade is complete, restart the service, when the service restarts, add it back to the monitoring system etc).
All of these can be modelled in these config-management tools, as explicit dependencies within your manifest file - as opposed to implicit and ad-hoc dependencies within your script.
Perhaps I'm missing some obvious detail but I would think you would just update your repository with the new version of your application and/or Tomcat package and that would cause the service to restart on the next Puppet run, after the new version has been installed.
The app was already documented, source-controlled, version-controlled, and tested by developers and QA. There wasn't much benefit to adding anything to the configuration management repository unless there were platform/infrastructure changes.
Specify the service declaratively, and subscribe it to the package that provides the payload.
service { 'myservice':
ensure => running,
enable => true,
subscribe => Package['myservice'],
}
Then when you upgrade the package on that node, either by using ensure => latest, or by manually specifying the new version, Puppet will automatically restart the service.For example:
Package 'libapache2-mod-php' should be installed. File '/etc/apache2/sites-available/customersite' should be the contents of this file we have on the master server. File '/etc/apache2/sites-enabled/customersite' should be a symlink to above. Directory '/var/www/customersite' should be the contents of this git repository. Command 'apache2ctl graceful' should be run if any of the above change.
With these five rules, you can turn a default install of Ubuntu into a webserver providing client files in a matter of seconds.
Once you have those rules written, you then apply them. You can wildcard, so that hosts matching 'www*.yourhostingco.com' get these rules.
Now, when you want to add a new webserver to your cluster, you install a new Ubuntu instance, you register it with the Salt master server, and then trigger a state update, and you're done. No SCP'ing files, no manual git checkouts, no copy-paste-edit configuration management.
Then, once you've got all that, you can get into templating. You can do templating both on the contents of files (e.g. for memcached config, insert the machine's internal IP address instead of having one config per server) as well as the rules themselves to avoid having a ton of boilerplate (for module in 'list','of','python','modules', install the module via pip with these rules).
One of my favourite benefits of something like Salt is that if you use Salt to do all of your configuration, then just back up your salt config to github or wherever, then you always have a record of how you did things. You don't get the incomplete documentation or missing steps that happen with most approaches.
As a sysadmin managing a cluster of systems, it can be life-changing.
In combination with service.running you can require that, for example, package x is installed, and its service is running, before service y is started.
Meanwhile, a disciplined system of designing and documenting tests and rollback procedures before committing changes will accomplish this without the need for additional software.
My guess is that if you are not disciplined enough to design tests and rollback procedures on your own, then you are probably going to just switch off or work around those features if they are built into the software.
Vagrant actually works with Chef, Puppet, CFEngine, Ansible, Salt, and more.
From http://www.vagrantup.com: "Create a single file for your project to describe the type of machine you want, the software that needs to be installed, and the way you want to access the machine. Store this file with your project code"
That, and http://docs.vagrantup.com/v2/why-vagrant/, sounds quite similar to a configuration management tool to me (I believe you that it's not, I'm just saying I wouldn't have understood that). Seems geared to launching instances/VMs too perhaps. Maybe it's my lack of domain knowledge and maybe I have unfair expectations but I think that the sales pitch should be accessible to a developer who's not solely in the "DevOps space".
Accessibility is good because I'm sure it is incredibly awesome and may well be useful to a broader set of people who aren't yet in the loop.
I agree that we should be evangelizing DevOps and bringing more people into the loop. I was simply saying that Vagrant itself doesn't really need any marketing. It's hard to read much of anything about the field without seeing it mentioned.
Vagrant doesn't need more marketing, but it -- and the other tools -- do need to explain clearly what it does (and does not) do in a way that non-DevOps folks can understand. Perhaps the easiest way to do that is with a simple use case.
I can then use "vagrant up". This creates a new VM using Virtualbox, VMware, or whatever I have installed (I think there's even EC2 providers now!). It configures that VMs hardware as above.
Then, vagrant can use one of many tools in the configuration space -- Chef, Puppet, Salt -- to configure the software in the VM.
I guess the distinction lies there: vagrant configures hardware. Salt et all configure the software.
Provisioning tools let you create it, and incrementally update your image in a way that lets you redo it from scratch at any time.
That alone is the reason why I like the idea, not necessarily the resulting applications that have been created so far for the task.
Shell scripts could do the same and have for years if your only interested in from scratch setups.
SuSE Studio lets you do exactly that. (Though exporting to AMIs is just a feature, not the main objective.)
1. Write a script to configure an instance and run it when the instance starts.
2. Clone the instance to an image
3. Run instances based on the image.
It's quite straightforward to do. See http://github.com/gyepisam/fcc-textify for more details.
The actual configuration of your system will drift further and further away from that templated AMI, leaving you to constantly have to build a new template every time you deploy a new machine, or manually make all of those differential changes to your new system.
Config management can be tailored to rapidly update the configurations of certain classes of systems at a greater interval than the constant cycle of blowing away machines, updating templates, redeploying, etc , etc.
If AMIs are your deployment method of choice, you still need to build a repeatable process for applying updates and changes to your previous base AMI, testing that it still works under all your realistic configurations, and then gracefully deploying it to your production infrastructure.
Which is all doable, but not trivial. And the processes you build will be very Amazon-centric.
However, there are some common sense tiers:
3rd party dependencies that changes infrequently = put in ami Common ops daemons and boot scripts = put in ami your application code and your configs = deploy via python/ssh client side
Also, one simply has to use ec2 instance tags to name things.
Result, you can have bit-for-bit identical instances fired up in no time with high confidence without need for a full OS package mirror. Without, of course, a ridiculously over-engineered configuration server framework, DSL, SPOF, security holes...
I asked myself this same question several years ago. The basic advantages chef and puppet (and salt, etc.) offer over client-side python alone are:
A thin abstraction layer for basic system/platform information (facter/ohai). System information that you'd obtain using various tools like dmidecode, uname, df, /proc, ip, are all gathered and made available in a single data structure. Overkill? Maybe, but makes for much cleaner scripts.
A library for performing common tasks: package installation/upgrade, directories, files and templates, services, and shell execution. A full list: http://docs.opscode.com/chef/resources.html#resources.
Finally, a framework for organizing your configuration in a standardized way, with most of it in one place. With chef, for example, the idea is to put all the configuration for a single application in one "cookbook" and then if you have two different servers that both use that application in different ways, you define two different roles that pass different parameters to the basic recipes. Obviously the extent to which this one a benefit depends somewhat on your needs and personal preferences.
But I do agree that both chef and puppet are over-engineered. Puppet's DSL and chef's server are mostly overkill that I don't have much use for. But the DSL can be dealt with and the chef server is optional. There doesn't have to be a single point of failure.
After having used it for a while, I adore it. One minute and one command to go from a new project to a full deployment, what more can you ask.
Jesus, and I thought the CDCs "big vault full of things that will kill you" was a terrible idea...
At this point, the only thing that could save ordinary users from the melee of competing Open Source configuration management frameworks is an illustration by Randall Munroe.
(My hacker senses are tingling. If you listen very carefully, you can just barely hear a new eponymous law being born.)
Checkout fabric.py and cuisine.py, both are small, simple, elegant approaches.
Combined with a native package manager (like apt), fabric+cuisine is nearly as powerful as puppet or chef.
The ZeroMQ stuff makes sense if you're pushing configurations inside a data center, but it's a dealbreaker for us having things hosted externally.
(Then again, that's how it usually goes anyway.)
This library does not protect against timing attacks.
Do not allow attackers to measure how long it takes you
to generate a keypair or sign a message. This library
depends upon a strong source of random numbers. Do not
use it on a system where os.urandom() is weak.
I'm not saying Paramiko (or its patch sets) are insecure, just pointing out that the same arguments can be made against the libraries and code that Ansible is based on.So, don't use it in the cloud? [1]
The other big win for me is I can read their code, I understand python & have a number of items I'll be able to contribute to upstream that will help others use the product.
Also, Ansible does not require you to mess around with dependency lists to ensure that packages/files are installed in the right order - the order is built into the yaml config file. You don't lose any capability, you just gain (and this is my personal opinion) ease of understanding the order that operations will occur in.
A: "I just set up a cloud instance by running some shell commands by hand."
B: "You shouldn't do that, because of X and Y and Z. You should learn Puppet or Chef."
A: "Wait... did you just tell me to go spend thirty hours banging my head against solid objects, in exchange for nebulous benefits that I can't even perceive yet?"
B: "Why, yes, I believe I did!"
Ansible feels much less embarrassing to advocate.
I've tried B, C, D, and E, and I like E.
And if you criticize me, I down-vote you. Classy.
There are ways to configure salt with masterless or behind vpcs or with syndics that are perhaps an enhanced security model. But defaults are for regular use cases, and for most cases the defaults are fine.
I suggest checking it out.
You don't need to do this. SSH is secure enough. Require key-based authentication and leave SSH on port 22.
Sure, those drive-by attacks are pretty weak and not much of a threat against a hardened configuration and the real attackers will find the new port anyway, but it reduces noise significantly.
Cleaner logs are easier to parse, so the net result is that you can spot attacks that you care about much more easily.
1: at which point you'd most likely need to be a lot more involved anyways, and you should most likely be running something more serious in front of your server anyways...
Why are you exposing VPN ports in public anyhow? Put them behind a VPN.
Seriously though, there isn't anything to say a VPN is any more secure than SSH.
There is a longer conversation that can be had to explain this. There are lots of mistakes one can make in managing a box with SSH that will easily lead to an account compromise, not even considering universal methods like 0-days, stack attacks, mitm, phishing, etc. The simplest defense to all these problems is a hardened VPN device on a separate network segment with separate authentication and strictly defined ACLs.
If you just run a personal VPS, I wouldn't worry about it. But if you ever start handling customer data, get serious about security and don't run SSH in the open.
Something concrete please..
SSH contains many features which may make it vulnerable to attack. If configured incorrectly it will expose the machine, and depending on the machine this could expose a big portion of your network. I could write 5 pages on abusing SSH options for entry.
After SSH itself, there's the operating system. PAM has had holes for years and yet people everywhere rely on it for their authentication. Their OS might not be patched up, and probably has had no hardening of the kernel or userland against typical attacks.
The box is also probably on a shared network segment, meaning your attack surface is now a whole lot bigger: many machines and user sessions to hijack.
On the other hand, a VPN on a separate network segment provides protection for all these components, if you do it right. Usually the VPN should be an appliance of some kind, with service by a company who spends its time hardening the device and patching the software regularly. This makes sure the configuration of the software is correct, the system itself is resistant to attack, and authentication/authorization is managed outside the machine, which means limited access to the network for specific user sessions.
You cut the attack surface down by separating their network access away from a specific network and shared server. You also cut down on mistakes in managing said service because everything is a managed service, not a config file edited by one or more admins.
Also consider that SSH just doesn't have the features of a standard VPN that make it useful on a network level. Its tunneling feature is a joke; using ppp is the closest way to provide an actual remote vpn tunnel. Persistent connections are nonexistent. Network access control is nonexistent. Pushing routing, DNS and other network information is nonexistent. It is not meant to be a network-level tunnel, it's a single-session user tunnel for remote hosts, not networks. The extra options and features for more network access are hacks.
If you wanted to emulate a VPN using SSH, you could:
* put a machine on a separate network segment
* apply Grsec and other patches
* harden the filesystem and stop all services other than SSH
* disable user logins
* disable all features of ssh
* enable all strict ciphers, modes, options for ssh
* add ssh config options to run pppd as soon as the user logs in
* configure PAM to use RADIUS or SQL to auth with a remote database
* use iptables and some kind of packet-marking module to track network sessions by user id
* create a custom application that configures iptables rules per user
* constantly update all patches
That would get you close to a real, normal VPN. Unfortunately, it's also a metric shit-ton more code and services to do the same thing as one service, and considering you'd be setting it up for the first time, there's probably gonna be some mistakes made.SSH is not a VPN, and a VPN is not SSH. Like I said before, if you're just remoting into your personal box, I don't think you'll have issues, because nobody would care to break in. Bigger networks are a different matter.
Now, it seems like you're saying that vpn is more secure because it's often separated and more hardened. I guess that kind of makes sense, but isn't that really an argument for a single, hardened entry-point into the network? Could be anything.. And of course, this has its drawbacks as well =|
If you have a building you want to secure, you could make a lobby with a man-trap and a guard at the front desk and security keys for the lobby and on each floor. Or you could put a padlock on the back door. You decide whether it's worth the risk.
In terms of what to use to remote into your private VPS, just use SSH. A VPN would be overkill.
For many people ssh is considered secure enough. And any secure issues get immediate attention so you can be updated pretty quickly.
Now any software could have unpublished zero days, which is why I also do some other host based security.
Look, I can maintain my own ansible configurations (inventories, playbooks and roles), and the other groups need not be the wiser. They don't need to worry that servers are going to be reconfigured out from under them, since I am running my configurations explicitly. All they see is that I ssh'd in and did a bunch of stuff really fast. To work with my other groupmates, we just keep the files in git (just flatfiles, yes: just flatfiles). With git I have an audit log of everything that has changed (hosts that have moved environments, configurations that have been updated, etc). And like git it is as distributed as you want it to be. Want to run it from a central server only? Fine. Want to have your admins all run it from their workstations? Fine. Go for it.
Even if the server is ancient or weird (ie. no python), I can still manage it with the raw module.
Ansible gives me everything I want, with no fuss. It is so basic that I can do things the way I want to.
Personally I've found dealing with Ruby and the various versions rather irritating. I kind of wish I had started with Salt or Ansible because I know far more Python, than I do Ruby. I started down this bunny trail before Vagrant support Ansible and/or Salt. I guess it's not to late to start over... hahaha Maybe I will after the initial release.
I guess the only advantage of Puppet was that I'm able to do most of what I want using major modules...
IMHO, the overwhelming problem with salt/cfengine/puppet style solutions (which I will refer to as 'post-facto configuration tinkerers', or PFCT's) is that they potentially accrue vast amounts of undocumented/invisible state, therefore creating what I refer to as configuration drift.
IMHO, a cleaner solution is to deploy configuration changes from scratch, by deploying clean-slate instances with those changes made. In addition, versioning one's environment in this way creates an identifiable point against which to execute automated tests. (This class of solution I refer to as 'Clean-slate, Identifiable Environments' or CSIES.) Examples are Amazon AMI's, and any other kind of versioned/identified VMs.
PFCT's deployment paradigm tends to be relative slow and error prone. CSIE's tend to be fast and atomic. PFCTs are headed for the dustbin of history. They are temporary hacks that clearly grew from old-school sysadmins' will to script. CSIEs embrace modern day devops, as more holistic entities that embrace virtualization and recognize the integrity of the environment as critical to preventing ridiculous numbers of environment-induced, service-level issues that are an expensive tangent to service development, testing and deployment. Thus, I would argue that what we are looking at with PFCT's is a failed paradigm, and with CSIEs, the now real and current opportunity for something far more elegant.
(Disclaimer: Haven't tried ansible or vagrant first hand, but they do seem to be PFCT's to me.)
I personally believe that this area is going to expand rapidly, building off of the trajectory begun by present first-generation of cloud infrastructure. Perhaps for someone who wanted to practice thinking in this area, I would a say exercise is to limit yourself purely to third party and cloud based infrastructure but demand high performance and global (multiple cloud provider) availability, and challenge yourself to write a multi-service system including a deployment tool that automates your solution and actually produces maintainable infrastructure. If you follow through, you will understand the problems.
Salt as an 'ecosystem' has capabilities that extend more into this area. Salt Cloud and the Salt Bootstrap script are 2 of the things that I feel invalidate some of the argument that tools like Salt ( I wont comment on the Chef & Puppet ecosystems ) are capable of operating much closer to your aim than you give them credit for.
For starters, I'd quit before letting a boss tell me their precious snowflake AMI was 'more stable' and I shouldnt waste any time trying to ensure I can recreate it should someone delete the image. In my own workflow, Salt is step 1 in any system. I build a stand alone salt configuration that can from bare OS image on first boot, initialize salt, and then pull the system forward to the desired state. Then for cloud roll out, the next step is to take an image of that VM/Container which I can then replicate much faster.
And one of my long terms plans is to hammer Buildbot, Salt, Salt Cloud, and a lot of time into a system that gives end to end control. Now if only I could code a decent UI for it... putting bootstrap on months of work would feel cheap lol.
Right. It's probably fair to say that lots of present-era PaaS doesn't give you a large bounding space. Perhaps cloud providers always limit your space. Your own hardware can even provide limitations. But within that which you control, it's extremely important to version, package and test configuration sets .. or you wind up with a wide class of tangential issues.
Salt Cloud and the Salt Bootstrap script [are] capable of operating much closer to your aim than you give them credit for.
You may be right.
... waste any time trying to ensure I can recreate it...
If you can't automate the generation of your environment, you want to maintain the systems that you produce, and they of reasonable complexity, then IMHO you are asking for trouble in the long run. This is something that took me awhile to learn.
> IMHO, the overwhelming problem with salt/cfengine/puppet style solutions (which I will refer to as 'post-facto configuration tinkerers', or PFCT's) is that they potentially accrue vast amounts of undocumented/invisible state
I think you mean "post-facto" as in: "run after everything is done"? This is not the way that people would advocate Puppet should be used. Puppet should be used from the start, not added as an afterthought once you are done.
> IMHO, a cleaner solution is to deploy configuration changes from scratch, by deploying clean-slate instances with those changes made
This isn't a cleaner solution, this is almost the solution you get when you use Puppet. With puppet the development workflow is like this:
- Span up a vagrant VM
- Run your manifests against this VM to test them
- To check in, run your manifests against your staging environment
- To deploy, spin up new clean production VMs and run puppet manifest against them
- Use a reverse proxy to route all traffic to new production VMs. Terminate old production VMs
> PFCT's deployment paradigm tends to be relative slow and error prone.
This much is true. Puppet's slow speed is particularly galling, but maybe that's just because I use it at work.
The difference betwen deploying an instance of a stored environment and generating that environment from some prior state is the generative process, which can fail or change in unexpected ways due to network conditions and other factors.
More importantly, PFCTs enable and to some extent encourage modification of generated environments remotely, en-masse, without any significant capacity to ensure that individual instances within a group have not subtly shifted in configuration. This is what I meant by configuration drift.
CSIEs, by contrast, are essentially the complete product of the entire generation process, thus ensuring that future instantiations are identical. A subtle difference, but an important one.
I forgot to mention this before, it is strange that you credit yourself with defining this term when it has been well defined for some time in ops.
> More importantly, PFCTs enable and to some extent encourage modification of generated environments remotely, en-masse, without any significant capacity to ensure that individual instances within a group have not subtly shifted in configuration.
But this adds a significant weight over and above the "generative process" of running manifests. Yes, running manifests against your VMs can "fail or change in unexpected ways due to xyz" - don't do it against VMs that are currently in production! I'm not sure you've ended up with anything less error prone and you're still going to need a way to get from a fresh VM image and your output images - which is where Puppet would come in.
I'd really rather not make the entire disk image my build artefact, for fairly obvious reasons (ie: size).
You might like this, which is written by a colleague of mine, except that it is not in "opposition" to Puppet/Chef/etc:
Frequent refreshes are great, but your system is only doing something useful once it has "mutated" (i.e. accepted external data to operate on).
The tradeoff system designers have to make is frequency of refreshes vs the cost of transferring interesting data to that server.
Seems like your organization has just coined a new synonym for "gold master"
CSIEs are still a generative process. The difference with what you call PCFT is that the generative process isn't swept under the rug and codified into a versioned image unless it's really necessary for performance reasons.
The result is that it's easy to maintain a clear distinction between machine state and human instructions. For a trivial example: a list of packages that humans decided are necessary for the system versus the final output of 'dpkg -l' after all dependencies have been resolved.
With chef/puppet/etc. the code used to generate instances represents a human-created description of what the environment is supposed to look like, with as much version-control and referenced documentation as is necessary. With a versioned-image approach, all you have is the one-dimensional history of the image in question.
1) You have not dealt with large enough data, since you advocate just creating VM copies or snapshots. Try that on 10 or 100 TB of data.
2) You haven't thought about what and how those initial CSIE configurations are generated. Do you hand tweak everything, make && make install onto a particular installation of a particular OS all the software then just spawn those? It seems that should go to the dustbin of history. You know essentially have a black-box that someone somewhere tweaked and you have not recipe on how to repeat it. If that person left the company, it might be tricky to understand what and where and at what version was installed.
If you have "configuration" drift there needs to be a fix to the configuration declaration and people shouldn't be hand editing and messing up with individual production servers. If network operations fail in the middle then the configuration management systems needs to have better transaction management (maybe use OS packages instead of just ./configure && make && make install) so if the operation fails, it is rolled back.
There are many ways to take an image of an environment, not only VMs or snapshots. But if your system image includes 10-100TB, it could be argued that the problem of size really lies in earlier design decisions.
You haven't thought about what and how those initial CSIE configurations are generated.
On the contrary, generation should be automated. In the same way that a service to deploy to such an environment is maintained as an individual service project, the environment itself is similarly maintained, labelled, tested and versioned as a platform definition.
Without disagreeing with your conclusion about the design process, it's useful to note that this situation simply isn't a problem for a conventional configuration management tool.
> On the contrary, generation should be automated.
One could argue that Puppet and Chef are ideal tools for performing that automation.
Sure. But loads of other stuff is. The weight of tradeoffs is clearly against PFCTs here.
One could argue that Puppet and Chef are ideal tools for performing that automation.
Absolutely agree - but not within live infrastructure. Only build.
A mix of two. Use salt/puppet/chef etc to bootstrap a known OS base image to a stable production platform VM for example. Then spawn clones of that. I would do that and I see how it would work very well with testing.
All of us who have built big cloud-server clusters have dreamed of this plan at least once. But there are big practical problems.
Relaunching infrastructure is easy in theory, but from time to time it becomes very difficult. There is nothing like being blocked on a critical upgrade because your Amazon region has temporarily run out of your size of instance, or because the control layer is having a bad day, or because you've accidentally hit your instance limit in the middle of a deployment, or...
A much bigger issue is that bandwidth is finite, so "big" data is hard to move. This is a matter of physical law. It's all well and good to declare that you're never going to apply a MySQL patch in place: You're just going to launch a new instance with the new version and then switch over. But however fast you manage to launch the new instance (and you will be hard put to launch an instance faster than you can apply a patch and restart a daemon...) you will be limited by the need to copy over the data. Have you ever tried copying half a terabyte of data over a network in an emergency while the customer is on the phone? It is very annoying. Because it is often physically impossible to do it quickly: Cloud infrastructure isn't generally built for that, and when it is it costs money that your customer will not want to spend for the luxury of faster, cleaner patch-application.
A solution to this is to use cloud storage like EBS. Now your data sits in EBS and you just detach its drive and reattach it to a new instance. That actually works okay, provided you're happy with the bandwidth and reliability of EBS, which lots of people aren't – and, as those people will cheekily point out, you have now solved the "relaunches are slow" problem by replacing it with an "everything is uniformly slow" problem. Moreover, detaching and reattaching EBS volumes isn't instantaneous either. You have to cleanly shut down and cleanly detach and cleanly restart, and there's like 12 states to that process, and all of them occasionally fail, and if you don't want your service to go down for thirty seconds every time you apply a patch you need a ton of engineering.
Which brings us to the other problem: Complexity. Most programmers are not running replicated services with three-nines-reliable failover that never breaks replication. But even if you are, because you've got the budget for excellent infrastructure and a great team, it will always - for values of "always" measured in several more years, anyway - be more complicated and risky to fail over a critical production service than to apply, say, a security patch to 'vi' in place on a running server. 'vi' is not in your critical path. If you accidentally break 'vi' on a live server (and you won't, because vi is older than dirt and solid as a rock), you will have a good laugh and roll it back. Why risk a needless failover, which always has a chance of failure, when you could just apply the damn patch and thereby mitigate risk?
At Google scale that argument probably stops applying. But most people don't run at that scale and it will take decades to migrate everyone to a system that does, if that even happens.
So, "dustbin of history", maybe, someday, but in the long run we are all retired, and I will be retired before our dream becomes reality. ;)
Your fifth paragraph describes problems related to operations process, which are entirely avoidable.
The very purpose of most of these tools is to make state visible, to be able to see what has been poured in. That visibility contributes to extensibility, the ability to take that known configuration starting place and to be able to branch and create new configurations.
If you have infrastructure in mind for how your CSIEs are constructed, I'm all ears, but I'm envisioning more sh scripts posted on the corporate wiki as your CSIE implementation.
Sure.
The very purpose of most of these tools is to make state visible
Right, but they do a poor (ie. post-facto, limited granularity) job of it. This is why IMHO their paradigm is inelegant.
infrastructure in mind for how your CSIEs are constructed
We use scripts with exit values that execute within the target environment to bring it from a base configuration (eg. some AMI or some distro) through to the desired config, plus validation tests. I think most people's approaches will be similar. PFCTs essentially provide this, and could be used for this step without issue.
This type of design leads to tooling which is relatively slow and often obtuse, but there are well-reasoned (and researched, start maybe with [1]) assumptions behind those tools. In particular: systems modeled by these tools are designed to run for years, include a large variety of hardware and software, to be worked on by a lot of different people, and to be able to grow to massive size.
The "clean slate" idea is a good one, and has been in use for a long time (the idea of gold images goes way, way back). But that's for initial system state, not dynamic system modeling. The cfengine family of tools grew out of limitations in the "clean slate" approach for real-world system models.
If your concern is "how can I deploy my Rails site to EC2 quickly and reliably," you'll have different goals and assumptions on the tools you need, than if your concern is "how can I grow a resilient infrastructure". (edited to add: both are completely legitimate goals)
All that aside, I definitely agree that the tools still suck :)
[1] http://cfengine.com/markburgess/papers/sysadmtheory3.pdf
Try telling a client who wants 100% uptime that! Computer systems can be many things - it could be (and probably has been) argued that we as progammers fundementally aim to reduce random qualities and increase stability within our programs.
there are well-reasoned (and researched) assumptions behind these tools
I can see where PFCT's came from. I took a look at the paper (which is exclusively about policy-driven management systems), but don't think this invalidates my points, re: paradigm failure. Taking the same policy based management systems and applying them to a CSIE environment (say, with corosync/pacemaker) is one way to combine their benefits. (That's what we do on our own infrastructure, actually.)
If your concern is...
Right, differing concerns. You also have different life expectencies for instances, purposes (dev/staging/prod), availability expectations, etc. But these are tangents! Developing good software comes down to consistently carrying out fundamental practices regardless of the particular technology. - Paul Duvall.
We attempt to remove an entire class of issues related to deployment environment by changing our platform engagement paradigm to one that is less procedural/'stochastic' to something that is more atomic/reliable. At the same time, this facilitates a clean segregation (and versioning!) of deployment environments versus service development (which may target one or more CSIE).
> Try telling a client who wants 100% uptime that!
A client who asks for 100% uptime will end up disappointed.
> We attempt to remove an entire class of issues related to deployment environment by changing our platform engagement paradigm to one that is less procedural/'stochastic' to something that is more atomic/reliable.
I'm afraid I don't get the difference. You've argued that configuration drift is a major problem with configuration management tools, but I don't see how any solution based around deploying full system images couldn't also apply to stepped upgrades applied by CM. Allowing config changes to be made directly to production servers instead of going via the deployment tool is the problem there, not anything fundamental to which deployment style has been used.
With puppet configuration (for instance) in version control, it's not a problem to test against a known, versioned, identifiable and auditable environment. As long as you're not applying config changes to live servers, the switchover to a new environment is equally atomic either way.
Given that the trade-off is building, storing and deploying full system images, I don't think there's a fundamental advantage to full-system deployment that can't be matched by a thought-out application of conventional configuration management.
Sure.
[...] I don't think there's a fundamental advantage to full-system deployment that can't be matched by a thought-out application of conventional configuration management.
OK, on the face of it, this is a fair line of reasoning. If we assume, however, that we are looking for ... replication (say, for the purposes of regression testing, etc.) then we really do need to know that the entity in question is the same as the last time it was .. err .. generated/instantiated. With the process you propose, there is clearly higher risk here. That is a paradigm weakness.
For me, however, it seems like it would be impossible. My team has to manage 10,000+ physical servers. We can't possibly wipe and provision from scratch every time we change configurations.
In other words, if your servers are ON Amazon's platform, use CSIE. If you are RUNNING the Amazon servers, you need something else.
If you're not copying the system image and simply network booting a remote image, then that doesn't apply, and yes, that can be fast.
I'm in the midst of a large puppet enterprise deployment at my day job, and it seems they've taken great pains to prevent any drift from happening. You get a dashboard webapp that shows you every puppet run on every machine, and a large overview that gives you states like "nonresponsive / changed / unchanged / pending / error".
When a puppet run makes a change to a systems, that run is marked as changing the server. If you're getting lots of "changed" runs on a system, you immediately know to go look at that server because something is making that box deviate from your defined baseline.
You write your configurations (manifests), add them by name to the webapp, add those to groups, and then add machines to the groups, which define what set of manifests to apply.
On my admittedly newbish level, it would seem this system is completely immune from any drift-over-time, provided you bother to glance at the dashboard occasionally. We're going to put it in our NOC as soon as the deployment is done :)
http://nixos.org/nixos/ https://github.com/NixOS/nixops
It's a young project but it already solves most of these issues.
https://github.com/ansible/ansible/
It doesn't require any deamon and does all its work over the good old unix fashion way: SSH. And it's python too.
"Chef works atop ssh, which – while the gold standard for cryptographically secure systems management – is computationally expensive to the point where most master servers fall over under the weight of 700-1500 clients. Salt’s approach was far simpler."
Does that assertion about Chef somehow don't apply to Ansible?
On the use case:
"I have this command I want to run across 1,000 servers. I want the command to run on all of those systems within a five second window. It failed on three of them, and I need to know which three."
[1] http://www.ansibleworks.com/docs/playbooks2.html#pull-mode-p...
[2] http://www.ansibleworks.com/docs/playbooks2.html#fireball-mo...
http://jpmens.net/2012/10/01/dramatically-speeding-up-ansibl...
However, you're not forced to use this. In the beginning, you can just seed your CentOS or debian with a Kickstarter or seed file and then run your inital thing with ansible simply over ssh (using all the goodies, ssh-agent, password less ssh etc..).
One huge plus for ansible is also that it used yaml which is rather simple. I've been following both project for >1 year and it seems that recently ansible has picked up a lot and will probably make the "race" (IMO).
> Chef works atop ssh, which – while the gold standard for cryptographically secure systems management – is computationally expensive to the point where most master servers fall over under the weight of 700-1500 clients
That said, unless you must have 700+ simultaneous slave connections, you should probably make life easier on yourself and choose ansible.
> We’ve recorded a free 2-hour presentation designed to ...
Wait. The quickstart presentation is 2 hours?!
Looking for an easier way to get started...
Salt is all of these things as well.
http://www.steve.org.uk/Software/slaughter/guide/
Policies (read "perl scripts" + "file templates") are fetched from a central location, which could be a git repository, a SSH server, an rsync export, or similar. Then they're compiled and executed locally.
Surprisingly powerful, plus you get the power of Perl + CPAN.
It turned out to be a rabbit hole. As soon as I thought I learned just enough to get it running, something else popped up that stopped me.
That's why I created PuPHPet [1]. So far the reception has been fairly positive.
At one point in my learning, I got fed up and tried Salt. I couldn't get the Salt hello-world running. I followed the directions to a T. If your tutorial is incorrect, or hard to follow just to get the most basic version up and running, it will turn people away.
Also, this was all on top of Vagrant.
[1] PuPHPet - https://puphpet.com
* Ubuntu Precise 64 Bit (12.04.2 LTS)
* Apache
* PHP 5.4
You can switch between Apache/Nginx and PHP 5.4/5.3 (5.5 coming today!)
I am generally able to SSH into a box and get things configured the way I need. However, I have huge amounts of trouble translating that into salt scripts.
Consider logrotate. Here is the only documentation I can find on the topic [1]. From this, I have no idea what to put in init.sls to make sure a given log file is being rotated correctly. It seems this would work on the cmdline, but not necessarily in a salt script.
And that's just for logrotate! My uswgi + nginx configuration - translating that into salt - I don't know where to begin.
How do I make sure things get installed in a certain order? (Answer seems to be having 10 directives, for 10 packages, each depending on another, to enforce order.)
Is there anything that more closely mirrors what I actually do when configuring the box? SSH in, set certain values, etc? I guess I could write a shell script (or use fabric) but then I seem to have lost the point of configuration management.
[1] http://docs.saltstack.com/ref/modules/all/salt.modules.logro...
You will be very well off if you read and 'digest' the Salt docs on States first, before moving on to modules, pillars, grains, custom returners, etc.
What you probably need to do with logrotate is take the configuration that you normaly setup on your servers, then add it to your salt system. So top.sls calls 'logrotate' running the 'logrotate/init.sls' and that has a definition that says "I want logrotate installed, I want it running as a service, and by the way take the file 'logrotate/config.conf' and shove it in /etc/logrotate.d/ as <correct filename>, p.s. If i change that file, restart logrotate"
With States & the requisite declarations to enable salt to know what order things need to be in, you shouldnt have much trouble adding a simple service like logrotate along with a specific config file to use for that service.
>And that's just for logrotate! My uswgi + nginx configuration - translating that into salt - I don't know where to begin.
For nginx you just need a package: nginx installed state and a few states for the conf files.
uswgi will need something similar.
Using Salt's logrotate module might be more than you need. You can always just have salt manage a file in /etc/logrotate.d/ and keep it up to date, which would be a lot easier.
The salt module is meant for more complex, fine-grained tuning of logrotate files, but I'm not really sure why I would use it.
> My uswgi + nginx configuration - translating that into salt - I don't know where to begin.
The approach that I use: check your entire nginx configuration into git (e.g. your whole /etc/nginx directory) and then you can configure salt to check it out and keep it updated.
Also, if you're custom-compiling nginx into /opt/local/nginx or something, you can actually have salt recursively copy the entire directory to your slaves.
edit: link
It's as much a semantic thing as a technical thing. Instead of thinking about an unconfigured node as a "minion" awaiting orders and provisions from central command, I prefer to think of a node like a stem cell, fully capable of differentiating itself based on signals that it receives. You need a way to update the DNA and a way to send the signals, that's it.
This may seem like a meaningless difference, since there is still value in centralized services (package repositories, security, reporting, monitoring). But it's still a subtly different focus and over time yields different results.
For my part I think the distributed, organic "stem cell" way of thinking will win out over "master/minion" in the long run.
If you want to get a feel for how salt looks like when managing some servers and laptops you can take a look at my states: https://github.com/uggedal/states
http://docs.saltstack.com/topics/releases/0.15.1.html
https://github.com/saltstack/salt/commit/5dd304276ba5745ec21...
Here's a good article( http://missingm.co/2013/06/ansible-and-salt-a-detailed-compa...) with a comparison to Ansible that others are also mentioning here. Ansible uses KeyCzar which which seems more sane than rolling your own crypto as many readers here on HN know.
http://gigaom.com/2013/06/20/devops-player-saltstack-wins-st...
(I'm a SaltStack employee)
Funny, that's how I feel about Ansible compared to everything else including Salt.
That was handy for me too as while I'm somewhat familiar with Ruby, I'm no expert at all. I can read Ruby no problem and I can write ruby that's not-quite-idiomatic and I'm terribly slow at it.
Point is: Python is the number one scripting language (after bash) for sysadmins just like Perl used to be.
took you 3h? It may take slightly longer if you want 1.9, but it still exists in fedora so it should not take 3h to solve.
https://access.redhat.com/site/documentation/en-US/Red_Hat_D...
With that said you can use Omnibus Chef Installer now which includes a copy of Ruby just for Chef. Good for servers where you don't need Ruby or small servers that would take awhile to compile a newer Ruby.
This matches my experience with a significant amount of software on CentOS. Since CentOS is just a rebuild of RHEL, and RHEL is extremely conservative when it comes to new software, CentOS tends to be out of date at release and get progressively worse.
That being said, we use third party repos for big projects like mysql and php on the assumption that a project that size is going to be thoroughly tested.
Also, Salt is written in python, for which the same argument holds as well for almost all environments.
That said, I also disagree that puppet "sucks." It's good at what it claims to be good at so long as you can deal with its quirks.
The important thing is someone uses some form of Configuration Management. If people find Salt easier than Chef/Puppet/CFEngine/Ansible then great. At least they have something to build/scale their infrastructure.
It doesn't have to be this way. The situation where one host repeatedly needs to talk to hundreds via SSH is precisely where the SSH ControlMaster socket shines. This saves you a ton of overhead by not having to start up and tear down the session every time you want to issue a command via SSH.
I often use this trick on busy Nagios servers that execute many active checks via SSH -- it works well.
What, specifically, "sucks" about Puppet and Chef and what is so much "simpler" about Salt or Ansible? As an Ops guy who has been running Puppet since 2008 (and Chef most recently) against hundred of servers, I don't see the simplicity reflected in the documentation, nor do I find Puppet or Chef particularly complicated.
(Ok, Chef's attributes system is a bit confusing at first, but it is hugely powerful.)
> Chef works atop ssh, which – while the gold standard for cryptographically secure systems management – is computationally expensive to the point where most master servers fall over under the weight of 700-1500 clients. Salt’s approach was far simpler.
I think I'm with you (without the experience): I find "it works over ssh" a lot simpler than "we wrote a custom protocol on 0mq." Simplicity apparently has lots of interpretations. I couldn't care less if ssh performs well enough to support a trillion connections. In practice, you only need a handful, usually one.
Maybe Salt is fantastic. I'm not really in a position to judge. The article made it sound interesting to me, but I'm not sure attacking Chef/Puppet was really necessary, especially since it wasn't really expounded on.
1) It's slow. Puppet runs take forever and use a lot of resources.
2) The Puppet language is terrible. Basically it looks like someone took Ruby and hit it with the stupid stick. Trying to build reusable components nearly impossible without resorting to hackery.
I have no idea is Salt/Ansible/Chef are any better, but I'd be willing to give any of them a try based on just these issues.
Chef doesn't run over SSH in any environment I've used that wasn't a toy (vagrant w/ chef-solo). Please fact check.
Fanboy article.
The notion is deeply flawed to me: using an image as a precondition for making an image, over time, becomes an intractable mess and requires very careful supporting documentation to prevent the scheme from devolving into a bunch of "buckets of bits" with no transparency into what work has gone on to make it that way.
The #1 thing that I enjoy about automated tooling is that I can take a bare OS and spin up a complete new system in a matter of minutes, and I get to watch that entire process happen before my eyes. There's no mystery, no external dependencies, no existing work I'm riding off of: everything that happens is visible to me in an immediate way.
There's a value & use to immutable images, but please decouple your image making from past images made: no one wants to root around to figure out what twelve horrible things you did to install Java 9 image instances back whenever, nor are they going to have any fun reproducing it on the twenty nine active variants of that ancestor image when there's a security fix to be done.
What I really like about salt is that everything is in one place and all goes towards building the same data structure that everything runs off.
However, I am not sure how would one use Ansible where VMs get launched dynamically (private cloud/virtualization fabric where devs can instantiate systems) and then receive their configuration without any manual steps.
For example, one can create kickstart/VM-images which get a hostname based on certain regex pattern, register with a Puppet master, the Puppet master auto-signs certs matching this specific hostname pattern and then client nodes receive their catalog. This is really useful pattern wherein systems pull their configuration state almost immediately after boot. It requires manual setup only while writing kickstart/VM-iamge profile and Puppet master configuration.
Ansible's SSH keys setup requires manual intervention, however, I think it can be automated using pre-defined keys in kickstart/VM-images. Haven't tried it yet though...
We tend to destroy and recreate servers more often than we scale out, so we haven't bothered to remove the manual step of adding the server's hostname to the ansible inventory_hosts file. However, that's easily automatable...
Ansible will _execute_ your inventory_hosts file if it's executable, and IIRC it just needs to return a JSON or YML data structure representing all your servers and the groups they're in. So, as long as you have a library which can query your infrastructure (e.g. boto for EC2 etc) it's not hard to automate this.
Our greatest challenge has been coming up with a tool which can manage images for both VMWare and Microsoft Hyper-V. This article introduced a web integration between Salt and libvirt called Salt-virt. Has anyone tried this interface for managing images? Does it work better than the young integration between Vagrant and libvirt?
So - is it just me or is there seem to be a big/huge learning curve for all of these dev ops technologies?
As current 'devops guy' on a django project myself, salt works wonderfully. Salt has states available that let you setup all that software, create the virtualenv you need (including telling it you want to use the requirements.txt that you pulled down with your django project source code - Salt gives me my own little Heroku :D ) and for anything left in those wgets you can throw a block of salt cmd.run calls using specified ordering to enable them to run neatly in the sequence you desire.
I didn't find MCollective hard at all - you just install some debs, a message queue server (Stomp was easiest at the time - it's now deprecated, but surely is not much different to RabbitMQ?) and it Just Worked for me. And there was a great screencast.
Did it get far more complicated since I used it last?
RabbitMQ works but is slightly problematic since the authors have a morbid penchant for not wanting to support anything but Apache ActiveMQ. Ask a question about MCollective and RabbitMQ and the answer you get is 'switch to ActiveMQ'.
Come to that I'm hoping we can start building out a full set of mcollective modules to replace the existing ones that will be fully supported and kept up to date so that getting mcollective running will be as easy as including a class and waiting.
A huge part of this job is ensuring community patches get merged in and contributors get treated as I would like to be treated when contributing to a project. I hope we can reverse your experience with modules within a few months (I took this job because I've been in exactly your position, grabbing official modules and having them not work at all!)
The unqualified assertion that Chef uses ssh is inaccurate. You can run chef-solo via ssh if you like, but you'll run into the same scalability ceiling as with any other ssh-based solution.
Powershell works great for executing commands on arbitrary servers (which sounds like the basis of Salt), but it'd be great to declaratively say "I want the server in this state" like the config management side of salt. I assume there is a tool built atop of Powershell like this somewhere?
http://projects.puppetlabs.com/projects/1/wiki/Puppet_Window...
http://wiki.opscode.com/display/chef/Fast+Start+Guide+for+Wi...
It is easy to fall down the rabbit hole of trying to implement things in a CM tool/Powershell combo that could be done in AD far easier.
I am dying for the chef/puppet/salt/ansible/cfengine recipe that will let me fully configure this stuff, including the domain memberships.
That said, on-topic, I just wanted to say that having tried Puppet, Chef and Salt, I've found Salt the easiest to use. Straightforward installation (no messing with Ruby versions/rvm/etc.), really simple setup (systemctl start salt-master; systemctl start salt-minion; salt-keys -L; salt-keys -A yourbox; done), and the YAML-based configuration syntax has been a breeze to work with.
Really quite pleased with it; it's made getting a few of my hairer boxes under control much easier than I expected (and much easier than I found with Chef or Puppet).
[ControlTier](http://www.controltier.org/) had (don't think it's actively developed now) options to execute general system commands, configure systems and application deployment. But it was fairly complex and required [ant](http://ant.apache.org/) skills.
So, take ansible. Primary use: push. But has ansible-pull.
Look at puppet.
Primary use: pull. But has mcollective.
IMO, and I am not there yet but soon to be. The gold standard is to combine two strong players that specialize one each in push/pull. For me, it is looking like ansible/puppet.
In the past few months I've been slowing converting to SaltStack and it really is everything I ever dreamed of for a CM system. Fast, easy, real-time. Lovin' it.
http://docs.puppetlabs.com/references/latest/type.html
http://docs.opscode.com/resource.html
https://cfengine.com/archive/manuals/cf-manuals/cf2-Referenc...
require 'rake/remote_task'
set :domain, 'abc.example.com'
remote_task :foo do
run "ls"
endIf you're just running shell commands, it's easy to screw up and waste your time or break your server by accidentally having the same commands run twice.
http://sysadmincasts.com/episodes/8-learning-puppet-with-vag...
Puppet has MCollective (with ActiveMQ) to implement the similar feature.
1 - http://docs.saltstack.com/topics/tutorials/states_pt3.html
In fact, it is the most flexible software I've seen recently.
There's even a renderer for pure python, allowing one to write config files in python for more control and flexibility.
One wonders if you've ever actually tried Salt. It really is OK to have an approach of: "I like Chef and it works for me." More power to you. Everyone should use the tool that works for them.