Ansible Simply Kicks Ass
devo.ps
devo.ps
It is throwaway lines line that where you really need to be careful since, no, you don't need to RabbitMQ, solr, couchdb etc. You can just use chef-solo which can also be installed with a one liner (albeit running a remote bash script)
When comparing two products (especially in an obviously biased manner) you need to make sure you are 100% correct. Otherwise you weaken your case and comments like this one turn up.
scp locally dev'd (with vagrant ideally) cookbook to server
Run cookbook
3 steps that can easily be wrapped in a little script (which I know a large company does because I saw their presentation about it on confreaks. Sorry I cannot remember the name of the company or presentation but it was a chef related one).
Still not exactly killing it in terms of complexity. I would avoid comparing ansible to chef-solo in that respect and focus on bits where ansible has a clear (IMHO) win.
Having said that I should say that I have not used ansible and am basing this on what I have read about it.
I wonder if anyone has done a http://todomvc.com/ equivalent for cfengine, puppet, chef, salt, ansible etc.
Something like a simple webserver running Apache with mod_xsendfile, Passenger, Ruby 2.0.0, postgreSQL and a few firewall tweaks (why yes, I AM mainly a ruby developer, why do you ask?)
https://en.wikipedia.org/wiki/Comparison_of_open-source_conf...
https://ops-school.readthedocs.org/en/latest/config_manageme...
I mean all my servers are 70% identical and have the same base, and the DB have a few different packages than the app-servers for an example.
I would hate to have two different chef-solo installations that basically does the same 70% of the time.
If you're interested though, I recently released our stack as a chef solo project at: https://github.com/FineIO/fine-io-chef
I think you're thinking of Ben Rockwood from Joyent: http://www.confreaks.com/videos/1684-chefconf2012-chef-behin...
If I recall, he made a comment to the effect of "Chef solo is great for really small deployments and really big deployments."
chef-solo is a limited-functionality version of Chef and does not support the following:
* Node data storage
* Search indexes
* Centralized distribution of cookbooks
* Environments, including policy settings and cookbook versions
* A centralized API that interacts with and integrates infrastructure components
* Authentication or authorization
* Persistent attributes
(This may have changed since last time I tried it.)
It might be more direct to replace that page with a single sentence: "chef-solo: We're Not Interested And You Shouldn't Be Either."
Or, might we consider compiling a friendly list of the great things chef-solo could do for you? I was going to work on that, but now I'm tinkering with Ansible instead.
But maybe it's more of a standalone tool for others?
You do not need to bootstrap Chef 11. You download a single omnibus package and run a single line to install it. All dependencies are already embedded into the package. Under the covers, it bootstraps itself, but from the end user's perspective, It Just Works.
But yes, you don't need these, Ansible can exchange variables between hosts in memory without needing a database, etc.
I made a similar clarifying request on the Ansible user group forum a few weeks ago but my message was not approved. I believe that these are fair requests and deserve clarification.
Ansible is GPL, and we're totally going to keep it just the way it is, and continuing to make it more awesome. (The application is just invoking ansible-playbook).
Why is it commercial? A lot of the stuff we are building is mainly interesting to enterprises -- RBAC, reporting, things like that.
The core things everybody needs to use Ansible today, all the modules, are free software and we will continue to add loads more there. Ansible received over 50 new modules in the last 2 releases.
Chef [1] and Puppet [2] both chose the Apache license. Salt did as well.
If you were worried about competitors offering hosted versions of Ansible, I could understand going AGPL, but not GPL3.
Thanks!
[1]: http://www.opscode.com/blog/2009/08/11/why-we-chose-the-apac... [2]: https://puppetlabs.com/apache/
Ansible being GPLv3 doesn't affect folks much. Your library modules can be licensed in any way you like, and it's fine to shell out and call ansible-playbook. AWX also exists as a nice API layer if you need a web services interface that abstracts you from the license question, so folks looking to do commercial integrations can tap into the REST layer.
AGPL is more restrictive than GPL, in some ways, as you can't use things in a hosted service. We didn't want to be that restrictive.
As long as there is a REST bridge/layer, any service shouldn't be considered a derived work -- only changes to "core" Ansible would be "covered" by the AGPL.
I'm sure you know this, but it is an important point (just like the GPL never dictates that you have to contribute changes upstream, just to "your" users (which of course might include upstream, and most people feel it is more constructive to give back in the more general sense...)).
I've seen some strange (imnho incorrect) interpretations of the AGPL, hence this comment.
The point is, the beginners never find the tree. They get lost, kind of like I've gotten lost in this metaphor, and they go try something else.
Where we really need to be is a place where configuration management is just another part of application repositories (there's a parallel here to db migrations), so that web applications are completely self-contained. That includes versioning, unit testing, etc. I haven't yet seen any of the configuration management tools come forward with a solid unified best practice solution for treating applications this way, but there are a lot of people working in that direction, which makes me very hopeful for the near future.
http://docs.opscode.com/install_server.html
but when you run:
sudo chef-server-ctl reconfigure
Chef goes into a frenzy of bootstrapping itself , installing a huge stack of dependencies, takes ages, only to fail with an obscure message in the end about one of its dependencies failing to start (RabbitMQ, if I recall correctly).Can I work around the failure? Sure! But what does it tell you about a piece of software that its installer fails on two pristine different and common linux distributions? To me, it tells me that it is over-complex and poorly engineered. So, from a different perspective, I really identify with the "gazillion smaller dependencies" comment.
I helps me little that chef-solo is simple. I don't want chef-solo. I want a full config management distributed solution that is dependable (which probably requires lean).
My next candidate is Salt, and so far it looks good.
SL is basically the same as CentOS: a community version of RHEL.
(I don't think we work at the same place, there aren't any likely Sergio's in the corporate directory)
Stop by the list if you'd like more info, but it's easy to generate nagios configs from templates this way.
But yes, it's easy to do that. Assuming you meant a template, you could do:
{% for host in groups["datacenterX"] %} {% if host in groups["Y"] %} {{ ssh_host_key_rsa_public }} {% end if %} {% endif %}
Ansible is also good at carving up groups based on these things, like if you just wanted to talk to those hosts:
hosts: datacenterX:&Y
note that Hacker News ate my newlines.
etc
If you want to edit the post to include them, indent by at least four spaces the text for which newlines are important. Maybe like this:
{% for host in groups["datacenterX"] %}
{% if host in groups["Y"] %}
{{ ssh_host_key_rsa_public }}
{% end if %}
{% endif %} [sic]What happens if the server crashes? Can a subsequent run before all the data is repopulated cause the derived files to have invalid or missing data?
When you run again it gathers all the facts from the hosts.
This video is a nice introduction: http://m.youtube.com/watch?v=jJJ8cfDjcTc
I haven't used it myself, but I've been reading up on Salt quite a bit in hopes that I can replace Chef with it soon.
As for installing a huge stack of dependencies: if we're talking about Chef 11, there's nothing to install. All of it is embedded into the omnibus package. It is configuring them.
It's too bad about the RabbitMQ. Have you submitted a ticket?
(Disclaimer: I used to contribute a lot of code to salt and still do when I have the time)
Chef client does not require RabbitMQ, Solr, or CouchDB. It too has an omnibus package that embeds all the dependencies and isolates it from the system.
Both server and client are now one-line installs as well.
Actually, can you? I don't mean can you give me an exact lay of the land, but what kinds of problems in this arena are you facing? I presently work in gov't contracting and so far you've defined my exact deployment scenario. That is, infrastructure architecture requirements (hardware/network config) are ever changing, and won't stabilize for some time.
If you can't be more specific as to your problems, maybe you could at least let me know what about Ansible has made this easier to deal with?
http://www.stavros.io/posts/example-provisioning-and-deploym...
Folks might be interested in checking out http://github.com/ansible/ansible-examples for some full-stack use cases showing playbooks that we do have.
site.yml
ansible.cfg
hosts
host_vars/*
group_vars/*
roles/*/{files,templates,tasks,handlers,vars}Just noticed these don't mention "vars" in the roles directory, do need to fix that :)
This article could just as easily have taken the complete opposite view of Ansible by saying things like "parallel ssh sessions don't scale, strong encryption costs too much CPU time; push can never work reliably therefore pull is the only viable model; etc. etc."
I feel one of Ansible's strongest points to champion is the low barrier to entry. It takes minimal understanding to get going, compare that with at least 1 month hands on with cfengine or perhaps 2 weeks puppet before you would consider yourself proficient enough. With Ansible it's 20 mins or so.
I agree with the complexity of similar-but-different hosts, and that's actually something we're set to solve with devo.ps (we'll see if we pull if off).
It definitely doesn't take Facebook sized infra to outgrow a technology though. What if the gamble doesn't pay off? What if you planned to scale the central NFS server dependency by just adding an extra NFS server but have now found there's no rack space left / no budget / a purchasing delay / insufficient network capacity / cooling capacity / power capacity / a.n. other unplanned problem.
For the SSH question, i couldn't reliably get more than 250 concurrent connections outbound from one circa 2008 blade server. From memory that would have had a spec of dual core CPU, maybe around 2.4Ghz with 8Gb ram using PAE as it would have been a 32bit kernel (our spec, the cores will have been 64bit). They were multiply-connected at chassis level on myrinet fabric in one DC and infiniband in another and the resource being exhausted was CPU.
These days all the blade servers are gone but we see an absolute explosion of virtual machines so it's a similar and still relevant problem in many ways.
As Michael (creator of the project) said here; it scales just fine even with serious players (thousands of instances). He has actual hard data and concrete use cases to back it up.
So far, all I've heard from detractors is the "OMG Chef is so hardcore Facebook uses it". Well, they use some. They'd probably do just fine with Ansible instead, provided they were putting half the brain power they put into Chef to make it work for their infrastructure.
You may get very big, just likely not Facebook big.
One of the things many people want to do is rolling updates too, and Ansible is remarkably good at them, having a language that is really good for talking to load balancers and monitoring and saying, "of this 500 servers, update 50 at a time, and keep my uptime". Folks like AppDynamics are using this to update all of their infrastructure every 15 minutes, continuously, and it's pretty cool stuff.
For those folks that do want to do the 'facebook scale' stuff, ansible-pull is a really good option. One of the features in our upcoming product is a nice callback that enables this while still preserving centralized reporting.
Happy to have the conversation, but definitely I've never heard the CPU time compliant. I think the one thing we see is a lot of users are happy that Ansible is not running when it is not managing anything, rather than having daemons sucking CPU/RAM/etc, and folks are actually getting a little better performance from avoiding the thundering herd agent problems.
In particular with virtualization, running a heavy agent on every instance can add up. (reports of the ruby virtual machine eating 400MB occasionally occur).
Many folks are actually not doing repeated config management every 30 minutes, though I realize that may be heresy to some Chef/Puppet shops, there's also a school of thought that changes should be intentional. So there is often a difference in workflow.
LOTS of folks are doing rolling updates, because rolling updates are useful.
Many folks are also using ansible in pull mode.
You could also set up multiple 'workers' to push out change, using something like "--limit" to target explicit groups from different machines.
What happens if you feed Ansible --forks 50 it's going to talk to 50 at a time and then talk to the next (it uses multiprocessing.py). If you also set "serial: 50" that's a rolling update of 50 nodes at a time, to ensure uptime on your cluster as you don't take the 1000 nodes down at once.
This is really more of a push-versus-pull architecture thing, while it presents some tradeoffs, it's also the exact mechanism that allows the rolling update support and ability to base the result of one task on another to work so well.
Ansible also has a 'fireball' mode which uses SSH for the initial connection for key exchange and then encrypts the rest of the traffic. It's a short-lived daemon that doesn't stay running, so when it is done, it just expires.
I think this is a false dichotomy. Those who believe runs should be performed frequently often implement this to revert manual changes performed by people operating contrary to site policy.
My only point is that it doesn't take two weeks to learn Puppet. I'm not saying that Ansible is worse or anything like that, rather I just wanted to contribute another data point.
Of course it's going well. But you should make sure to buy all the people who built all these well-engineered things a tasty drink at the next company get-together, because rest assured: Not every Puppet setup is easy to tinker with.
now I know that using VMs / vagrant is critical to sane server orchestration development workflows though, so there's that.
For instance, I worked for a major computer vendor doing an OpenStack deployment, and watched a simple deployment there suck up 20 developers for six months, where all of that time was in writing automation content.
Repeated hammering out of dependency ordering issues, coupled with the non-fail-fast behavior, and having to trace down where variables came from turned us into automation tool jockeys, so we couldn't focus on architecture and development. The project barely had deployments extending beyond 5 nodes in the end from all of the complexity.
Ansible already existed at this point, but it provided major fuel for me doubling down efforts into expanding it. The goal here is not just the basic language primatives, but making it really easy to find things as you have a large deployment, and making it really easy to skim/audit even if you aren't a really smart programmer.
That all being said, Puppet deserves major credit for pioneering a lot of concepts and revising CFEngine.
While Ansible aims to be a cleaner config tool, but also focuses on application deployment and higher level orchestration on top, so you get some capabilities not found in those other chains (like really slick rolling update support).
I know we were looking at Boxen as a way to roll out environments to our dev machines, and we are hoping that our existing Puppet configs will work well with that effort (since Boxen uses Puppet). Do you think it's at all possible that there will be some sort of adapter to allow Boxen to use Ansible? I have no idea if that would a good idea or not (I haven't looked in to either Boxen or Ansible enough yet) but that's the sort of thing that would likely help steer our decision process.
In the 1year since, Ansible has exploded in capability. It isn't quite as low barrier anymore (more to config language) but still gobs lower than others. While adding much sophistication.
I wouldn't even consider Puppet or Chef anymore. Salt or Ansible. Although, Python guy, so biased more than a little.
Also, none of the "VM lifecycle management" tools we needed like Foreman or its commercial equivalents had any integrations with Ansible.
These concerns left Ansible dead on arrival for us. It sounds like a decent tool if it matches your needs.
Yeah, there's no Windows now. If we do it, I'd want it to push powershell over native means. I'm still not 100% convinced it's something Windows admins want to do in a large enough number to focus on it, but input is always good and would be interested in hearing more about use cases folks want.
As for lifecycle, Ansible has significantly less provisioning integration required than some of the other systems because all you need to do is get your SSH keys in there -- less bootstrapping. It also has a pluggable inventory system so you can get your list of hosts/variables from things like EC2 or OpenStack, so with things like cloud-init, that is usually what folks need. And there are some modules for spinning up new guests in various systems. (I'd like to see a vcenter inventory plugin too).
I'm also the original author of Cobbler and there's integration with that on the provisioning side, and for something else similar, it's largely the case of writing an inventory plugin and making sure the key is installed.
I think this is much more common now that windows in the cloud is quite easy. (My company seems to have as many windows instances in EC2 as linux.)
I do agree with your preference for powershell over cygwin.
Our Windows folks are big Powershell fans here. You wouldn't find too many complaints with it. The biggest problem going with Powershell would be bootstapping it into the environment if it's not already installed or if you depend upon a more recent version.
Since we're actively evaluating lifecycle management tools right now, this is our main concern. Many of the integrations to provisioning tools flow downhill from frontends which manage the state of many systems. We find ourselves choosing the best-integrated stack rather than evaluating provisioning software. The plumbing is very important, but a well-written lifecycle tool should make the copper pipes less visible.
Powershell is the obvious choice, and, as of v3, I think it would be a viable transport for Ansible. If you or anyone on this thread wants to put some ideas on paper for a PoC, I'd be glad to help. Email in profile.
PS. Thanks for your work on Cobbler. We used it recently for automating VM template builds.
I want to choose one of them but can't decide which one to pick.
I would try writing the same recipe in each of them and see which works best for you. All of them can now support easy, domain-specific languages in JSON or YAML.
Puppet is the oldest and has the largest community. Chef is almost as old and large and seems most popular with the Rails community. Salt is the youngest and has the most active contributors.
Foreman, Redhat's ManageIQ EMC, UrbanCode's Terraform, CloudBolt C2, vCloud Automation Center, BMC's Cloud Lifecycle Management, and IBM's SmartCloud Orchestrator.
Most of these use libvirt and one of Puppet, Chef, or Salt under the hood.
CFEngine is by far the most pull-based tool as it is underpinned by a theory mandating this behavior. Puppet is pull-based but with more push, Chef is even more push based, and Ansible and Salt are (I guess) mostly push.
In the end it depends on what is more practical for your situation. If you have a few machines, then more push will be practical, but the more machines you have, pull-based solutions scale better.
If you want solid reporting with pull mode, we're going to have a product release in the next month that has a cool callback based feature, where you get centralized reporting and nodes can still phone in any time and request configuration of themselves. Easy to bootstrap, it will just require a one-line wget to invoke a configuration request.
I would not characterize SS as being push based in the same way, one of the things Ansible is really good at doing is orchestrating things "like an orchestra", so if you want to talk to the woodwinds, brass, percussion, and then go BACK to the woodwinds, you can. It's not simply saying this-than-that, etc. This enables things like the rolling update feature, and doing some pretty intricate multi-tier work.
One of the major inspirations for Ansible was when I was building Func at Red Hat and wanting to build a multi-tier architecture management system. I still strongly believe you need a push based system for that, because you don't want to wait 30 minutes + 30 minutes + 30 minutes for all of those changes to stack up.
Some other tools try to solve this by glueing layers on top, with Ansible, at least the intent is to build that in.
So yeah, I think there are definitely times when you need both.
A senior sysadmin once told me that he asked all incoming recruits about what this proportion generally should be. The answer: 1/6 (16.6%) push and 5/6 pull (83.3%). I found that answer slightly amusing as it seemed overly general. On the other hand, he felt very strongly about that point and was quite annoyed with juniors' natural inclination to use push by default when pull could practically be used.
edit: docs indicate it's still a problem http://www.ansibleworks.com/docs/modules.html/#apt-repositor...
I assume this is a rhetorical question :-) A number of folks on HN are either in a full-time or part-time dev/ops role and so tools that are developed in this space are interesting.
Unlike other sites there aren't specific subcategories of interest on HN so you get the devops stuff with the politics stuff with the new language stuff with the economics stuff with the funky physics stuff with the the i-did-a-thing stuff.
Ansible is an example of what is part of the Web 3.0 stack, dynamic but durable provisioning of resources out of a third party infrastructure.
It's also the first open source project I've committed code to. Huge fan. I use it to run a small 8 node openstack system, as well as deploy Apache VCL.
I'm using a single set of playbooks for my infrastructure with Jenkins to check out changes and push them to production machines. The same infrastructure definitions are used with Vagrant to give each developer an exact copy of the production system.
Ansible also has a --ask-pass and --ask-sudo-pass if you want to provide a password.
Using ansible from a local machine is fine, because you can make your devs type in passwords and etc, but I can't think of a secure way to do it with continuous integration.
It use agents ("minions") and is architected very differently. Ultimately, we settled for Ansible because it uses lean and out-of-the-box tech, doesn't push dependencies on remote hosts and uses YAML extensively (close enough to JSON, actually I believe JSON is a subset of YAML).
So far I'm coming down on the side of ansible. The config is a marginal improvement on salt (which is already pretty good) but I think the clinchers is the explicit push nature rather than the declarative state declarations in salt.
I'm far from expert mind you and not using either in a serious production setting, so don't take my word on anything!
Ansible advantages:
+1 no client required. Python libraries may be required on the client machine to take advantage of certain modules (e.g. Python's postgresql library to use some parts of the postgresql module)
+1 roles. A nice, best practice way to package up logical configurations (e.g. files, templates, tasks).
+1 separate host files for prod/uat/dev/whatever you like. Keep all generic configuration tasks/roles host-agnostic.
+1 very easy to write (and distribute) additional modules
+1 uses SSH for all transport. I love that it uses something tried and true, secure, and robust like SSH. Based on comments on Salt's Github, Salt may be receiving something similar soon. A ZeroMQ mode (aka Fireball mode) is available in Ansible for speed.
+1 sticks to Linux (Unix?) only. I like that Ansible is not targeting Windows (for now?); The Windows ecosystem is a very different community; with its own mindset and tools, and they'd prefer to use their own vs. something from Linux-land. This also allows Ansible to concentrate on getting really good on Linux/Unix.
Salt stack advantages:
+1 (+1 +1 +1) allows you to use Python (with Jinja2 templates) to express logic in configurations. Some people are against this (and tools like Puppet and Ansible try to hide programming from the user), but I feel this is a terrible idea. All that you end up doing is writing a configuration language that's Turing complete [joking!]). Some people will say: logic doesn't belong in configuration! But when you need to express conditions or loops, let me use a (standard, i.e. Python) programming language. I don't want to have to learn your DSL (which only your specific community knows about).
So in summary, I'd love a tool, that has:
- Ansible's "no client" setup
- Ansible's logical way to package up configurations (roles)
- Ansible's separate host file setup
- Ansible's simplicity of writing (and distributing) modules
- uses SSH (currently only Ansible, but might be available in Salt shortly)
- Salt's templating being available in Salt states (Ansible's "playbooks"). This point alone puts Salt on even ground with Ansible. I cannot stress this point enough. This is the reason one of the main reasons I don't use Puppet and Chef.
We had one manager describe it as "automation even a manager can understand". I'm a developer, but I'm jaded about software development, and I definitely agree with the approach Ansible (and Puppet) take about describing infrastructure with a model, rather than making it a program. I'm not really even trying to make it turing complete, I want to create a language that allows modelling how you want to describe the datacenter, but it's not trying to be a programming language. In fact, I'm rebelling against it trying.
There's a difference to a developer vs sysadmin mindset, but as something in between, I want to write applications code when I'm writing code, rather than writing infrastructure automation code. Code should be minimized, and that makes things more reliable in the long run, and enforces consistency IMHO.
I'm a little unclear about what you say about putting host specific information in Ansible roles. Pretty sure you can't do that, but I may be misunderstanding the question :)
About the second point (host specific information in Ansible roles): I was referring to this: http://www.ansibleworks.com/docs/playbooks.html#roles -- I was (incorrectly) under the assumption that you could put host file variable overrides (as you would in host_vars) under the vars/ section, but it looks like that's more of a generic, non-host, non-group specific variable directory? I have edited my comment and removed the incorrect section.
[1] http://www.ansibleworks.com/docs/playbooks2.html#conditional...
[2] http://www.ansibleworks.com/docs/playbooks2.html#loops
[3] http://www.ansibleworks.com/docs/playbooks2.html#variable-fi...
Can ansible do this? Can I easily migrate a playbook from using say apt-get to pacman?
To see the package managers Ansible supports OOTB look at:
http://www.ansibleworks.com/docs/modules.html#packaging
For examples on how this can be abstracted to handle multiple platforms see: http://www.ansibleworks.com/docs/bestpractices.html/#operati...
I was trying to set up https://github.com/pas256/ansible-ec2 to manage some EC2 instances and it seems like I needed to replace a static config file with a script? Wouldn't that prevent me doing anything that wasn't EC2 with that instance of Ansible?
Debian also moved the inventory files around in their package so that made things confusing.
I wanted to make a language that was good for both, and easy to do ordered tasks, and a lot of app deployment requires ordered tasks and being able to do things on the result of other things. OTOH, sysadmins want a language that is easy to read and has the declarative state features, so it's kind of both at the same time.
Chef (and formerly Puppet) had (for us) a much higher barrier to entry when it comes to getting an unfamiliar dev on board. We basically had to get reqs from the dev, and then we would build the recipes.
I love how the Ansible "Getting Started" guide [0] first says you need to install Puppet 2.6, as 2.4 on CentOS / RHEL is not supported :)
There's a very big difference given the nature of this article. I thought that Ansible was dependent on Puppet and was quite confused.
Python 2.4 is all you need on remote nodes, but on the control machine you do need Python 2.6.
Source ./hacking/env-setup from a checkout and you are running live from there.
So you may wish to just install the python package and then set up your library directory.
Gcc is sometimes needed to build (often optional) c extensions.
Not a lot of data to go on in the above, but if you are specifying -c ssh, ControlPersist is good! Perhaps you were also trying to manage an AWS cloud from outside AWS. Better to install a control node inside the cloud in that case. Otherwise, I don't know, maybe a DNS issue? I'd need to know where your issues were to dig further.
One thing ansible is pretty sharp at is grouping yum transactions, so if you had every package in a seperate line, with_items is a good thing to use too, that can make deployments extra zippy.
Another thing you might be interested in is 'fireball mode' which uses SSH for key exchange and uses a more direct transport for making subsequent changes.