Ansible open-sources Ansible Tower with AWX
jeffgeerling.com
jeffgeerling.com
Your Ansible master is one of the most critical (if not the most critical) machine in your network. It has unrestricted access to everything, including a very detailed list of what and where everything is. If it's compromised, it's over.
Something like Ansible Tower adds a LOT of attack surface. Instead of a locked-down server exposing a public key-only SSH port, you suddenly have a whole web application stack in there. Your browser and every Chrome extension with full page access now has root access to your network (and don't get me started on potential CSRF or XSS vulnerabilities...).
If you don't need any of the enterprise-ey features, just might be better off with plain Ansible and Ara[1], with the latter running on a separate machine. A "sudoers" rule is much more secure than any ACL in a web application backend.
If you do want to use Tower, you need to think about these risks and how you mitigate them. Of course, this also applies to any similar tool like Jenkins, Rundeck, CircleCI or whatever if you give them production credentials.
It only has unrestricted access to what you give it unrestricted access to.
If your ansible playbooks use a user that doesn't have root access and limited sudo powers, then it's not much different from using Ara.
If you're squeamish about that, on principle, then Salt, Chef, and Puppet would be problematic as well. The only logical choice would be to provision and bake all of your images and push the AMI up to AWS...in which case Ansible would be an excellent utility anyway.
Also, this assumes access to Tower is unlimited. ACLs applied in Tower should be able to limit what user has what access.
https://threatpost.com/seven-more-chrome-extensions-compromi...
It's absolutely possible to use Ansible Tower securely, it just takes some effort.
In my experience, ansible playbooks are great when run from a more general purpose task runner like jenkins, which then has permissions to access/modify one's production environment. I don't think I would personally use tower unless it provides something much much better than running ansible tasks in Jenkins... it would be just too much of a hassle to get the security/compliance aspects right.
You can do it with Ansible, if you're going home-roll it, but I haven't seen too many people do so.
(Edit: my bad, ignore this comment, I misphrased my question; the real question should have been: can you run Ansible locally on the target server, just like chef-solo/chef-zero?)
"Serverless" would be more like what I described with regards to Chef Zero, where a machine, as it bootstraps, is able to fetch its playbooks from somewhere and self-execute with some sort of sourceable configuration data. The standard Ansible workflow is not only not "serverless", but it is antithetical to cloud-friendly scaling and fault-tolerance practices. (Think about how you're going to manage auto-scaling groups with it. It hurts.)
1. https://github.com/edx/configuration/ 2. https://github.com/edx/configuration/blob/master/util/vpc-to...
For my purposes, Packer isn't always available, which is a bummer. The Chef Zero approach I described above is nice for my purposes because it works with either an AMI or a live instance; when I write cookbooks I break them into "bake" and "configure" recipes and sub-recipes, and the "bake" steps are effectively memoization of steps I also run (idempotently) when a machine comes up.
It also means if any of your servers is compromised, the attacker will be in possession of your whole playbook.
Which should, if you are designing your systems wisely, not matter at all, in any way, because configuration management systems should not contain secret or sensitive data (which should be provided from a more secure option--Credstash, Vault, whatever).
I could open-source all of my Chef cookbooks or Ansible playbooks and not have a security or correctness concern; this is pretty straightforward design.
But, as far as serverless/self-bootstrapping deploys go, it's less common. Ansible has less of a "culture of dependencies"; the simpler, more approachable-looking nature of the Ansible playbook format seems to lend itself to people one-offing whatever they need rather than looking for best-practices solutions that already exist. Because of this, there's no real Berkshelf equivalent for Ansible. The tooling doesn't exist, outside of Tower (sorta), because nobody wants it, and nobody wants it because the tooling doesn't exist. So the people who are doing with Ansible something similar to the Chef Zero stuff I mentioned above are mostly home-rolling it. (I just use a S3 bucket as a Minimart berkshelf endpoint and move on with my day.)
Last-mile configuration is also tricky. In my Chef Zero stuff, I use CloudFormation metadata to provide Chef attributes. You can do something similar with Ansible...but it's duct-tapey. There are times when simple is better; IMO, Ansible's core tooling errs too far on that side and the ecosystem has not caught up to make more rigorous approaches really viable.
https://galaxy.ansible.com/intro
There are tons of roles available, for just about everything, and the quality isn't always great, but still higher on average than what I've found for most Chef recipes and Puppet modules.
And that's not to mention the very high number of high quality modules that are builtin to Ansible.
I would strongly, strongly disagree as to the quality of most Ansible modules that I have dealt with, but it's probably more based on exactly what you need than anything else.
We use CloudFormation to deploy, so in the instance metadata we have it run Ansible locally to bootstrap and return the exit status to cfn-signal.
We retrieve secrets via Parameter Store. For environment specific configs that are not secrets (ie passing in vars from CloudFormation), we have cloud-init write a json file that we include with our ansible-playbook command.
The command ends up looking something like:
ansible-playbook -v -i 'localhost,' -c local /some/path/playbook.yml --extra-vars '@/some/path/vars.json' && /opt/aws/bin/cfn-signal -e $? --stack ${AWS::StackName} --resource ${AsgName} --region ${AWS::Region}
In my mind, though, the biggest reason I used ansible was it's simple push-through-ssh nature. I knew that if there are some maintanance tasks I would have to do by sshing from my machine to some server, I could as well create and run an ansible playbook against it. This was espacially useful for configuring ephemeral boxes spawned on i.e. openstack where I know it will have ssh running and that is it.
Ansible also works well over ssh too. It’s pretty flexible.
There's also a balance to be struck between pre-baking AMIs and running config management at instance launch time.
This really is what I do for a living. I'm speaking from a position of entirely too extensive experience when I say that Ansible has no good solution here in common use. If I thought Ansible was good enough for me to be spending a lot of time with (it's not, and I advocate that clients not use it if they have a choice), I'd have probably already had to write it.
As far as machine images go, they are an optimization, not a core system. Your configuration management systems need to be able to bootstrap from either an AMI, to lay on last-mile (configuration, as opposed to setup, stuff) and converge any updates since the last AMI build, or to start from scratch. And that is another weakness of Ansible; writing idempotent Ansible scripts is significantly harder than it needs to be.
I'm having a bit of a hard time 'getting' the handling of secrets, etc. in such a setup.
(We're currently using Ansible, and frankly it's turned into a huge pain -- even on just a few hosts it's incredibly slow and horrible to debug. I'd really like to eventually transition to a more "build-a-pristine base system image" + "self-by-applying-more-recent-playbooks" type approach.)
I find that this is one of the most frustrating things about the Brave New World of immutable/container/VM/cloud/etc. It's actually REALLY hard to separate the hype from actual working things because all you seem to find is the hype and... 50 page "tutorials" on how to set Kubernetes[1] up.
[1] Random example, but the mere fact that they use a "simplified" admin program (minikube) for the tutorial tells me volumes about how fun it's going to be and how little administration it's going to require in production. /s
The fun part is, I think most of the sexy tools are lousy. Ansible demos well but works poorly; Terraform has bitten me so many times I don't trust it; Kubernetes doesn't make sense to me in a universe where I am buying already-provisioned-and-segmented resource bunches (we can call them "virtual machines", maybe) where I have to incur extra deadweight loss because fault-tolerance implies requiring extra space in case any existing node goes away.
I have spent enough time thinking about this that I am reasonably confident in my approach, and I would like to share it. It's a little more than a blog, though.
To quote the infinite wisdom of The Simpsons: "Do what you feel" [is right]. :)
With the Jenkins model, can't someone just add a job that dumps the execution environment, and get all your credentials?
There's no webapp stack to attack if you're only able to access it via a tunnel. If you're assuming the machine you're tunneling it to is compromised, there are bigger issues at play - ones that would compromise even a plain ssh link.
I'm talking a direct tunnel from your ansible master to the host you're planning to use it on, not say, into your company's network at large.
It can then do everything required with a few exceptions.
The exceptions are creating init/systemd files (associated startup/shutdown) and creating the necessary top-level directories for new apps.
Those can be allowed with a few careful sudo rules.
It can be a little convoluted but much more secure.
* "gpasswd -A <user>,,, <group>"
Note: to be a group admin doesn't necessarily mean you're in that group; it means that you can put yourself in and out of that group as required.
Couple in a sound architectural model, and you're good. Unless you're working on some super duper top-secret stuff...but if we're talking about compliance and security at the CIS standards level, to the best of your ability and knowledge, you should be good. Keep in mind, of course, that large, blanket vulns show their faces every so often - like the KRACK vuln - which renders all of our preparation and paranoia mostly null.
Example (of a neurotically paranoid architecture): Dedicated VPN resides in a dedicated AWS account and fronts all traffic to all hosts in all of your organization accounts across AWS. The only services exposed publicly other than your VPN service are the public-facing services you run if you're hosting some SAAS, for example. Even then, your LB's better be public, and your EC2's behind them better be private. Yay port 443 and LB to instance certificate encryption.
Ansible Tower lives in some other "Internal Services" AWS account, and all traffic ingressed to hosts in that account is fronted by a bastion host and traffic proxy. The bastion/proxy is governed by a set of strict VPC route tables, and security groups that are set to only ever permit or accept traffic that corresponds to the IP addresses of your VPN.
If you wanna get even crazier than that, you can also strictly control egress rules for your security groups - so even if someone got in, they'd be hard pressed to get the data out of your systems without doing some acrobatics.
In general, if vulnerabilities @ the Tower server level are your concern, addressing that with architectural best-practices and network-level controls for locking up access and traffic to Tower is reasonable and gets the job done. It would be the same as securing anything else that's sensitive running in your infra, like Jenkins or RunDeck.
I think RedHat's done an excellent job, and a huge service to the community by finally making good on their promise of open sourcing Tower.
Full disclosure: I'm a Software Engineer at Red Hat.
Tower has a lot of great features but I hate to concede that it does come with some of the drawbacks you mentioned. I think it comes down to weighing the pros and cons, making sure you are aware of the cons and putting measures in place to avoid problems.
It also heavily depends on your use case, Tower provides things like RBAC/ACL, auditing, scheduling, online editing and execution, an API, etc. If you happen to need none of that and you're perfectly happy with just using Ansible from your command line, there's probably little incentive for you to use it.
If all you need is reporting, ARA is simple, easy to install and setup, doesn't get in your way and just records things transparently.
I figured I might as well take the opportunity of posting to leave a video demo [1], even if a little outdated, as well as an example of live report that ara provides [2].
Let me know if you have any questions !
[1]: https://www.youtube.com/watch?v=k3i8VPCanGo [2]: http://logs.openstack.org/21/474721/7/check/gate-openstack-a...
That said, I'm having a hard time seeing how a compromised browser access AWX could give anyone full access to everything. Attackers could effect anything a playbook has access to, but considering how varied playbooks are, it would be really hard to guess the right values to pass into a job to get what you want.
Can someone elaborate on how an attacker could do real damage using a browser based attack?
So, can you explain how an attacker would figure that information out? Assuming the playbooks are stored in source control like the Tower docs recommend, and all your variable values are mostly in source control as well.
That being said, I think you are overstating some of the risks. You don't need to grant every Tower user root access to everything on your network. If you're at a scale where Tower makes sense, you probably already have some sort of separation of privileges. I agree that a malicious Chrome extension could do a lot of damage - just like it could with all your other management tools like DRAC/ILO, network equipment GUIs and so forth. Yes, every web application carries the risk of CSRF/XSS or other vulnerabilities, and Tower is not immune to this, but we do spend a lot of time worrying about it, conducting audits, etc.
If your operation can succeed with nothing but a sudoers rule and command-line Ansible, then by all means use that. Nobody wants to force AWX/Tower on people who don't need it. But if you do need the feature set of Tower, I think it's one of the safer options available.
- Github Project: https://github.com/ansible/awx
- Project FAQ: https://www.ansible.com/awx-project-faq
- Demo Video (10min): https://www.ansible.com/tower-demo
- Remove angry potato logo: https://github.com/nanobeep/awx-logos
I do the curation for the 'Ansible & Friends' newsletter and we've been following this closely for a few months...
If you're interested in alternatives to Tower/AWX, see our listing in issue 64: https://hvops.com/news/ansible/64
Kudos to Jeff, he's quite prolific in the community. We've highlighted him recently in the "Community Heroes" section: https://hvops.com/news/ansible/64 And in our September issue, we had to give Jeff his own section because there were just too many great articles he had put out: https://hvops.com/news/ansible/67
(Full disclosure: Ansible Inc hired me a few years ago to update and automate the Tower documentation systems)
This is strictly speaking correct but I think the author is implying that AWX code will get proprietized in Ansible Tower, which is not correct. Normally, Red Hat does not alter upstream open source licenses in its product offerings.
I gave Jeff's (author of this piece) Ansible role a try and had an AWX system up and running quickly. He's making some really nice Ansible roles.
But I had no luck with AWX. I think if you are using entirely platforms that are supported directly by AWX (OpenStack/AWS/Azure), you might be fine. We run 80% of our stack on Ganeti, and have a custom inventory script that AWX seems entirely unable to use. Even distiling that dynamic inventory down into a static inventory isn't working in my testing because it wants our Vault password to load the inventory, and the inventory task doesn't allow for associating vault credentials. Nor does it seem to provide the boto credentials needed by our dynamic inventory script, even though it has a spot for providing AWS credentials.
I want to try it one more time to see if I can find a combination of static inventory that doesn't need the Vault.
But at the moment, the Rundeck weirdness with our inventory is less than the AWX weirdness, so that is our solution. I'm sure I'm just not getting something about AWX, but I've spent a day with the docs and made no real progress.
It's awesome to have available as an option. I wanted to get Tower but could never secure the funding for it, largely because it was a big unknown.
You _should_ be able to pass in a vault password as well, as part of any job? Are you doing something with vault in your inventory?
There are still a couple small bugs with my Docker image (if you're testing with that) that I just haven't had the time to work out; you might be running into one or two of them :(
And if you do use a cloud provider, you likely want to run Ansible in the same cloud, just for performance reasons.
Arrogant as it may sound, I'd much prefer to take ownership of that, personally. I'd be happy to self-host it and take the operational pains that come with it, and sleep (at least slightly more happily at night) knowing that I'm managing my secrets myself.
I'm guessing AWX wants to be the central "source of truth" for all of that stuff instead of just keeping it all in a single git repo.
(edit: disclaimer .. I am a Red Hat employee!)
I use Ansible in my day-to-day job (I'm a full stack dev learning how to do sysadmin stuff in scale). We use Ansible in a custom webapp to orchestrate bare metal clouds. We build and operate custom High Performance Computation Clusters and this is our SaaS offering.
But the code relies heavily on cron and legacy perl code.
I wanted to rewrite the code and use a queue for a long time, but never found enough time to actually do it. It works good enough and we have customers relying on it.
Click on the first link in this article, "Ansible Tower" and you'll be taken to a page [0] where (above the fold) there are two videos: a two-minute overview and a 10-minute demo.
I used to send a whole bunch of fixes to Ansible as they were so responsive in merging my stuff extremely quickly. But after they were bought (or /incredibly/ close to the same timing..) my PRs stopped being merged in a timely manner, often requiring many rounds of tedious rebasing as they looked at it 2 months after submission when the code had moved on. Naturally, I prefered to spend my time elsewhere...
I've seen some discussion in GitHub issues about that; hoping we see it supported soon!
So in a very basic use case where you want to configure a new compute instance in your environment, there's not much difference. However, in cases when you need to do things like hardening enforcement and remediation, end-user self-provisioning with surveys, workflows involving disparate tools which Ansible can coordinate (which may not even involve your hosts), reviewing/analyzing results of a play across 100s or 1000s of hosts, or auditing/rbac controls, Tower/AWX will be your tool. These features will most likely not make their way into Satellite because that's not really the focus of Satellite.
[1]: https://access.redhat.com/articles/1343683
Quick correction: after looking at some docs on the RH site, it seems the Ansible Foreman plugin may not be included in Satellite at this time. I suspect this will change in the upcoming versions, but not really a way to tell at this point :(
Edit 3!!: Ansible support is coming but not for a while. See discussion regarding the roadmap presentation from Summit here: https://www.reddit.com/r/redhat/comments/6lb460/satellite_63...
- self healing - automatic failover - Docker centric tools
You can have a python s*it show where you use containers like processes. In fact they are processes that probably can't handle signals or pipes. Maybe use tini and that at least works, but it misses the point. Ansible and docker are fundamentally at odds.
When ansible was created it had the promise of a simple layer of abstraction over ssh. now It just does all the things because it jelly. Don't give into the h8
http://docs.ansible.com/ansible/latest/docker_container_modu...
Ansible fits all the automation gaps that exist between all the different infrastructure tools.
I've had issues or a lack of features with ansible's docker support, but use it otherwise for pretty much everything.
I just don't see a point to moving to a wrapper than can never have all the latest support / options that upstream has.
The nice thing about using Ansible for all the management is consistency (especially in mixed container / not-container environments). The nice thing about templating a compose file is being able to track the latest Docker / Compose versions more easily.
Also, not every org has (or would want) to go 100% 'all in' on containers. Any hybrid infrastructure will always need configuration management tools like Ansible, Chef or Salt.
a tool for guys that setup your docker, k8s and Swarm.
curl -O https://raw.githubusercontent.com/geerlingguy/awx-container/master/docker-compose.yml
docker-compose up -d> To install AWX, please view the Install guide.
which links to an installation guide[0] with information on setting it up with OpenShift or Docker.
I'm not sure what part isn't clear.
You can use ansible.verbose="vvvv" if you're calling the play via vagrant's ansilbe provisioner.
In that case, running ansible verbosely gave enough info to point me in the direction of fixing what was wrong with my ssh setup (can't remember now what it was).
I think the biggest advantage (in your case), is Ansible's idempotent nature. It takes way more work to make your shell script idempotent that it does writing a playbook for the same thing.
I'm also curious about your current dev environments. Do you call Ansible through Vagrant or are you using completely Ansible to spin everything up? Getting the runtime environment installed/configured correctly can sometimes be a challenge for people just starting out (it was challenging for me at least) and if you have 10s/100s of engineers trying to set up Ansible properly and run playbooks themselves, you may have a bad time.