- Running a job against a single host will finish in 3 minutes... running that exact same job against thousands will take well over an hour and max out your machine.
- Running against more than around 3k hosts will somehow consume all 60GB of RAM and trigger the oom-killer
- CPU usage on the ansible runner is absurd for a large amount of hosts. We're currently using a c4.8xlarge (our biggest box) just to run deploy jobs and have them finish in a reasonable amount of time (10-15 minutes)
Slicing up our inventory into chunks and running them on different servers sucks big time and is pretty hacky. How do I combine the results? Can't do orchestration like "Run X on these roles first, then run Y on these roles when you're done".
Most likely what I'm going to do is have a single server execute ansible doing only the following in async (aka CPU friendly) mode:
- Upload a current copy of ansible to S3
- Upload the configs to the target machines with ONLY the secrets that role needs in plain text. (I'm not putting my vault secret on every box!)
- Have the servers pull it down and execute in --connection=local mode.
- Wait until each remote finishes
All that said, I LOVE LOVE writing stuff in Ansible. It is so easy to read, follow, and understand. I picked up most of it in a day or two just by reading their "Best Practices" page. Getting it to work at scale hurts though :(