Fucking Shell Scripts
fuckingshellscripts.org
fuckingshellscripts.org
In many cases, Ansible is even easier than shell scripts. I wrote a post about this a few months ago: https://devopsu.com/blog/ansible-vs-shell-scripts/
I completely understand where the sentiment is coming from though. I wrote a comparison book on Puppet, Chef, Salt, and Ansible a few months ago and am currently finishing the 2nd edition ( https://devopsu.com/books/taste-test-puppet-chef-salt-stack-... ). Even for an experienced sysadmin, using Puppet and Chef to do even a trivial project (replacing a ~10 line shell script) took a painful couple of days. Why? They're overly complex and have confusing broken documentation (that mostly haven't been corrected even 6 months after I gave them a full breakdown on the issues). Salt was pretty smooth, but Ansible was downright easy.
Using a shell script to set up a server generally indicates that it will then be managed manually afterwards (sadness and despair!).
A huge advantage of using a configuration management (CM) tool is that they're "idempotent". Idempotency basically means that you can run the directives over and over again safely.
An idempotent command will verify that the system is how you defined it and will only make changes to bring the system back into alignment with what you defined. That means you can define your system in the language of the CM tool and use it not only for initial system setup, but also for monitoring, updating, and correcting a server's configuration over the life of the server.
A CM tool can ultimately act like a self-healing test suite for your systems - neat!
Your systems are the "app" that your app runs on. They're the foundation. Not using a CM tool is like not having any tests for your app. Sure, it might seem faster at first, but you'll pay for it in chaos, slowness, bugs, and misery later.
Shell scripts are a great first step, but if you're serious about your systems, you need to be using a CM tool. Modern ones like Ansible are simple, easy, and have great docs so there's little excuse left for not using one now.
I'm genuinely curious. Why? Ansible, in particular, provides the same value (easy) but has existing modules to do a bunch of the things you'd have to script.
If you deploy more than a few systems a year and don't have a PXE boot environment (or at least the equivalent of a network accessible kickstart config to be manually selected) or a golden image to deploy from, then I can see how CM tools may seem a pain, because you haven't tackled the initial manual pain point yet, the actual install.
I haven't really used any CM tools, so maybe I'm getting the point where kickstart would traditionally leave off and ansible or chef would take over slightly wrong, but I can't see it being all that complex to automate configuring one after install.
There's a difference if you want to run the scripts automatically on install (non-interactively).
Sometimes you're working on a scrappy prototype and it's just quicker to use shell scripts initially and then go back to solidify the setup into CM scripts afterwards, once you know what you actually need.
Of course that approach will fall to hell if one never gets back around to making the CM scripts.
For Ansible, you'd need python and Ansible. I'll assume the official iso already has python (it's a full iso after all) -- for ansible you'd need to install it to ram (eg: (optinally virtualenv) and pip install), or modify the iso (say with[1]). Then you'd need a (optionally set of) script(s). Then run that script through ansible. So essentially everything could be the same, but with a lot of the logic of your current script handled by ansible.
For many cloud environments, it'd be costly (in time and energy, not necessarily $) to replace all 1000 virtual servers for an update. With Docker, you can essentially do that in a trivial manner. I'm still learning Docker and my understanding is still a bit weak, but it's an exciting development in this regard.
Well not blow your brains out difficult but "learn a new syntax, behavior, rules, depend on a new package" difficult if a few shell commands is all you want to do.
I can see where they are coming from. I can go a long way just using shell script to configure (and yes, you can make them idempotent too).
From your site:
> You already have a religious passion for a particular CM tool
Aren't you doing the same against Chef and Puppet then? It sounds like. Oh "comparison" -- yeah just use Ansible.
Agree totally. If you're doing something tiny, then a few shell commands are what is needed, not a CM tool.
I'm speaking mostly about serious systems that businesses run on.
Ansible is not for everyone. Each tool has strengths and weaknesses. I generally push Ansible because it's the easiest to get started with, but can also scale to 10K+ nodes. If something simpler/easier comes along, I'll recommend that instead.
I suspect some combo of Docker and Ansible to ultimately be the simplest set up (in the near-term), but I'm actively learning that and not confident enough in it to be able to suggest it to newbies.
More focus on the "why", less on the "how" http://thechangelog.com/ansible-docker/
So I would suggest that it's natural and OK for someone's evaluation process to yield a favorite (apparently in Matt's case, it's Ansible), without it becoming a "religious passion."
I got most of my inspiration from Kathy Sierra and Why the Lucky Stiff. They're waaaay better than me in this regard, but they taught me that most brains love a little whimsy. I try to add a little since systems engineering can slip into dryness pretty quickly if you're not careful.
Little jokes like the borg cow can lighten things up: https://devopsu.com/newsletters/ansible-weekly-newsletter.ht...
An added benefit is that whimsy helps filter out trolls. Adorable puppies, kittens, squirrels, cows, etc have a way of turning trolls away or at least softening them up a little :)
They're idempotent only as long as the configuration stays the same. If you removed this "apt: package=XXX state=present" line, XXX is not going to be magically uninstalled, but you may have a nasty surprise the next time you attempt to provision a second server with the same configuration and realize you're missing a runtime dependency. That's what I find fascinating about NixOS/NixOps: you can describe the entire state of the machine in a single configuration file (if you use the declarative style of package management). Of course, it won't help you if:
- you already have a large number of non NixOS systems
- you need many packages not present in NixOS
- you'd like polished package management (the Nix package management still has a way to go before it reaches yum/apt ease of use)
http://stevenjewel.com/2014/03/puppet-undo-and-puppet-audit/
The more raw/command/shell modules you have to use, the uglier the whole approach seems to me - and maybe not worth the trouble in the first place. If it can't be done 100% 'right', should I bother?
I put my efforts on hold so far. Ansible improved greatly (especially the documentation) lately, but .. it still feels cumbersome and hackish for my usecases.
Yeah, that's a problem. It's gotten a lot better with the 1.5 release.
Example: we just made the light-speed jump from 1.1 to 1.5. While re-doing a role yesterday, I noticed we now have a module for ec2 snapshots.
So _next_ week I'm replacing a 50-line bash script with a five-line playbook. Which will be run by Jenkins.
A few thoughts: 1) This project is incomplete and is more of a concept than anything. We wrote it literally in an afternoon, and were giggling like schoolgirls the whole time.
2) There is value in what Ansible provides. It does a lot of things for you. FSS did not replace ansible where we work.
All that being said however, I feel like this (or a project like this) does have merit, if executed properly. I feel that once you get used to how FSS works, it might be a good solution for you. We intentionally tried to keep it as simple as possible.
For my team, I decided on Ansible because it was the simplest option (no agents to install, pure ssh).
Lately I've been researching a way to automate things for a bunch of vps I own (nothing important, mostly self-hosted services like tt-rss), and obviously I looked at Chef/Puppet/Salt/Ansible and the likes. While those are valuable and awesome tools, I feel that they're too complicated for my simple needs.
FSS seems exactly the right compromise between running everything manually and using a full blown management tool.
I wonder if this is what you guys had in mind when you came up with the name
i've used both chef, and puppet. and i still have a puppet standalone bootstrap script that i use every now and then. but I too created a simple shell script bootstrapping process for one of the projects i was working on.
it pretty much boils down to this:
tar cjf bootstrap.tar.bz2 VERSION library $2
scp bootstrap.tar.bz2 protonet@$1:/tmp
ssh -A -t myuser@$1 /bin/bash -c "
mkdir -p /tmp/runscript
cd /tmp/runscript
. ~/.profile
tar xjvf /tmp/bootstrap.tar.bz2
bash /tmp/runscript/$2 $3
rm -rf /tmp/runscript
rm /tmp/bootstrap.tar.bz2"
so far i haven't really found a good reason why we can't just have a library of generic bash scripts doing things.i have to apologize about the crappy function style, and the inconsistent shebangs, but meh, whatever.
It does in fact make me smile a little bit when the puppet installation script is longer than the script to do whatever it is that needs to be done to install the app itself. :)
Over time, it may make sense to move to a more deterministic solution, but to start there (IMO) is often a case of premature optimization.
> Over time, it may make sense to move to a more deterministic solution, but to start there (IMO) is often a case of premature optimization.
I mostly agree with this though.
It depends on what you're doing, I suppose. Me? I wouldn't want to try to use "simple" bash scripts to manage a large (or even medium) app farm, especially when those servers need ongoing management.
I sent this along as a tongue-in-cheek 100% Solution.
And.... I actually read the site after I sent this along and realized its actuallyfuckingcoolshellscripts.org
So... thanks!
Edit: removed speculation about an idempotent *nix shell.
The Unix philosophy is a set of tenets based on assumptions that held in the 1970s and 1980s, but don't really hold today. One of these tenets is, "make the implementation as simple and as correct as possible; it is better for an implementation to be simple than to be correct."
This may have been true in an era when every site had to roll a homegrown solution for pert-near everything that didn't come with the base OS, but in this era of open source it is far more important for the implementation to be correct, because the Right Thing can be written once and everybody can use it, and save themselves the accumulated hours of frustration incurred by simple-but-subtly-incorrect implementations.
This is why the Linux world is standardizing on systemd -- to get AWAY from Fucking Shell Scripts and towards a more deterministic, declarative model of what we want done. In the case of configuration management, what you want is a tool that accepts a description of what the system configuration should be, diffs that against the current configuration, and enacts a plan of changes to get from point A to point B automatically. NixOS seems to be a good step in this overall direction.
Fun fact: I used to do instancing of robotic control computers running a specialized version of Debian with Fucking Shell Scripts (FAI, to be exact: http://fai-project.org/ ). It was pure hell. We would have killed for a more deterministic solution.
Conversely, complex is almost always incorrect.
How many more times must it be said? How many more times must Apple win -- and win big -- before open source nerds get the message?
http://www.jwz.org/doc/worse-is-better.html
> This may have been true in an era when every site had to roll a homegrown solution for pert-near everything that didn't come with the base OS, but in this era of open source it is far more important for the implementation to be correct, because the Right Thing can be written once and everybody can use it, and save themselves the accumulated hours of frustration incurred by simple-but-subtly-incorrect implementations.
The BSDs and linux are open source, but we still must check for EAGAIN.
Having been a Unix admin since long before Linux existed, I find this statement to be complete bullshit.
[0] ("Correctness-the design must be correct in all observable aspects. It is slightly better to be simple than correct.")
Me? I like perl and I decided I'd write something that just executed "primitives" locally. Then setup a flexible system of pulling them from git, via rsync, etc.
My tool is largely ignored, but modelled after CFENgine 2.x, and is here:
"Wanna just use fucking shell scripts to configure a server? Read on!"
Step 0: Install the gem
# gem install fucking_shell_scripts
# gem: command not found
wat.If I want to use FSS to configure the server, I already do. I don't need the "latest fad in configuration tool" to do something FSSs already do!
Like fabric. If you're not using their advanced features, you're writing a shell script in python.
Thanks, I'll write a shell script
* Depends on Ruby.
_brilliant_ !
If someone doesn't already have Chef installed, why would they even have clicked on this link?
If somebody created a list of requirements for the perfect language for doing CM it would probably be similar to:
- Always available on every machine
- Proven reliable and stable
- Built to interface with the OS and other utilities
Hmm, that sounds a lot like the shell. If your application doesn't already depend on Ruby, bringing Ruby in as a dependency, is a lot of extra overhead. There is something to be said about this concept, but it should be shell all the way down.
The bash script itself is also extremely straight forward and is easily testable with a built in REPL (you know, bash). Anybody who vaguely knows unix is also going to be able to understand and maintain the shell script. Adding in new dependencies is simple. It's easily portable and you don't need to do any extra work to add in new features. I can't even think of a drawback.
Shell scripts are not easily portable either, unless by "portable" you mean "works in Linux". Node.js scripts, for example, really (mostly) work on Linux and other platforms like Windows.
Until you hit some nasty "path is longer than 254 characters" bugs. Oh well..
I can't really argue the point about complexity though.
There may be some overlap between things that must be done on both the FizzBuz and BazQux fleets, but that overlap is probably in simple tasks.
For something like this, you'd start up an instance, ssh in, set up the server using bash, grab your commands out of bash history into a script, and then deploy 100x instances and have them all run the script. Simple enough anybody who has used unix now understands your whole ops setup. It doesn't have the perfect rigor that other solutions have - but sometimes that high learning curve and perfect rigor means that you miss the forest for the trees.
Just because someone wrote something shiny and new that replaces 10 lines of shell with 20 lines of chef installation and another 10 lines of chef, doesn't mean you have to use it.
for box in box1 box2 box3; do
cat << 'EOF' | ssh $box
# do something
EOF
done
easy. Even better, just pass in some data into userdata with your autoscaling group. (That userdata field? just start it with a #! line just like a shell script... and it'll execute your shell, php, perl, python, etc script!)
The problem is so many programmers don't take the time (or care to take the time) to learn to use the well thought out design of Unix tools opting instead to see every problem as a nail corresponding to the latest trend in hammers (programming languages).
This has led a lot of programming types to create advanced tools for managing Unix systems which largely ignore the design of Unix.
http://discuss.joelonsoftware.com/default.asp?joel.3.219431....
Seems like he really did need some tools.
find -iname 'foo*' # [1]
... | sed -e 's/ab\+c//' # [2]
... | sed -i -e 's/abc//' # [3]
tar -xf some-archive.tar.gz # [4]
python -c 'anything' # [5]
Things like messing around with /proc are more obvious, but things like curl (is curl installed? what do we do if it isn't? try wget?) can be hard too.[1]: find doesn't assume CWD on all POSIX OSs [2]: "+" isn't POSIX. You have to \{1,\} that. [3]: -i requires an argument on some OSs. [4]: This is stretching the definition of portable a bit; I've worked on machines where you had to specify -z to tar, given a compressed archive. (tar has been able to figure out compression on extraction for well over a decade now, so -z is usually optional, but some places are really slow to upgrade.) [5]: Unless anything is a Python 2/3 polyglot, you'd better hope that you guess correctly that Python 2 was installed. (And it's really hard here: python is either python 2 or 3 on some systems, depending on age & configuration, with python2 and python3 pointing to that exact version, but on some machines, python2 doesn't exist even if Python is installed, despite PEP-394.)
Every script you write isn't going to be portable, but it's not that much of a stretch to endeavor to keep your script simple, not make assumptions, and be mindful of the potentially missing features of some implementations.
I take a special objection to [5], `python -V` isn't difficult at all to run, hoping and guessing are not necessary.
There's a good guide here: http://www.gnu.org/software/autoconf/manual/autoconf.html#Po...
YOU can control where the app is deployed (this is largely true even if you're selling your app just by having installation requirements or by selling appliances instead of installable apps).
I mostly meant that in a simple statement of:
python -c "code"
…you're probably forced to assume that it's Python 2 (or write a 2/3 code) and hope that your assumption is right. You can't run `python -V`: you're a script! The point is that it is automated, or we wouldn't be having this discussion.Of course, you can inspect the output of python -V (or just import sys and look at sys.version_info.major) and figure it out, but now you need to do that, which requires more code, more thought, testing…
As for 5; How many system does have python 2 installed, but no python2 binary/sym-link? (I've never had to consider this use-case for production).
Note a slight benefit of splitting tar to zcat and replacing python with python2, is that you'll get a nice "command not found" error. You could of course do a dance in the top of your script trying to check for dependencies with "command -v"[1]. If nothing else such a section will serve as documentation of dependencies.
Something like:
# NOT TESTED IN PRODUCTION ;-)
checkdeps() {
depsmissing=0
shift
for d in "${@}"
do
if ! command -v "${d}" > /dev/null
then
depsmissing=$(( depsmissing + 1 ))
if [ ${depsmissing} -gt 126 ]
then
depmissing=126 # error values > 126 may be special
fi
echo missing dependency: "${d}"
#debug outpt
#else
#echo "${d}" found
fi
done
return ${depsmissing}
}
deps="echo zcat foobarz python2"
checkdeps ${deps}
missing=${?}
if [ "${missing}" -gt 0 ]
then
echo "${missing} or more missing deps"
exit 1
else
echo "Deps ok."
fi
# And you could go nuts checking for alts, along the lines of
# pythons="python2 python python3"
# and at some point have a partial implemntation of half of
# autotools ;-)
[1] https://stackoverflow.com/questions/762631/find-out-if-a-com...Unix tools are the arguably the best tools available to a modern user. That, however does not mean that the Unix tools are well designed; many would argue that the Unix tools are extremely poorly designed or have no discernible design at all. S-expression are a much more powerful and useful abstraction than a "stream of bytes". POSIX was hacked on many years later in attempt to make sense out of the mess that shell commands had become. Shell scripts are very fragile and have never been truly portable across various *nixes, although the situation is better than it was twenty years ago, when it was enormously difficult to port scripts across the various commercial Unix installations, because they would break in many different and subtle ways.
I recommend reading the out of date but still useful Unix Haters Handbook: http://pdf.textfiles.com/books/ugh.pdf
Wrong, kind of. There are two variables in a VM: 1. Everything not in the VM and 2. The script itself.
1 is mostly dependent on what you're doing. If you're just calculating digits of pi, then yes, it's quite probably deterministic; if you're deploy software that's being pulled from github, running some initialization scripts, attaching some storage, then you're going to run into variables. All of those aforementioned actions have failed: github.com might be down (a rarity, but happened this week!), your scripts contain new code that's not quite up to par, and the cloud provider says the storage is attached to the VM, but it doesn't actually show up.
2 is that the script is probably in a VCS, and people are changing it. Someone is bound to write a line that doesn't work. (In fact this seems to happen quite often when tests are absent…)
> I can't even think of a drawback.
I can. The biggest one is that bash's arcane syntax is a deathtrap. It's a great shell, but for stuff that needs to work and work reliably, it's riddled with holes. Take the article's script:
sudo apt-get -y install build-essential zlib1g-dev libssl-dev libreadline6-dev libyaml-dev
cd /tmp
wget http://ftp.ruby-lang.org/pub/ruby/2.0/ruby-2.0.0-p247.tar.gz
tar -xzf ruby-2.0.0-p247.tar.gz
cd ruby-2.0.0-p247
./configure --prefix=/usr/local
make
sudo make install
rm -rf /tmp/ruby*
Several of these (apt-get, wget, tar, did you just install code downloaded over an insure channel onto a server?!, ./configure, make, make install) can easily fail; if they do, your fucking shell script will keep plowing along as if nothing happened. Depending on the next action, this can be meh, or WAT. Since it ends with "rm -rf ...", I think if it does blow up horribly, it'll return success. You can say "set -e" at the top to cause it to bail sooner, but `set -e` won't catch failures in all commands (false | true). Fucking shell scripts.Don't get me wrong: shell scripts are great, especially if you need it to work NOW. One off stuff especially. But if it's going to stick around awhile, having something that automatically looks at and raises exceptions/errors when stuff fails is great.
The thing I miss from a lot of these automation libraries is being able to annotate dependencies between commands. That wget and apt-get can run together. (The rest is must pretty much run in parallel.)
Libraries also allow people who really know how to make this stuff sing built the low level functionality in. That make could be make -j $(( $coeff * $number_of_cores )) ; make install could be similar. Maybe CFLAGS or CXXFLAGS could compile ruby with a bit more options for a slightly more optimized install. We might extract the tar in a directory where a rm -rf /tmp/ruby* won't inadvertently delete something (unlikely if you're on a new VM, but I find that's not always the case).
Shell scripts are a tool. They have a place. Nobody is saying get rid of them, nor is anyone saying get rid of them for deploys. I just want something a little more robust.
Bash is fine, but it's not a sophisticated high-level language. For doing sophisticated things, a simplistic tool is not enough. We need processes that are deterministic and that can act intelligently. The "bash is fine" crowd, in my experience, tends to be the same crowd that thinks that servers are special snowflakes that we must feed and care for. Those days are over.
I got frustrated enough with the Capistrano, Fabric, Puppi lot that I wrote a pure bash deployment tool with a fitting name htpps://github.com/gerhard/deliver.
Ansible on the other hand is something else though. There is some learning curve, agreed, but it's not as bad as awk or sed. And seriously, if you know your bash, you will know both awk and sed. I consider my shell scripting to be above average, and I've attempted a Docker orchestration pure bash tool https://github.com/cambridge-healthcare/dockerize, but Ansible just makes the same job easier. It's not everyone's cup of tea, but before you dismiss it, give it a real chance. I should know because I have initially dismissed it thinking that it's too complex, yet another Chef circus, yada yada, but trust me - it's worth it ; )
One of the most important lessons I learned: use operating system packages for production; don't compile from source.
Compiling from source:
- Wait minutes for each instance to come up, as each time requires a fresh compile
- Have to download from ruby-lang.org -> you have a dependency on this site being up, bad idea when you hit a load spike, need to scale, and ruby-lang.org is having a bad day
- Loss of dependency management. Now you can't tell whether your environment has the right packages installed, nor can other things express a dependency on your code being installed
- Very difficult to remove packages for upgrades/maintenance/security fixes
Packages:
- Host your own repo, removing the need for external dependencies (can do this with source as well, but it's much better when packages know to discover themselves from a repository)
- Much cleaner rollback -- almost impossible to trash a system with apt-get/yum, they'll always leave you in a good state in case the package fails to install, or other mayhem ensues
Does anyone else here see value in a configuration management system that centers around installing files and running scripts, much like FSS? I'm contributing to, and using in production, a configuration management system that takes this approach. It's a lot more mature than FSS, but unfortunately the primary author isn't ready to open source it yet. If there's some demand for it, maybe I can convince him to hurry up.
Bottom line, if you are just running a small startup with a few servers and a few people, then by all means use FSS, but eventually you will need real CM.
This system has been managing about 100 servers and several hundred desktops at a single site since 2009, so the approach does scale. It has some problems that declarative solutions like Puppet solve, but at the same time it solves problems that Puppet/Chef/others have.
Servers are now disposable.
* Upload/update a configuration file, and if it resulted in a change then restarting the affected service. * Append a line to a file, if it is missing. * Search/replace a pattern against a file and do something if that resulted in a change.
So yes, add the ability to install/remove a native package (be it a .deb, .rpm, or whatever) and you've got a lot of power.
I'm not sure you can genuinely have a good provisioning script that treats all of these things as variables in a matrix:
- OS (sometimes including windows as well as linuxes and bsds)
- OS version
- Using system packages or building from source
- Package/source version
- Every single config variable the program/service can have
Every time we use a community cookbook, which includes several hundred plus lines of code to deal with platforms we will never ever deploy to, it ends in tears eventually.
Having a 20 line recipe that works on the Ubuntu LTS we've standardized on is much closer to what FSS can achieve, but even then it does often make simple tasks considerably more difficult.
This fucking sucks.
That all being said, the thing missing here between any of these tools is most obviously the resource model -- and the templating system and where you put variables and things to manage variance between systems. Thus, it will blast out some commands for you, but that is the easiest part of the equation -- not too much different than say, doing something from Capistrano or Fabric.
Even in Ansible, that's the part we built first.
Achieving idempotence in shell scripts is the reason most people move away from shell scripts, and also ... well, the desire to program less :) Then you'll want the rolling update features, or provisioning, or dry run, or a way to pull inventory from cloud sources, or... and you'll streamroller a bit.
The balancing act for us is achieving the right level of features vs language complexity and keeping in that sweet spot.
I still think you should have an easy time getting started, see also things like http://galaxy.ansible.com for community roles to download to go faster -- most people should be able to do basic things in a few minutes. The script module in particular is a great way of pushing a f'n shell script :)
http://docs.ansible.com/script_module.html
(full disclosure: I created Ansible, but I did not shoot the deputy)
With Ansible, this is a single line.
I guess I could learn the "one-line" ansible way to do it, but I'd also have to learn to set up ansible. And I'm guessing it's not as flexible as, say, shell scripts.
As for not being as flexible as shell scripts, I'd say that's technically impossible, since it has a command for running shell scripts :) Personally, I never had to use it, but it's there.
2. I don't take your project seriously (perhaps unfortunately)
My surprise when I saw this is the exact opposite of that ...
The reason automation, CM, and devops have taken hold is because people are so goddamned tired of serving technology like it's some kind of god to which we owe worship. Screw that. I want the VMs to go make ME a sammich, not the other way around. Trying to build serious infrastructure using nothing but shell scripts is a quick trip to the temple of server-worship.
So far I only have experience with puppet, and it seems really annoying to tie phases like "install postgres" -> "Reinitialize the database in UTF-8, but known to work on ubuntu" -> "Update the configs so non-localhost connections can actually connect" -> "Okay, now it can start."
The above problem is something I do not want to deal with when it comes to a provisioning system.
I know that was only an example script, but we don't want to be encouraging bad practices here.
I think a huge problem with puppet/chef is that they try to do it all, ie deployment and OS state management and make it work with vendor specific package management systems.
Unfortunately rpm/debian based package mgmt systems are not well suited for complicated deployment strategies. Most companies I worked at come to a solution similar to:
* Install everything you need in /package_versionstr (except say glibc)
* Point the current version with a symlink
This enables you to do a simple atomic rollback/rollforward and is much easier to reason compared to complicated pkg state.For simple OS config management (say usermgmt, sysctls), we used a system similar to the ideas expressed by OP.
1. Keep it simple. puppet/chef have horrendous DSL's that make it really complicated to reason about. I shouldn't have to debug a backtrace 15 levels deep to understand why an useradd didn't work.
2. Server side logic. Don't try to do "intelligent" stuff based on client side state, it will be almost impossible to get it right. All data needed for state needs to derived by group membership. This helps in * validating all changes upfront * diffs for state changes across a group of nodes.
3. No orchestration. Except any changes to be applied at any time. This acts as an enforcing function to make your scripts idempotent.
The Ansible setup for my personal server started as two files: an inventory file consisting of an ip address, and site.yml. My ignorance (and some installation issues) notwithstanding, it has scaled pretty smoothly from there to copying up config files, templated nginx config, and so on. I don't see much room for anything between that and a literal "just f###### shell script".
1) Step through the program, try to think of logic errors etc.
2) Do a binary search through your latest commits and see where something broke.
Those are your two clues.
However, not having idempotent configuration files seems to be not only missing some of the power of shell scripting but also missing a huge principle of "good" server configuration.
Do the scripts reproduce with all that fscking going on?
'nuff said.
If you want people to use your software -- and you do or you wouldn't be promoting it -- you should consider how you present it. Bad words don't offend me, but they do suggest immaturity and inexperience... qualities I tend to avoid in systems orchestration.
> That's why we have perfesionalism, cause prfesionalism is good.