How to write idempotent Bash scripts
arslan.io
arslan.io
There are of some downsides to that approach which stem from the fact that you can’t possibly declare every possible thing. For instance, if you were to run the above declaration and then later remove the declaration, chef wouldn’t know to delete the directory; at that point chef wouldn’t know anything about the directory at all. So ideally you can build infrastructure from scratch each time (Docker, etc) rather than having to converge an existing machine.
If you are doing sysop^Wdevops work it is the single most important thing you can learn, once you have a good grasp of your shell and your editor.
The difference between a giant git archive of shell scripts that various people have modified over the years and state changes described in a configuation management language is the difference between fixing things in the small and reasoning about integrated systems. It's something that needs to be experienced to be appreciated.
Reasoning about system state as an integrated whole is just as relevant when you are shipping applications as containers, if not more so. It's not uncommon to start using something like Kubernetes without first being able to describe global state and the result is just as messy as before, if not worse. Something like Helm is impossible to understand unless you have complete control over your configuration.
I understand a shell script is not suitable for orchestrating a datacenter, but sometimes it is suitable for installing a couple of things.
I learned a lot about catching all the different bad things that can happen, adding prerequisite checks, rolling back changes automatically, etc. In the end, a tool like Puppet/Chef/Ansible would have taken a fraction of the time to develop scripts for.
Containers have definitely lessened the need for mutable configuration management, but they are still useful in the 90% of environments that are still working in a past decade.
Usually those cookbooks will want to have things in a particular way, which might not be the way you want.
Chef/Ansible/Puppet all have the problem of having many layers of overhead for the same thing.
This makes them slow, hard to debug, hard to read and even write (in my opinion).
Sure, bash is hell, but the go from bash to something like Rust, instead of Chef/Ansible.
Sure, again, this might sound rather unconventional, but you get safety, speed and convenience of a full fledged programming language, single binary output, etc. And you need to test cookbooks/playbooks too, so why spend enormous effort on cobbling together scripts and high-level abstractions instead of writing just what you need with the assumptions you really have - instead of what a general cookbook/playbook has to have (which is a vast difference).
Chef/Ansible/Salt could be described as rudimentary type systems for operating system states, that enable you to convert between different types/states.
It's conceivable you could build such a system in Rust, but Rust itself has no primitives that would make it easy. It certainly isn't a good starting point when you want to set up a database server.
I've used both Chef and Ansible. Their maintenance cost is pretty high unfortunately, and they are not that flexible, nor resilient to worth it.
No wonder the container/k8s boom is so big. Immutable infrastructure (docker images) with really powerful operational features is what can justify high maintenance efforts. Whereas fancy but brittle (due to the inherent problem of state discovery) CM systems are at best only useful for the initial platform setup.
Rust is just useful because you can quickly and relatively safely produce single binary programs. (Many people use go for this, but it's easier to skip error handling in go which result in runtime problems, which is very inconvenient in a provisioning/setup system.)
And the image build should be just a simple imperative install these packages, use this config, run this command on invocation.
Nix is rather amazing with its powerful CLI stuff it provides (S3 compatible dependency store, fetching via SSH, closures, etc).
My only problem with NixOS is that it's very much like Gentoo. It has infinite composability built-in, but it means you have to rebuild everything. In Debian/Ubuntu land you usually can simply enable/disable install/uninstall specific feature related packages. (For example postfix and postfix-mysql packages.)
Then they simply execute other programs to then parse their output. And/or fiddle with files (parse them, alter them, write them).
Sure fundamentally the syscalls and apt/dnf/yum will be the slow parts, but I found that development of CM scripts/plays/recipes are usually bottlenecked on the turnaround time of the CM system's own workflow. And execution time is a significant part. (The bootstrap, the transfer of whatever files, and so on.)
Rust would help with writing things that are relatively well error handled at compile time, and gives you single binaries.
And simple imperative script is a lot more readable than a custom DSL with who knows what ruby hooks.
Currently the CM system that makes sense on the long run is a git repo for Terraform. (Because everything else runs in containers anyway, and to set up the immutable images you don't need "state management". And where you need it you need active state management such as k8s and its specific operators.)
I wouldn't pick Rust for this but there are lots of examples of using Node/JS, or C# or Python. It's a much nicer environment and benefits from full IDE support.
For a good example, look at what Pulumi is doing for modern infrastructure: https://www.pulumi.com/
Chef and friends are blocked on disk I/O. How does a typesystem and/or thinner abstraction layer to disk I/O speed up the underlying expensive operation: blocking disk I/O?
It’s much easier to see the failure conditions in a Rust program rather than in bash. Also rust seems like easier to maintain, too.
How is a 100 lines of rust easier to maintain than 10 lines of shell script?
It's exactly what the other commnet said: "Writing a shell script in rust would be obnoxiously difficult for no real benefit."
>> How is a 100 lines of rust easier to maintain than 10 lines of shell script?
In 10 lines of _proper_ bash you won't be able to even check if the arguments supplied to the script even exist, let alone parse something more than simply subscripting argv[].
I don’t think there would be anything wrong with a rust implementation of infrastructure management, but I also don’t think it’s the silver bullet to solve what plagues the space.
But it simply is a solution in search of a problem in today's container orchestrator/platform world.
HAproxy is amazing, but a k8s Ingress service [0] has a nice API, so I don't have to run Chef to add a new vhost. Yet I can persist the config in a YAML. (Or json, or whatever.)
[0] which can of course be backed by HAproxy down below, but traefik has dynamic config; though haproxy2 will do that too
You could rewrite your bash scripts in rust but while it would probably be safer/more correct than bash, the whole paradigm of writing imperative scripts to manipulate existing machines is flawed, whichever language you write it in. And executor performance is almost never the bottleneck in my experience.
I agree that manipulating existing machines is folly, but that's why people use containers, and immutable images. And then you can prepare the images anyway you like, but since they are very deterministic (modulo the external repositories/packages/curl-bash-piped-scripts), you usually use something simple (eg bash), because there's no need for that state management.
And actual state management should be left for the platform (k8s, swarm, or some cluster manager).
I found the chef workflow slow, not just the executor itself. (Uploading cookbooks, fiddling with dependencies, testing them, debugging them, etc.)
Ansible was even worse in my experience. Much slower execution, harder to debug, opaque python blobs, extremely confusing and fragile "YAML programming" coupled with hostile variables files. (All the things that are tenfold more intuitive in Chef.)
Then you have scripts that you want run once and if run again not have any more effect. FOr example a script to be run by ops on production servers, what happens if they accidently run it twice! An easy way to cover that is to have the script create a file, which you check at the start of the script and if present you exit. For this you can use the date/time, PID and script name for a unique file name, create in a /tmp directory that you cleanup weekly via sculker or whatever frequency you require. Handy way to handle scripts that you want run successfully once and never to be run again.
Another approach is to create a unique directory (named by hash or date) with the results, and once it's good, symlink the real directory to it. (Full disclosure, I stole this approach from nextflow.io, a great workflow system aimed at bioinformatics).
ln -sfn "build-cache/2019-07-07-21:23:04" "build"You can use numerous tools to compare/manage branches; md5sum / shasum (md5 is usually useful, though not entirely safe), diff and kin, vimdiff, rsync, fdupes, jdupes, git, hg.
Traps are underused. They solve a lot of error-handling and cleanup problems really easily!
Here[1] is a simple example of using a trap to cleanup a tempfile, that safely handles pipeline errors and unexpected exit/return. The same basic idea works for many kinds of error handling, cleanup, or similar work.
If you ever need to cleanup something at the end of a function
foo() {
setup_command
# ...
cleanup_command
}
it using a trap might be as simple as moving the cleanup command to trap at the earliest possible time in the function the command can run: foo() {
setup_command
trap "cleanup_command" RETURN
# ...
}
[1] https://gist.github.com/pdkl95/61fc242e7961cc2584a787ed1760c...Of course, the better solution to those situations is to make the cleanup command idempotent ...
(that said, any order is much better than the unfortunately-common style of simply ignoring errors and exceptions.)
This does sound a good solution on paper, but "checking at the start" and "creating a file" are two different steps (i.e. not atomic) and will cause trouble eventually if your system has the tendency to run the script twice. A better solution is to use the `flock` command before creating the flag file. For example,
exec 999>/var/lock/my.lock
flock -n 999 || exit 1
if [ ! -f flag_file ]; then
echo "script not ran before, running"
touch flag_file
else
echo " already ran, exiting"
exit 1
fi
# do stuff hereThe first safely creates a temporary file. The second (with noclobber set or -n argument) will atomically rename a file, but not overwrite existing contents.
Linux uses symlinks to system library files in order to allow in-place atomic replavement without affecting running processes with open filehandles to earlier versions. MS Window's lack of this feature is what makes (or made, I'm very out of date) malware removal and system updates so painful.
Nope. If you rm a library, the processes using that library keep using the old library. Symlinks have nothing to do with it.
Trying to do this with a filesystem instead of a real database, and things get messy.
I'm into using transactions as explicit guards around chunks of known-non-reversible logic as a robust way to warn that manual intervention may be necessary, but using them to replace lockfiles seems to have the worst drawbacks of both systems: it's inherently at-least-once, plus it's substantially more risk surface than a lockfile to establish and maintain a database connection/transaction.
Other concerns:
What if the database has gone away or crashed by the end of your possibly-long-running script (unless you're using sqlite, in which case, why not use a lockfile? On second thought, sqlite is probably better at lockfiles than flock...)?
What about the tech choices that using a transaction precludes? You can't write ops scripts in pure shell any more (without some seriously weird file descriptor munging and job backgrounding to keep the transaction open while you do stuff). Installing your runtime of choice plus a database driver may be unnecessary and time/space-consuming on a lot of systems you need to manage this way.
You now also have a bootstrap problem: your database driver and associated tooling may not already be available on your target system, so using such transaction-managed scripts to provision clean/empty systems is another area where additional challenges emerge.
All of those can be mitigated or worked around, and this shouldn't be taken as advice to never use an RDBMS's transaction system to manage ops tasks, but I think the cases where it adds more than it costs are pretty rare.
Bash scripts can't be idempotent because they operate in an external environment that can't ever be. The better option is to just be extra safe.
See:
> http://redsymbol.net/articles/unofficial-bash-strict-mode/
The article is about writing bash scripts that don't care if you've ran them once, or 10 times.
set -eu set -o pipefail
Anything else I'm missing?
Is the idea that participants don't know what "idempotent" means in a thread with that word in the title?
Learn to handle errors, don't learn to brute force your way through them.
There is a good reason that many of these flags are not default behavior, and it’s that they can be quite destructive.
If you’re writing a script, as the beginning scenario says, that has errors in it, and you expect to have more, do you really want to tell all the commands to power through and delete and overwrite stuff if it encounters an unexpected state? Telling people to just always use these flags and not only when they’re quite sure what their script is doing, is probably playing with fire.
This may make the script idempotent in repeated runs, but the first run might not be what you expect at all.
Where are they saying this? These flags are about making sure the script wouldn't error out right away if you run it again, eg how a regular mkdir for a path that already exists would not exit 0 and thus end script execution (given set -e is active, which the article seems to imply). Also it is not so much about purposely writing a script that contains errors, because that would just mean the error is encountered every time, but more about transient errors like running out of disk space, a curl call failing because of network hiccups etc. Everything your script did up until that failure shouldn't prevent a second call from succeeding.
Every call to touch will alter the modification time of the file (example.txt in this case).
It is very difficult to write truly idempotent bash scripts (consider log files). Creating sort of idempotent or at least rerunnable scripts would be nice but I think even that would take more care than this.
I occasionally do full-text searches on system files, and this resets access times on all files. Thus, I cannot imagine using a script which cares about it.
The advice such as replace rm <FileName> with rm -f <FileName> could lead to Disaster depending on the scenario. So a HUGE YMMV
Removing something and adding it back seems to be, by definition, not idempotent. What if something tries to access that file in between? What about the filesystem timestamps (same issues as with the "touch" claim)?
How do you guarantee the replacement symlink had the same chmod flags and grp flags as the original? (chmod -h for reference*)
I don't think these commands are idempotent in the wide. They satisfy only the narrowest "yea, most normal-ish cases are sort of maybe ok" but since we lost atomicity, there are significant windows of time in the filesystem where inodes are in flux, and the new thing has a different inode to the old thing, if you delete and replace it. The disk is different.
ln -sfn target new && mv -Tf new old
It would be nice to have a toolset of commands that are all idempotent to make that type of task easier.
And if one wanted to limit my shell scripts significantly, I'd limit myself to commands supported by busybox's "ash" shell -- this is a limited shell environment which is pretty widely used in initrd's and embedded devices.
echo -n > example.txt
Both fit the description of "Creating an empty file". Is it idempotent, or more or less so than `touch example.txt`? If there is already an existing file, then touching it obviously will not empty it, so the end state is not "empty file exists" like one would expect from operation "Creating an empty file"Ironically article calls "This is an easy one"
> example.txt
without prepending the echo command. This will create an empty file or it will truncate an existent file. The touch command could be preferable to preserve the content, if needed.
>> example.txt
this one (with a double > ) could replace touch in this application case: it's less readable but more efficient.
The workaround was `printf "%s" foo`.
Compare:
{ a; b && c; } ≫ log.txt
With the same written in Python: from subprocess import check_call, CalledProcessError
f = open("log.txt", "a+")
try:
check_call(["a"], stdout=f)
except CalledProcessError as e:
pass
try:
check_call(["b"], stdout=f)
check_call(["c"], stdout=f)
except CalledProcessError as e:
exit(e.returncode)
exit(0)
(Note: I took these examples from an article I wrote a while back about programming languages https://innolitics.com/articles/programming-languages/ )Of course, others would prefer, say, Clojure, or Kotlin (although the JVM startup time is the worst of all), but therein lies the issue... We're stuck with Bash scripting because everyone (more or less) knows/understands it to some extent, it's installed everywhere, and it has essentially no startup-time cost.
I was surprised to find that Ubuntu’s handler for unknown commands that tries to suggest what you could apt install is written in Python. I’ve never noticed it being slow.
Python might not exist on your windows machine.
Etc.
For example, there is https://docs.bazel.build/versions/master/skylark/language.ht... which is a subset of Python.
Bash might not exist either.
It can be learned, and one can even learn to love it. But scripting is often done by people that very seldom write scripts, the barrier to make decent scripts is much too high and readability isn't much better. The nonsensical syntax is also easy to forget, the result is just an awful lot of headache and poor scripts floating around. Surely we can do better?
To some degree, this has been also solved with containers. It's much easier to create new container image, then create proper idempotent script. And you can create create image from running system too. So you can just "bash" commands and don't care about the state of the system. When done, you just create final image and you can be sure, that state of the system in the container would same as you intended.
I'm talking mostly about install scripts. Of course there are valid use cases for idempotent bash scripts.
Terminology is broken. Correct is "ln -s TARGET LINK_NAME" per its man page.
Yes, ultimately you have to test your scripts on the actual systems, but that is something you have to do anyway. For example, when you run scripts on MacOS and you run into old Bash bugs because Apple refuses to ship an up-to-date version, those are issues a standard can't solve.
However, I have no experience how much you can count on POSIX when it comes to C APIs and the like.
> but that is something you have to do anyway.
Exactly my point.
Works with GNU mktemp, the old BSD variant on macOS and also Busybox. Current Android uses Toybox, I think? Its mktemp implementation looks like it takes the same flags as Busybox mktemp.
Instead, I could use --tmpdir, but somehow that one seems to be buggy:
$ mktemp -d --tmpdir name.XXXXXX
mktemp: Failed to create directory name.XXXXXX/name.XXXXXX/tmp.sL2WbI: No such file or directory
The macOS version, on the other hand, does work with those options, but it creates a file like name.XXXXXX.veCNnwkX (instead of name.veCNnwkX) so not a deal-breaker, but it is certainly not what you would have expected. And using --tmpdir with macOS doesn't work either.So yes, in theory, it shouldn't be too hard but sadly, the reality is often buggy and outdated :-/
I.e. this is about `mktemp(1)`, not `mktemp(3)`.
mountpoint -q $MOUNTPOINT || mount ...
Instead of
if ! mountpoint -q $MOUNTPOINT; then mount ... fi
It's much more compact in the case of having a lot of commands that need to be checked in case they don't need to be run again.
mountpoint -q /proc || {
mount -t proc none /proc
chown 0400 /proc/slabinfo
}
if you need multiple statements.https://unix.stackexchange.com/questions/228597/how-to-copy-...
You can delete an email or mark it as read even if it's been done before (on another open screen as an example).
Interestingly in the 'old days' of elevator operators hitting it multiple times would be either a 'hurry up' for the operator or an annoyance that maybe made it come slower!
Unless you have processes that depend on modification time.
> ln -sfn
If you don’t mind the possibility of ownership changing, this is OK.
if ! mkdir .lock; then
printf >&2 "Already running?\\n"
exit 1
fi
Some network file system implementation do not guarantee atomic mkdir, so you still need an extra caution with this method.A better system needs to be able to track directories creation, package installation, filesystems creations and mounting, service restarts and so on.
something_else && touch something_else_doneMake is much less well known than shell script. In fact, I'd go so far as to say that make is infamously poorly understood and makefiles are infamously poorly written. That makes it less maintainable, or otherwise you'll need to put "experienced with GNU make" on your job description.
If you've got the script scheduled, then you're going to have to document it everywhere that it's not actually a build script so please don't disable it because it doesn't obviously look like it needs to be running every 6 hours.
And because it's a DSL and you're somewhat going out of scope, you're more likely to have a problem like needed to extend your script to do something you can't do as easily in make. Worse, you might want to do something that you're expressly not supposed to do [0].
You're going to have to defend your decision every time you present the script, too. And if you ever need to pass it off, the first thing you're going to have to do is explain why you used make. Whether or not the person you hand it of too is an idiot or not, you know the first question is going to be "why isn't this a shell script or Python script?" You're going to have to defend your decision to use make instead of shell script because of idempotence, even though you can write idempotent bash scripts. And if the person you're talking to is an idiot -- and let's be fair that there's a good chance that that is the case -- then they'll never figure it out.
So, for me, you've got to go beyond "this language can technically do the task" in order for me to understand why you'd want to use something that's generally understood to be used for build scripts for, well, anything other than that. By choosing make, you're doing something unexpected. That's a bad idea.
[0]: https://www.gnu.org/prep/standards/html_node/Utilities-in-Ma...
Irrespective you have a good point that touch does have side effects.
My understanding is that "hermetic builds" refer to the kind of thing that tools like Guix, Nix, etc. go for. Something like "Declare all the inputs and dependencies to get a reproducible output".
Basically, if your build process requires you to pull code from any repository that you do not own, it isn’t hermetic. If you have to `pip install` or `go get` a third party (or even in-house!) dependency from a source that you do not control, your build is not hermetic.
Effectively, this means that you have to have versioned copies of all of your third party dependencies, and version-specified build graphs. Very hard to do without a mono repo and a build tool.
In the context of the GP, I’d say that deterministic would have been a better word choice. A reliance on time stamps technically wouldn’t make a build process non-hermetic, but it would definitely make it non deterministic. It’s technically possible to have hermetic builds without having reproducible builds, although that would be a very bizarre org. :)
I work on things close to Bazel, and the word "hermetic" gets thrown around a lot. And because of that, hermetic in my mind gets translated to "how an ideal build of a project should behave" (which obviously is wrong).
Hermetic implies that unexpectedly different input artifacts are not possible, and builds are deterministic across machines/environments.
A deterministic process always produces the same output given the same set of inputs.
A hermetic behavior ensures that indeed you'll always have the same inputs. It can just be a set of best practices (e.g. being very careful of not depending on external inputs that might change outside of your control) or it can involve an active barrier that sandboxes your environment in order to ensure that you indeed always have the same inputs.
A reproducible process is a process that can be repeated later in the future. There are various degrees of reproducibility you might be interested in. For example, you might want "bit for bit" reproducibility (important for security) or you just want to make sure you can rebuild something functionally equivalent (e.g. the compilation or link phase might not be fully deterministic in the order and layout of compilation units).
Reproducible processes usually rely on a deterministic system and leverage hermetic behaviours to ensure reproducibility (over time or across locations)
> touch -- change file access and modification times
If one were to argue side effects of touch, it would be that non-existent files are created. The purpose of touch is to update access and modification times.
https://en.m.wikipedia.org/wiki/Side_effect_(computer_scienc...
Idempotency requires that the function map from a space back onto itself. Touch is only roughly idempotent in that you can define a statistic S on the overall system state discarding filesystem stats such that S(f(f(x))) = S(f(x)).
Sorry, what does idempotency have to do with convolution?
There's a wider idea of idempotency in computing - that state changes in general respond to operations. Repeating the same function more than once (because the network stuttered or the sender resent the message) and arriving at the same state is a common example.
With the default behaviour of rm, ln -s, etc you know neither the state before, nor after.
For the case of a setup script you'd need to update your state file after every step and in turn when running it again see what the file says and resume at whatever point it says. But then you risk running into above problem.
I try to write idempotent scripts whenever possible, combined with general sanity checks specific to whatever environment the script expects.
We know what the state is right after the command for an undetermined amount of time, but we have no knowledge of what’s changing it. Why not cover that? There’s no reason not to other than it hasn’t been worked on enough. Consider a database schema migration, there’s a reason those are not idempotent, you know each state. Seems quite solvable.