When Bash scripts bite
blogs.janestreet.com
blogs.janestreet.com
However, don't let it give you a false sense of security.
Shellcheck probably won't say anything about the "biting" parts from the article – because they are valid, just behaving a bit different than the user expects…
That goes for any linter or static analyser i guess.
E.g. "if (x = 1)" is valid C but i'd be pretty unimpressed if a linter didn't flag it up!
Perhaps it's time for a new shell language that is less arcane, with fewer gotchas and simply less bug potential?
Unless you mean something entirely different than POSIX shell family, e.g. Powershell, Python, Ruby, or JS.
set -euo pipefail
yes | head
That will consistently exit because `yes` gets sigpipe and quits. Which is expected, but triggers a script exit. But more exciting is that something like: generate_data | head
only _sometimes_ fail. It's a race that depends on whether generate_data is able to stuff all of its data into the pipe buffer before head calls close().EDIT: I seemed to remember sharing this bug not too long ago, and indeed I did. pixelbeat responded with some interesting links: https://news.ycombinator.com/item?id=13940628
As another poster mentioned, set -e and set -o pipefail are crutches and tell me the script it sloppy.
When the author wrote "a particular production bash script (if that doesn't sound horrifying, hopefully it will by the end of this post)," I couldn't help but smile...
I interviewed there, but unfortunately didn't make the cut.
Do you happen to have some recommendations for similar blogs?
> set foo [exec /bin/true][exec /bin/false][exec /bin/true]
child process exited abnormally
while executing
"exec /bin/false"
> exec true | false | true
child process exited abnormally
while executing
"exec true | false | true"
> set sp "hello world"; exec echo $sp; # No need to quote $sp there.
hello world
Not all shell-like things are as convenient to do in Tcl as they are in sh (the most significant difference to me is that you cannot pipe to or from functions), and it is more verbose, but because everything is a string in Tcl I find that it integrates with *nix (or Windows!) command line programs better than other scripting languages. E.g., > lmap x [split [exec ps | tail -n +2] \n] {lindex $x 0}
4540 5767 5768 31161
What happens here is that you take the output of `ps | tail -n +2`, split it on newlines, then map over it treating each line (that is something like "5814 pts/0 00:00:00 ps") as a list and taking the first element in it. The result is a list of PIDs (a string containing the PIDs separated by whitespace).I can recommend trying Tcl to anyone fighting Bash who doesn't want to replace it with Python/Ruby/etc. If you try it, though, use version 8.6 or at the very least 8.5. The previous versions are EOL but are still common in the wild. If a recent Tcl is not available on your system, you can build a self-contained static binary interpreter with http://kitcreator.rkeene.org/kitcreator.
What I consider to be the biggest benefit of Tcl in general compared to the more mainstream scripting languages is the pervasive immutability. In Tcl values are strings and strings are immutable. It encourages you to write programs that consists of a reusable, easy-to-test collection of pure functions and a task-specific procedural or (more rarely) OO core.
You, like many others, underestimate Tcl. I present to you, command pipes for Tcl:
Tcl can be any language you want, because it's already the language you need.
Not in this case, at least :-), since on that wiki page there is one implementation of command pipes that I wrote and a link to another (in fptools).
What I mean when I say that "you cannot pipe to or from functions" is that while you can pipe data from one external process to another in an [exec] command, you can't pipe it from an external process to a proc or vice versa, which would be analogous to what the POSIX shell can do. Functions and external processes are less similar in Tcl than they are in sh.
The best you can do with a simple command pipe implementation like those on the wiki page is pipe data from [exec foo] to [bar $arg] to [exec baz << $arg], which, unlike [exec foo | baz], will read foo's entire output before passing it along. The only real solution that I know to allow you to treat snippets of Tcl code and external processes as basically interchangeable in a gradually read pipeline is rkeene's pipethread library. I am a big fan of it, but using an external dependency that isn't in your operating system's package repositories adds considerable friction to writing and distributing a simple script. (I do hope pipethread ends up in Tcllib.)
if res="$(ldap-query-for-valid-users)"; then
echo "($res)" > "/tmp/all-users.sexp"
else
handle_failure
fi
I second the recommendation to use https://github.com/koalaman/shellcheck – you really shouldn't be writing shell scripts without it – but in this case it doesn't seem to handle the issue (with default settings at least). result=$(ldap-query-for-valid-users)
echo "($result)" > "/tmp/all-users.sexp"
Decoupling the process substitution from another invocation allows the error to be detected.But still there's basically no way to make this consistently useful.
Even if you religiously set -e/E in every scope just in case, if you're anywhere in a scope inside the non-final operand of a bunch of &&/||s, or as the conditional expression in a control structure or whatever, -e/E will just do nothing, you can't turn it any more on, you just don't get early termination on errors no matter how many nested function calls you're actually removed from the original ||. It's not great.
Related HN thread from almost 5 years ago: https://news.ycombinator.com/item?id=4530897
EDIT: just learned that PBS is now sh.py, so I removed it from the comment.
http://www.haskellforall.com/2015/01/use-haskell-for-shell-s...
ocaml isn't that far off, just even less popular...
Regardless, I'd actually prefer Haskell's do-notation for both asynchronous commands and shell-scripting tasks. Even though Haskell isn't my favorite language, do-notation is quite nice for this.
But we don't do it often. Each language has its strengths and weaknesses, and the constraints of a small script tend not to change that much over time. The script either starts other executables and pipes some text around, or it does some computation. The choice of which to use is often (but not always) clear.
That said, there are still little things we use Bash for. But our tolerance for large bash scripts has diminished greatly over the years.
After the meeting I said to colleagues, 'I quite like bash scripts, actually', and they all said 'I thought that too...'
So it's really easy to go from
$ command1
$ command2
to:
#!/bin/bash
command1
command2
As a result, everybody "likes" bash scripts since they're so easy to write, initially.
And these shell scripts will work well in 80% of situations and you can get them to 90-95% just by sprinkling a few ifs around.
Once thes script becomes longer and more complex and you want to make them easier to read and DRYer, that's when the pain begins. Or when you need to handle a bit of logic which would be trivial to do with better data structures such as arrays or dictionaries or sets...
You can work around these issues, of course, but almost every workaround is either ugly or a hack (as someone was joking, "elegant hacks" for older Unix hackers, "gross hacks" for younger ones :) ).
POSIX sh doesn't have these, but BASH actually does have arrays, and even dicts (see https://stackoverflow.com/questions/688849/associative-array...)
Or you could just learn Perl...
Or if you're just stringing commands together, traditional simpler Posix sh (note: 'bash --posix' isn't).
set -e
export x=$(false)
echo ok
prints ok, but set -e
export x
x=$(false)
echo ok
exits early because of the `false`. Line 4:
export x=$(false)
^-- SC2155: Declare and assign separately to avoid masking return values. echo ... > "/tmp/all-users.sexp"
No, this is not a secure way to create temporary files.Please use mktemp(1).
I know of no good reason to not use mktemp.
eg:
set -euo pipefail
foo() {
false
echo "hello world"
}
variable=$(foo) # failsfoo | do_stuff # fails
My preference is to handle things as streams
echo ($(ldap-query-for-valid-users)) > /tmp/all-users.sexp
should be something like x=$(ldap-query-for-valid-users);
test ${#x} -gt 0||exec echo no valid users >&2;
echo \("$x"\) > /tmp/all-users.sexp;
This way they would get the message "no valid users" to stderr and the script would exit. According to the blog post that is what they wanted.Alternatively,
x=$(ldap-query-for-valid-users);
test ${#x} -gt 0||exit 100
echo \("$x"\) > /tmp/all-users.sexp;
if they prefer a nonzero exit code to a message to stderr.So most of the time you should do "set -e; for XXX". Otherwise your Makefile loops will "succeed" incorrectly.
Really... Check your damn return values. set -e is a crutch of a sloppy programmer.
Ok, yes; you can do this really slick thing in one line by stringing together a bunch of commands. However, just because you can does not mean you should.
Bash makes it simply with 'if ! <command>; then <failure commands> fi'. Try not to string ten things together. Keep conditional true state in your scripts execution flow.
Part of it is the problem domain. Rewriting any slightly complex bash script in e.g. Python with robust error handling can be quite challenging. You have to make a number of very difficult decisions.
Using the example from the post,
echo "($(ldap-query-for-valid-users))" > "/tmp/all-users.sexp"
you have at least a couple alternatives to an additional temporary file. if ! valid__users=$(ldap-query-for-valid-users)
then
... failure case ...
fi
echo "($valid__users)" >/tmp/all-users.sexp
or {
echo -n '('
if ! ldap-query-for-valid-users
then
... failure case ...
fi
echo ')'
} >/tmp/all-users.sexp- How readable will your script be?
- How likely are you to miss a few checks?
- Will you be careful enough to also check commands like "echo" and "rm"?
"set -e" makes error handling implicit instead of explicit, and it solves these problems.
This is a bad argument -- our tools shouldn't make the "easy path" a path that's littered with subtle bugs! It should be easy to do the right thing. See: PHP hand-coded apps that led to SQL injection often versus Python's DBAPI which makes it harder to make mistakes, Signal versus some PGP GUI, etc. People make mistakes when their tools make it easy for them to do the wrong thing.
I have the same complaint against much of the standard C library too, but for a language that the user interacts with on a daily basis, this is unacceptable.
[0]: http://xon.sh/
[1]: https://william-droz.com/xonsh-a-modern-shell-that-enable-py...
>>> out = $(echo @(x + ' ' + y))
>>> out
'xonsh party\n'
>>> @("ech" + "o") "hey"
hey
In raw python, you have to play with subprocess by yourself.It's usually a bad idea to begin writing the result before knowing that all input is there. That's like a server announcing a 200 OK and beginning to stream, only later detecting I/O error. It's difficult to deal with such a server as a client.
More generally, in pipelines we deal with pairs of programs that are only connected by a text stream, with no possibility to communicate out-of-band conditions. In `PRODUCER | CONSUMER`, PRODUCER can't tell from a sigpipe whether CONSUMER crashed or has read all the data it needs. And CONSUMER can't know whether PRODUCER crashed or if it should take action on the results.
Most scripts don't really need the take-action-immediately level of concurrency. It's sometimes nice to be concurrent (use multiple CPUs at once). But the cases where the processed data can't be buffered at least in a temporary file before taking action are really rare.
Not that I'm not ok with short bash scripts or for systems were you only have bash... but if your Unix compatible system only has bash there is something else wrong with your system in my point of view.
I have recently rewritten a lot of Python and Perl (over 10kLOC) to a small collectiom of 10-ish bash scripts that weigh in at about 700LOC total. There are some tasks where Bash is the only appropriate tool.
Much simpler code, much easier to read, and a lot faster too.
Other than that, the simpler design also allowed greater parallelism.
I'd also go as far as saying that if you're using bash because it's faster than python, you're also using the wrong tool for the job.
The code is much, much easier to read now. You can see a script on a single screen height, which works wonders for overview. The 3x performance gain spawned from enlightenment this overview granted us.
But don't get me wrong, in the very same rewrite, I rewrote 100 lines of bash + 300 lines of Python into a single 200 lines of Python, which when combined with PyPy, was 50 times faster than the original setup. Python doesn't do bash's job very well, and bash doesn't do Python's job very well either.
Once you go down that road, you're into magic switches at the top of your script (which TFA demonstrates aren't well understood), and subtle control issues where bash's flow control doesn't quite do what you think it should, and (for me) this is where bash stops being superior.
But bash is definitely not something where complicated logic belongs.
But doing stuff like executing commands with standard output or standard error redirection, with pipes, etc. is something it's not strong at.
That's what a shell script is strong at.
To clarify, I'd: use bash to create an execution pipeline, which itself may contain a Perl script which munges some output.
Bash in this case is more of a "pipeline and execution orchestrator", and Perl is used to "properly munge the output" to make stuff happen.
use IPC::Run qw(run);
run(["cat", "/etc/passwd"], "|",
["egrep", "-v", ':/bin/bash$'], "|",
["wc", "-l"],
undef, \$out);
print $out;
Not the best way to do that sort of thing (useless use of cat for instance), I was just trying to come up with an executable example. run returns success or failure of the pipeline as the return code, which I don't show here. I don't think you can pick out which command failed.I also find that once you're in a real programming language, you have much less need to have deep pipelines anyhow. Consider the pure-perl equivalent
open PW, "<", "/etc/passwd" or die;
while (<PW>) {
next if $_ =~ m|:/bin/bash$|;
$count++;
}
print "$count\n";
Of course "real perl" would suggest some form of "use strict" and the corresponding need to declare the $count var, but if we're discussing shell script contexts it is at least fair to consider ignoring that, since it would be a silly objection that the perl equivalent doesn't turn on checks that shell scripting doesn't even have.But if you're in a real programming language you don't need to shell out to wc just to get a line count, or write complicated awk scripts, etc. It's much more common for me in those cases to execute a single program and control its STDIN/STDOUT/STDERR and watch its exit code.
However, most bash scripts I interact with will take at least 2-4 times the amount of code to replicate as a proper application. The line count example is a good pro-shell one, with 6 not particularly nice lines of code replacing a single, simple line:
egrep -v ':/bin/bash$' /etc/passwd | wc -l
"Complicated" AWK scripts are also much more compact than their equivalent. Say, `xyz | awk '/thing/ { getline; print $0; }'` to get the line after a pattern. I wouldn't write an AWK script much more complicated than that, though, and for me, Perl and AWK belong in the same bucket.Your example isn't even equivalent to my rather sloppy Perl, since, for instance, it will do something quite different if /etc/passwd doesn't exist. Shell only has the size advantage as long as everything goes perfectly correctly.
Shell is only good for two cases: 1. You don't much care about the integrity of your input or output, and you don't much care about what happens if something goes wrong. 2. When you don't care about the aforementioned issues because you're right there on the spot and can fix things if they do go wrong, because you're using shell interactively.
Having fiddled around with designing a "new shell" every so often my current conclusion is that there is no way to bridge the gap between the optimal interactive use shell and the optimal shell script with one language, and modern shell scripting, for all the features they have apparently for shell scripts, are firmly for the interactive case.
Bash is very good for the things its good at, but you still need to know what you're doing to not fuck it up. Some people don't know how to consider error handling when they code, but that's the fault of the programmer.
Shell goes even beyond that, though; as evidenced by articles like this you can't even get good agreement on how you should write safe shell. When even experts can't agree on what's safe, that's just something unsuitable to any serious task.
You are presenting a false analogy. I was not, am not and will never argue that one needs to be an oracle that foresees all errors (which would be absurd), but I am arguing that poor error handling is a programmer error due to the situation not being Bash specific. To reword my statement from the previous comment: If your error handling in Bash is insufficient, your error handling would also be insufficient in most other languages with implicit error handling (e.g., those with unchecked exceptions).
Once again, your Perl blob is a great example. You manually added the optional `or die` to deal with the case where the file wasn't present - perl's equivalent of `|| exit 1`. If you had not remembered to add that check, then your program would fail silently, with $count being an undefined variable. You could try to enforce better checking in the local block scope with some hacky header at the beginning of the script ("use strict;"), but that wouldn't stop you from having poor code in a module, and it would only ensure that count was defined, not that the program was well-behaved. That sounds awfully sloppy, doesn't it? Almost like what you were complaining about for Bash!
Very few languages provide you with any assistance to remind your of error checking (failing either silently or by crashing), mostly due to the awful concept of exceptions (unchecked, specifically, but checked exceptions are only useful if unchecked don't exist). Go, Rust, and - ironically - C helps you by forcing you to think about error handling at every single call-site, with lazy error handling sticking out like a sore thumb (explicitly ignored return values, which look weird in these languages), while C++, Java, Perl, Python and Bash all expect you to decide what you feel like handling today.
Language-wise, Python is far superior and has a large ecosystem and these days it has ubiquity too.
I'd be interested in the argument for Perl over shell scripts that does not apply to just writing the program in a modern language.
$ perl -pi -e 's/MY_NAME/Marco/g' ./*.txt
... would go through all "*.txt" files in the directory, and globally replace the token "MY_NAME" with "Marco".Great for a {poor man's,simple and easy to understand} basic template system to be used inside a bash script ;)
$ sed -i 's,MY_NAME,Marco,g' *.txtUse a C local (e.g. LANG=C or LC_ALL=C) for dramatic speed up of tools like grep/sort and probably also sed. If you don't do that, usually the UTF-8 parser kicks in, which is a lot slower.
Perl is actually well known to have a non-optimal regex implementation: https://swtch.com/~rsc/regexp/regexp1.html That said, I don't know how fast it is (what optimizations it has) at simple string substitutions.
I would love to see Python utilized, and maybe with VSCode and MS integrating open source into the .NET world it will eventually come to enterprise (where I work).
And it's a dream in VSCode + the most popular python extension https://github.com/DonJayamanne/pythonVSCode
Also, every system may have some Python installed, bud by the time I've dealt with versions (2.6 is the CentOS 6 default, 3.6 is the current version) and dependencies (virtualenv? system wide pip? PYTHONPATH? RPM Python modules?) I'd better get a significant improvement.
1) command substitution. you can catch this with an explicit error handler inherited by child processes. 2) pipefail SIGPIPE false positives. pretty hairy, command dependent whether this is a "real" error. often not, so you can work around it by ignoring SIGPIPE 3) process substitution. As far as I know, there is no way to workaround this whilst still using the convenience syntax. You have to carefully use explicit named pipes and carefully use wait on the PIDs (carefully!). Maybe still with races...
In my experience, you can write moderately robust shell scripts if you care enough and use all these flags and linters. But by the time you are at this stage you probably shouldn't be using shell scripting. More like training to spot problems in other people's code.
set -e
foo() {
/bin/false
echo "foo"
}
echo "$(set -e; foo)"The inheritance of these flags by subshells, which may have been written assuming different semantics, would be potentially even more problematic, so I think bash is right in this case, though arguably the inheritance could be limited to subshell code defined in the same file.
This is the time for a full scale programming language such as Python or perhaps Groovy on the JVM or Go language. When you need to write robust code, use the tools that were created for writing robust code.
Bash just has too many quirks.
Note that this is related to the most common way that people build a Big Ball of Mud. You have a simple app and you need a couple of features so you add them on. Rinse and repeat. Before too long you have an app that does too much and was never designed/architected to do that much stuff. You are probably ignoring a number of techniques for integrating functionality in large apps such as message queueing, microservices, separate libraries or packages, multiple languages.
Shell scripts suffer the same trajectory towards too much complexity. When you see it happening, and before the task gets too complex, replace the script with an app and apply all the normal software engineering techniques to make it robust.
-c Read commands from the command_string operand instead of from the standard input. Special parameter 0 will be set from the command_name operand and the positional parameters ($1, $2, etc.) set from the remaining argument operands.
Everything bash can do, ansible can do it better.
Or seconds, after you learn the search term "Perl one-liner" and add it to your queries for all your edge cases.
Start thinking of perl -we as a Bash built-in function.
Suddenly edge cases go from hours to your typing speed into Google.
Don't forget to leave a comment with a link to where you found the solution, what it does according to that site, and an apology for the unreadable line noise. If you make changes to what you copy and paste, explain your changes in a comment.
(Just in case it sounds sarcastic or ironical, this comment is serious and how I really work. I am really glad for my comments whenever I revisit old scripts.)
Actually, quite the opposite. All that shell scripts do is glue together other programs they depend on.
That makes them not portable. You can't just move a bash script from Linux to macOS because it uses ancient pre-GPLv3 versions of bash and all the other GNU utils.
For scripts where you need something unstandardized, say, imagemagick, you would wind up with up with a dependency in most other languages as well.
Also, other languages suck at doing small things quickly and simply. Try rewriting a shell script in go with little experience and see how long it takes.
I recently built a fairly complex system entirely in Bash, not because it was the "best" language for the job (it would have been much easier to do in Ruby or Python), but because the client didn't have any permanent staff that could maintain Ruby or Python scripts.
As a contractor, one of the major factors in the technical decisions I make is: how supportable is this technology for the client, once I have left? Of course the answer to that question will vary on a case by case basis, but Bash is usually a good lowest common denominator in a Linux environment.