Bash Debugging
wizardzines.com
wizardzines.com
Hitting ctrl-t on our main menu will, when booting with debug logging enabled, show a screen like this: https://i.imgur.com/Ge75zkP.png
We also have a flamegraph profiling mechanism that can be enabled with https://github.com/zbm-dev/zfsbootmenu/blob/master/zfsbootme... . That will dump data to a serial port, which when re-assembled, can be used to produce a graph like https://raw.githubusercontent.com/zbm-dev/zfsbootmenu/master...
Bash is suprisingly flexible.
PS4='+ ${BASH_SOURCE:-}:${FUNCNAME[0]:-}:L${LINENO:-}: '
When using `set -x` this makes it so that it shows the filename, function name, and line number. Which in larger Bash scripts can be quite handy in debugging.Also I recommend: rewriting scripts in another language. At work we are converting bash scripts to rust and while it’s a high ramp-up time, the resulting code is much easier to maintain and I have a much higher level of confidence in them. Bash is still good for quick scripts but once you hit 100 lines or so you really deserve a language with stronger guarantees.
# die: print error message to stderr, then exit with error code.
# example: die 69 "Service unavailable."
die() {
n="$1" ; shift ; >&2 printf %s\\n "$*" ; exit "$n"
}
Many more shell script exit codes and helper functions:https://github.com/SixArm/unix-shell-script-kit/blob/main/un...
__errex() {
printf 'Fatal error [%s] on line %s in '"'"'%s'"'"': %s\n' \
"${1:-"?"}" \
"${2:-"?"}" \
"${3:-"unknown script"}" \
"${4:-"unknown error"}" >&2 ;
exit "${1:-1}"
}
alias die='__errex "$?" "${LINENO}" "$0"' some-command || fail "message"
to produce a stack trace and exit the shell in case of non-zero exit status from some-command, or write some-command || softfail "message" || return $?
in case you want to produce a stack trace and return from the function.What I find compelling about bash is its position in relation to other languages and tools. It's ideal for tying together other tools and is close enough to the operating system to make that easy whilst also not requiring libraries to also be installed (c.f. python).
I often hear the opinion that more complex scripting should be moved to a language such as python, but that adds a layer of complexity that is probably not helpful in the long-run. I can take a bash script that I wrote twenty years ago and it'll still work fine, but a python programme from twenty years ago may well have issues with versions.
I think bash/sh’s key feature is that they are anti-entropy, there’s no development or evolution so there’s no chance you need to mess with dependencies or new features, the stuff that worked 20 years ago will continue to be the “bread and butter”. By design, this results in a system that’s averse to change and incentivizes people to reach outside of its limits when they are met.
In my opinion, bash has two things (at least vs NetBSD's shell, possibly a few more vs POSIX) that make the average shell script (that I write) much easier. The first is &> which makes it easy to redirect both stdout and stderr to a file for logging. The standard 2>&1 can work but needs to be placed correctly or it doesn't work. That place isn't always the obvious place like it is with &> and running bash seems much preferrable to me than figuring that one out.
The second is ${var@Q} which prints var quoted for the shell, which is nice to use all over the place to make sure any printed file names can be copied and pasted.
My sense is that targeting POSIX is usually done for maximum portability or for use on systems that don't have bash installed by default. However, bash is quite widely available even if not by default and very widely used so I wouldn't say it is unreasonable to look at bash as the de facto standard and POSIX and other shells as being used in more limited circumstances.
I feel like when I see a shell script in my work, which is not in operating systems development of course, people are targeting bash. I agree many things are careful to target sh for certain reasons (e.g. a script that runs in a container where the base image doesn’t have bash installed) but i still think GP’s question is interesting because it’s not common to see, say, a zsh shell script, but seeing #!/bin/bash is super common.
Android and Apple are in similar situations.
Bash/sh is good for when you need to combine some commands and what needs to be done can be accomplished mostly by CLI commands with a little glue to tie them together. Some times it is surprising what can be accomplished. I wrote a program to import pictures from an SD card on Windows using C#, copying pictures to C:\Pictures\YYYY\MM\DD according to the EXIF data or failing that, file time stamp. I tried to port it to Linux but ran into problems trying to connect to the EXIF library. After struggling with that, I rewrote it using sh, some EXIF tool and various file utilities. It took 31 lines, about half of which were actual commands and the rest comments or white space.
A much bigger project is a script to install Debian with root on ZFS. It's mostly a series of CLI commands with some variable substitution and conditionals depending on stuff like encrypted or not.
Once I’ve learned bash, I realised how much more problems i could solve, in addition to a majority of old ones. It’s an entirely new level of “computer literacy”; and a more genuine one.
But bash is so bad I wrote a ton of namespace shortened utils for using groovy scripts.
Sooooooooooooooooo much better. use IDEs for dev, save library system, groovy smoothed almost all Java annoyances
#!/bin/bash
die() { echo "$1" >&2; exit 1; }
cat myfile | while read line; do
if [[ "$line" =~ "information" ]]; then
die "Found match"
fi
done
echo "I don't want this line"
..."I don't want this line" will be printed.You can often avoid subshells (and in this specific example, shellcheck is absolutely right to complain about UUOC, and fixing that will also fix the die-from-a-subshell problem).
But, sometimes you can't, or avoiding a subshell really complicates the script. For those occasions, you can grab the script's PID at the top of the script and then use that to kill it dead:
#!/bin/bash
MYPID=$$
die() { echo "$1" >&2; kill -9 $MYPID; exit 1; }
cat myfile | while read line; do
if [[ "$line" =~ "information" ]]; then
die "Found match"
fi
done
echo "I don't want this line"
...but, of course, there are tradeoffs here too; killing it this way is a little bit brutal, and I've found that (for reasons I don't understand) it's not entirely reliable either.An arithmetic expression that evaluates to zero will cause the script to exit. e.g this will exit:
set -e
i=0
(( i++ )) # exits
Calling a function from a conditional prevents `set -e` from exiting. The following prints "hello\nworld\n": set -e
main() {
false # does not return nor exit
echo hello
}
if main; then echo world; fi
Practically speaking this means you need to explicitly check the return value of every command you run that you care about and guard against `set -e` in places you don't want the script to exit. So the value of `set -e` is limited.I'd still prefer a `grep <args> || true` over not having `set -e` for the whole file.
First of all, sending sigkill is literally overkill and perpetuates a bad practice. Send `TERM`. If it doesn't work, figure out why.
Secondly, subshells should be made as clear as possible and not hidden in pipes. Related, looping over `read` is essentially never the right thing to do. If you really need to do that, don't use pipes; use heredocs or herestrings.
Fourth, if you cannot avoid subshells and you want to terminate the full script on some condition, exit with a specific exit code from the subshell, check for it outside and terminate appropriately.
https://unix.stackexchange.com/questions/169716/why-is-using...
https://unix.stackexchange.com/questions/209123/understandin...
Both of these posts must be read carefully if you really wish to write robust scripts.
Reasons of 1) performance, 2) readability, and 3) security are provided as points against the pattern, and the post itself acknowledges that the pattern is a great way to call external programs.
I'd think that the fact that one is using shell to begin with would almost certainly mean that one is using the subshell loop pattern for calling external programs, which is the use case that your post approves of. In this case, subshells taking the piped input as stdin allows the easy passing of data streamed over a file-descriptor, probably one of the most trivially performant ways of data movement, and the pattern is composable, certainly easier to remember, modify, and to extend than the provided xargs alternative, without potential problems such as exceeding max argument length. Having independent subshells also allows for non-interference between separate loops when run in parallel, offering something resembling a proper closure. In these respects, subshell loops provide benefits rather than pitfalls in performance and readability. Certainly read has some quirks that one needs to be aware of, but aren't much of an issue when operating on inputs of a known shape, which is likely the case if one is about to provide them as arguments to another command.
Regarding "security", the need to quote applies to anything in shell, and has nothing specifically to do with the pattern.
Why? I do it quite often, though admittedly usually in one-time scripts.
set -euxo pipefail
at the top of my bash scripts. It makes some conditional testing more difficult but it has paid for itself many times over just because of pipefail
I use functions.sh in all of my scripts that are known to be running on Gentoo only and it makes them feel Gentoo-y and is useful in general.
Why doesn't set -e (or set -o errexit, or trap ERR) do what I expected? https://mywiki.wooledge.org/BashFAQ/105
What are the advantages and disadvantages of using set -u (or set -o nounset)? https://mywiki.wooledge.org/BashFAQ/112
Safe ways to do things in bash https://github.com/anordal/shellharden/blob/master/how_to_do...
Better Bash Scripting in 15 Minutes https://robertmuth.blogspot.com/2012/08/better-bash-scriptin...
Writing Robust Bash Shell Scripts https://www.davidpashley.com/articles/writing-robust-shell-s...
Personally, I'm a big fan of BASH3 boilerplate: https://github.com/kvz/bash3boilerplate
It's fine for BASH versions above v3 and provides decent logging though I typically extend the script so that I can pipe long running commands into its logging framework. It also ensures that you specify the "help" options correctly as it parses the usage information to process the command line arguments with support for short and long options.
The ability to see all variable names and its value is priceless.
OP's comment is not unfunny and not 100% untrue either though. But not 100% true either. A single word script still needs to be debugged.
Over a decade into my career, and I've successfully managed to avoid debugging Bash scripts.
There is hope for people who don't want to.
I wouldn’t want to write an entire application in Bash, but equally I wouldn’t want to write a script which does relatively simple file operations in Python. Bash is a language which has been honed over many decades for precisely that sort of thing, and so can communicate what’s happening far clearer than Python does in my view.
And, unfortunately shell has become the norm in CI/CD environments, pipelines etc. Can be convenient at times but can also be inconvenient and confusing as these scripts don't run in interactive shells.
A pipeline which relies on shell is not worth using, tbh. That's how much shell sucks.
I agree that bash sucks, but have yet to find anything to replace it that doesn't increase complexity and version problems.
> Image of a comic. To read the full HTML alt text, click "read the transcript".
but I can't find any button relating to a transcript.
fail-unless() {
local result
"$@"
result=$?
if ((result != 0)); then
echo >2&1 "Failed ${result} with command '$*'."
exit ${result}
fi
}
That way, I know exactly what failed in the script.Had to change it to drop the parantheses, then it worked, like this
trap 'read -p "[${BASH_SOURCE:-}:${LINENO:-}] ${BASH_COMMAND:-}"' DEBUG