What I learned from others' shell scripts
fizerkhan.com
fizerkhan.com
Here's what is done there:
OK=0
FAIL=1
function require_curl() {
which curl 2>&1 > /dev/null
if [ $? -eq 0 ]
then
return $OK
fi
return $FAIL
}
(`2>&1 > /dev/null` drops stdout and writes stderr to stdout—not what was meant. `which` doesn't write to stderr, so I drop that part.)That can be shortened significantly by using the return code directly in the if branch:
function require_curl() {
if which curl > /dev/null
then
return $OK
fi
return $FAIL
}
Or by just using the return code directly: function require_curl() {
which curl > /dev/null
return $?
}
And as it will return the return code of the last statement executed: function require_curl() {
which curl > /dev/null
} command -v llh >/dev/null 2>&1 ; echo $?
0 # good
usr/bin/which llh >/dev/null 2>&1 ; echo $?
1 # good
however (zsh) shell internal which does: which llh ; echo $?
llh='ls -lh'
/usr/bin/ls
0 #good
(I'm not sure however when ZSH chooses to use its internal function or the one from $PATH) $ which abcdasdf 1>/dev/null
which: no abcdasdf in (/usr/lib/mpi/gcc/openmpi/bin:/usr/local/bin:/usr/bin:/bin:.....snip)
$ uname -srv
Linux 3.4.47-2.38-desktop #1 SMP PREEMPT Fri May 31 20:17:40 UTC 2013 (3961086) $ which --version | head -1
GNU which v2.20, Copyright (C) 1999 - 2008 Carlo Wood.
Or may be the output of lsb_release, although it may not be installed in your system.Is "which" behaviour dependant on the kernel version?
EDIT: formatting; also to add that which writes to stderr if the command you're asking for can't be found in the path.
function debug()
{
if [ -v VERBOSE ]
then
echo "${1}" # might want to add: >2
fi
}
function has_dep()
{
unset dependency
dependency=$(command -v "${1}")
if [ -v dependency ]
then
debug "found \"${1}\": ${dependency}"
return 0
else
debug "\"${1}\": not found"
return 1
fi
}
# example:
VERBOSE=1
if has_dep curl
then
echo "we're good"
fi
if has_dep notfoundthing
then
echo "not good."
fi Betty:srcs lelf$ perl -E 'say "out"; say STDERR "err";' &> out && cat out
err
out
Betty:srcs lelf$
GNU bash, version 3.2.48(1)-release (x86_64-apple-darwin12) APP_ROOT=`dirname $0`
filename=`basename $filepath .html`
should read: APP_ROOT=`dirname "$0"`
filename=`basename "$filepath" .html`
Plus if your script supports both `--version` and `--help` with proper formats, you can easily generate a manual page with `help2man`. The help output example is far from that. Also, he does not mention getopt at all.The magic line for that is:
eval "set -- $(getopt -o hV -l help,version -- "${@}")" || exit $?
echo "${*}"
It does this: $ ./test foo bar --help --version
--help --version -- foo bar
Which can be parsed using a `while` loop with `shift`.For other tips on shell scripting best practices I recommend Gentoo ebuilds. They tend to be quite well written. And always read scripts from other people before running them yourself.
Also I would like to point out that you shouldn't keep using capital letters in your scripts because it might clash with environment variables. Treat them just like any other programming language, underscores or camel case for example.
I'd like to say "What I learned from idling in #bash on freenode" instead. That's a much better lesson.
However, are there are any people writing lint or prettify or Perl Critic like progs or web services that will properly wrap and do things you mention?
Would there be interest in such a thing?
UPDATE: I found this stuff. Does anyone use these things?
https://trac.id.ethz.ch/projects/bashcritic/
http://man.he.net/man1/checkbashisms
http://stackoverflow.com/questions/3668665/is-there-a-static...
wrap script into a shell function
source function definition
diff defined function against script
with something like: { printf 'dummy (){ '
cat script
printf "; };\ntype -a dummy\n"
} > dummy
source dummy > defined
rm dummy; unset -f dummy
diff script defined
Issues include: comments are absent from defined function listing
trailing semicolons appear after each shell command / line
the first semi in the trailing "; };" may be a syntax error
The default function listing format is fairly generic, so personal
formatting preferences must be left as an exercise for the reader.MHO: people who insist on specific code formats should provide tools that can convert arbitrary code into their preferred format.
The most interesting so far: Sh.py [0] (mentioned here before [1]) and plumbum [2], which is surprisingly extensive.
[0] http://amoffat.github.io/sh/
Edit: Place a general disclaimer to weaken the absolute tone of the above statement. Rule of thumb, exceptions apply, yadda, yadda. In other words, I somewhat agree with most responses that disagreed with the statement above :-)
Shell is still good for one thing: piping commands together. You can do that in Python, but I think even complicated output coloring and debugging scripts should stick to shell when the focus is piping utilities in and out of each other. Python can do that, but it is not its sweet spot, in my opinion. I am sure others disagree.
Different tools for different jobs, blah blah.
http://www.scsh.net/docu/docu.html
http://www.scsh.net/docu/html/man-Z-H-3.html#node_chap_2
I especially love the acknowledgements section :) http://www.scsh.net/docu/html/man.htmlGuess I was naive to think there was not without a proper Google-fu session.
UPDATE: I take that back, it appears some still works on it. That page just seems outdated, by a while now.
Disclaimer: I haven't used it and the author says it is "still pretty new." However, it seems to be under active development still. At least there was a commit to the repository within the last year.
shell -> perl -> python.
On OpenWRT systems, sometimes even microperl won't fit. All you've got is shell. (and sh only, none of this fancy bash stuff(1)) Its good to be able to do things at many levels.
(1) http://stackoverflow.com/questions/5725296/difference-betwee...
I use awk from time to time, specially when I need to process a file line by line and the task is too simple to use python (ie. an awk one-liner will do!).
Perl regexes are rightfully known for being cryptic but those awk statements make me cry for my mama.
I typically switch to Python (or similar) when I need something like multidimensional arrays / dicts (not portable in awk), or in general want to run more complex parsing or transformations on input. The more advanced string functions of bash I always have to look up (is it ## or %% that removes from the start of the string?), and for/while loops are in general a lot slower.
But for file system operations, stringing together commands, keeping track of backgrounded commands (using multiple threads without even thinking about it), and even turning arbitrary programs into "daemons" (using fifo's), nothing beats shell scripting.
- Filesystem operations. How many times have we seen programs using "ls * .foo" when they needed to use "find -name '*.foo'" to avoid command-line length limits? I answered this very question on StackOverflow this week. How many seemingly-capable software shops will churn out shell scripts which misbehave when a path has a space or other "weird" character in it? Even Apple fell prey to that one, about ten years ago, and destroyed some users' data.
- Stringing together commands. This encourages abominations like "cat myfile | grep foo | awk '...'". Just use awk if you're into that, but the shell has a knack for "tricking" people into spawning extra processes that are not really needed (indeed this is one of the most frequent performance sinks in shell scripts). And what about error handling for the several subprocesses? It's usually ignored for N-1 of them.
- Keeping track of backgrounded commands (using multiple threads without even thinking about it). Yes, you can use multiple cores without thinking--that can be cool. But what if you want to do N units of work on many fewer than N cores? You ought to use a pool, but there's no such thing in Bash. Maybe you're clever and use "xargs -P" for this, but most people don't.
- Turning arbitrary programs into "daemons". I use start-stop-daemon for that (it's included in Debian, and I easily wrote a workalike in Python when I had to use a system that didn't support it natively).
Just about the only thing here that shell scripts are really good for is doing things "without even thinking about it." Once when I was asked why shell scripting was not a good idea for production programs, I reviewed a smallish sample Bash script that had been deployed. I found a dozen latent bugs, 50% of which would have never have happened with Python (or Go, or...).
Re performance, I usually run Cygwin on Windows. Starting up too many processes is not a mistake I tend to make, because forking in Cygwin is hopelessly slow. Similarly, spaces in paths are common, and I've learned to be fairly religious about quoting, to the point of using print0, xargs -0 etc.
My scripts regularly deal with millions of files, and pipes that transfer tens of gigabytes. I rely on being able to string together sort and uniq to do set operations over multi-gigabyte files with constant memory usage; such scripts are not trivially rewritten in languages like Python without using non-standard libraries. When performance becomes an issue, the solution is a lot more heavyweight in terms of development time.
By the way, your "ls * .foo" does not even do the same as "find -name '*.foo'" due to a superfluous space ;-)
But the fact that most people don't know about xargs -P (or gnu parallel), or spawn too many processes, or use ugly hacks like pidfiles/start-stop-daemon, is not a reason to throw out all the good stuff shell scripting has to offer.
> "cat myfile | grep foo | awk '...'" […] indeed this is one of the most frequent performance sinks in shell scripts
Is it really? I would've thought loops were a more common performance sink. I can't imagine how that useless use of cat has _that_ much of an effect, unless you're running a whole bunch of copies of this script. It looks ugly in the process table and it does not let the real command (here: grep) move back and forth in the file, but I've never noticed performance improvements from removing uuoc's.
Not that I think it really makes much of difference in most cases.
Even this left me a but confused like how do the functions get the input parameters does the * in echo -e "$RED$*$NORMAL" have something to do with it?
There's just so many obscure ways of doing the same thing not to mention there's very few good places to teach you shell scripting and shell scripting best practices.
I'd rather just go with python. It's easier to read.
TLDP's Advanced Bash-Scripting Guide is a great resource for learning bash.
For example, this section talks about the meaning of "$*":
http://www.tldp.org/LDP/abs/html/internalvariables.html#ARGL...
For learning the differences between POSIX sh and bash (which can trip you up if you want to run on systems without bash, or if you simply don't know that #!/bin/sh is different from #!/bin/bash), see http://mywiki.wooledge.org/Bashism
And regarding how to parse your example, the $* in a string expands to all input parameters to the function, separated by the first character of the IFS variable (e.g. space).
I am sure no one will get that reference. :-)
The moment your script needs any more than 10 regular expressions, or dealing with >3 files all in a complex interplay- use of Perl becomes inevitable. And that is something like the very utmost basic thing you can do with Perl.
Python is more like tried-to-be-scripting-but-is-a-web-framework language.
In fact Django, Twisted and Zope is all the Python code there is.
Scripting was never Python's forte. Scripting is all about succinctness, power and providing as much power with fewer constructs and restrictions. Which happens to be exactly the very opposite of Python goals, and some thing which more or less as a mission statement Python tries to achieve. This is why Python will likely never be a very successful scripting language.
Python was always a web language for frustrated java programmers who couldn't put with java's problems anymore. Much of Python's success is in web frame work area. Which was previously Java's territory.
It's even more absurd to claim that "Django, Twisted and Zope is all the Python code there is." That's utter nonsense, in fact.
There are numerous non-web applications that use Python extensively, whether they're partially or fully implemented in Python, or whether they can be scripted using Python.
Then there are the numerous libraries and frameworks for Python, from GUI toolkits through to scientific computation packages.
Many, many organizations use Python internally for a very wide variety of tasks and systems, without broadcasting such use loudly, if at all. This ranges from one-off scripts up to entire multi-million-line software systems.
I sure hope that you're joking, but it just isn't coming across as a joke. Nobody can seriously claim that "Python will likely never be a very successful scripting language" when it has undoubtedly been one of the most successful, and versatile, programming and scripting languages around for many years now.
Of course Python is used for scripting purposes. But that amount of code is no where close the web code that is written in Python. It all depends how much code in ratio is written for what purposes and not the total amount of code written for that purposes.
Python's glory days came with web frameworks and continue to be the reason for its fame and wide spread adoption.
I will take Python seriously if it offers the same capabilities as Perl at least on the command line.
Python is a awesome general purpose language. But so far as scripting is concerned it is still no match to Perl
Many of the rest of us have been happily and very successfully using Python 3 for years now, including for the scripting tasks that you incorrectly claim we don't use Python for.
Your other claims like "Python's glory days came with web frameworks and continue to be the reason for its fame and wide spread adoption." are truly absurd and outright wrong. Python was very popular and widely used well before the mid-2000s, when many of the web frameworks you're referring to were first released.
I think you vastly overestimate the amount of Python used for web applications, as well. This is understandable, as such uses are often more visible than the other behind-the-scenes uses. But there's a staggering amount of Python code that you don't see, and that isn't used for web applications.
That is a pretty big assumption to make, in my opinion.
When we're talking in the context of scripting then we should be talking about what distros have easily accessible python3 packages/libraries, good python3 community support, and whether libraries relevant to sysadmin scripting tasks have been ported to python3 or not.
Regarding traps, I struggled for a while with getting a robust and simple way to kill backgrounded scripts on exit. I would have a script do stuff like "sort bigfile & pid=$!; runlongtask; wait $pid", and on Ctrl-C I wanted the sort command to stop too. The trick is to use "kill 0", which kills the non-interactive script and its subprocesses (see "man 2 kill"). So say you want to both remove some temp directory and kill all subprocesses on Ctrl-C, put this at the top of your script:
trap 'rm -rf "$tmp"; kill 0' EXIT
Usually, you never want your script to continue running after a command unexpectedly fails or you use an unset variable. Saves so much time debugging!
As an example, `sed -i` is used for in-place editing. BSD/OSX variants expect an extension argument to be provided after the flag, with an empty string for no backup (sed -i '' 's/foo/bar/' foo.bar)
On the other hand, gnu expects the backup extension in the flag (like -W in C compilers), so the command is parsed as if 's/foo/bar/' was the file to open
ssh remotehost "
$(declare -p var1 var2 var3)
$(declare -f func1 func2 remotemain)
remotemain"
In this example, var1, var2, var3 and func1, func2 are support variables/functions for the function "remotemain". This pushes all those to the remote side, then calls remotemain. $ ssh remotehost "arecord | gzip -c" | gunzip -c | aplay
records raw pcm stream an a remote machine, gzip's it and gunzip's and plays on the local machine. function require_curl() { which "curl" > /dev/null; }
function require_curl() { which -s "curl"; }
function debug { ((DEBUG)) && echo ">>> $*"; }
Last one is bash, which -s is BSD'ish debug() { [ "$DEBUG" ] && echo ">>> $*"; }This is not going to do what you expect if the current script is run via a command with spaces in one of the preceding directories' name.
It should be: APP_ROOT="`dirname "$0"`"
Even this would fail if your immediate parent directory's name ended with a newline character, but should handle any other whitespace without issue.
APP_ROOT=`dirname "$0"`
(See http://www.tldp.org/LDP/abs/html/varassignment.html#EX16)Also, even if ^J or ^M were present in a parent directory's name, the command above would still work. Whitespace-like characters take some getting used to with Bash.
To those curious: https://github.com/jalcine/dotfiles
which curl >/dev/null 2>&1
?
(what I learned from man bash)
One shouldn't use which at all. It forks an external process, it is vulnerable to PATH issues unless you specify what will surely be a non-portable path to the executable, often isn't included in chroot'd environments, and doesn't actually return an exit status on many platforms. In short: it sucks in almost all possible ways you could hope to suck.
The proper expression to use is:
command -v curl >/dev/null 2>&1
If you know you are in bash, type -P and possibly hash become acceptable alternatives, but then... why not just use command -v right?http://stackoverflow.com/questions/592620/check-if-a-program...
command curl # optionally 2> /dev/null
# but there really shouldn't be any
# output -- so any error you probably
# want to see...
do what's intended (exit with non-zero return if curl isn't available as a command)?[edit: Because command runs the command by default, it outputs info with -v!. So if the command does something by default (eg: command halt) -- then:
command halt #halts!
command -v halt #outputs something like /sbin/halt
Whops.]This is why (to redirect both stdout and stderr to $file) you have to do e.g. 1>$file 2>&1 rather than the more obvious-looking 2>&1 1>$file.
`> /dev/null 2>&1` will drop both stdout and stderr.
For myself, with `which` I would just use `> /dev/null`, as `which` writes to stdout and not stderr: if anything comes through stderr, you probably want to know about it.
PS. I think "which" will never produce anything on stderr