Use long flags when scripting (2013)
changelog.com
changelog.com
I think I'll continue to write these:
grep -i
rm -rf
ln -s
gzip -v
sed -e
instead of: grep --ignore-case
rm --recursive --force
ln --symbolic
gzip --verbose
sed --expression
If you're working on shell scripts, you probably know certain short flags well enough that long flags just add clutter and don't improve readability.For scripts that are read and executed more often than they're written or changed, the long flags really help ensure you don't fat-finger.
I'm accustomed to typing `rm -rf` when really I just need `rm -r` in a lot of cases, and in PRs it's very easy for your eyes to glaze over when there's more than a single short flag.
mkdir demodir && touch demodir/{foo,bar} && rm -r demodirI could've sworn that -I was the default......
I feel it does more harm than good by normalising rm -f when you want to recursively delete a folder, but with CD deployments these days it's less of a deal.
To be fair though they were referring to CD deployment(s), plural, which is a bit redundant and just CD would have done just as well in this case.
alias rm='rm -i'
alias cp='cp -i'
alias mv='mv -i'
I've also encountered one job which adds this to their base ubuntu image, and my current employer uses Macs which have it added to bashrc by their mdm software on initial install.So what's your justification to dismiss the idea that this is a common practice?
Thank you for the info, though.
The long flags are less likely to be changed over time. The script will usually fail gracefully and I don't also need to relearn what I was thinking 25 years ago when I first did it. (long flags are like in-command comments too)
And as others have pointed out -v(ersion) or -v(erbose) can happen too.
Many programs use -v as a short flag for --version, but some (such as curl) use it as short for --verbose
Probably something you'd catch pretty quickly, but still.
The default command was to just print the line. Hence the name grep
g/re/p
[1]: https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
Though I would make the list of "allowed" short flags very short. I can't think of many more than the ones you listed:
mkdir -p
sh -c
cp -r
tar -xaf
tar -cafI gave up and added this to my .bashrc:
function extract()
{
if [ -f $1 ] ; then
case $1 in
*.tar.bz2) tar xvjf $1 ;;
*.tar.gz) tar xvzf $1 ;;
*.bz2) bunzip2 $1 ;;
*.rar) unrar x $1 ;;
*.gz) gunzip $1 ;;
*.tar) tar xvf $1 ;;
*.tbz2) tar xvjf $1 ;;
*.tgz) tar xvzf $1 ;;
*.zip) unzip $1 ;;
*.Z) uncompress $1 ;;
*.7z) 7z x $1 ;;
*.zst) zstd -d $1 ;;
*.xz) unxz $1 ;;
*) echo "'$1' cannot be extracted via >extract<" ;;
esac
else
echo "'$1' is not a valid file"
fi
}Another "recent" (much more recent IIRC) change is that you don't need the dot anymore in find to search in the current directory.
extract() {
bsdtar -xf "$1"
}
you can decompress things which are not files, such as stdin or device files. in addition, your code does not cover *.tar.xz (more common than tar.bz2 nowadays), or lzma, or lz4, or tar.zst, or many many other formats. further, it's not even consistent: bunzip2, gunzip, and unxz remove the input file, but tar, unrar, 7z, and zstd do not.Personally, I consciously avoid using the internal decompression features of tar, both out of habit and to avoid unexpected results.
As such, I generally use command lines like: bunzip2 -c <tar.bz2 file> |tar xvf -
instead of relying on tar xvjf
I don't see it as wrong or bad to use tar's decompression features, it's more about my own preferences and experience. Being able to perform similar actions in multiple ways is one of the things I've always appreciated about the shell and the Unix/GNU userland.
Anyway, my advice to you and anyone else having trouble with tar is that "tar caf" and "tar xaf" are the only things most people need to remember about tar, with or without compression.
(In this case, most people means people who use tar, but so rarely that they have trouble remembering how to use it; they probably never use it for anything else other than (un)archiving. Also, xkcd isn't gospel.)
tar cf dir.tar dir # mnemonic: 'create file' <tar-file-name> <dir-to-tar>
tar xf dir.tar(.gz|bz2|...) # mnemonic: 'eXtract file' <tar-file-name>
On systems with "modern" versions of tar `-x` is capable of recognizing which compression format is used and doesn't require the explicit `-j/z` flags you usually see. $ tar --help
Obviously. :) $ tar --help
tar: unknown option -- help
usage: tar [-]{crtux}[-befhjklmopqvwzHJOPSXZ014578] [archive] [blocksize]
[-C directory] [-T file] [-s replstr] [file ...]
$ $ list '#! /bin/sh' 'echo "tar: invalid command" >&2' 'exit 64' >/tmp/tar
$ chmod a+x /tmp/tar
$ PATH="/tmp:$PATH"
$
$ tar -cf a.tar a/
tar: invalid command
$
0: any program that doesn't support --help is defective, which was rather my point.Fair enough. My point was that there's a category of operating systems where long options (including --help) isn't really a thing. You may of course consider them defective, though I'm not sure everybody agrees.
Even better. :)
Case in point, when I was learning, I'd copy and paste snippets I found online without understanding what some flags did. At first I didn't even know how to look up docs, and googling for the meaning of shorthand isn't always fruitful, especially given the "don't know how to find docs" limitation.
I've even ran into cases where the next person is myself. For example at one point, I "knew" docker flags while it was fresh in my mind, but then forgot what they meant a few months later...
Kids today no longer know the shell. They think just writing some yaml files in ansible is sufficient. Giving them a shell script won't help at all.
Sounds like Kochan and Wood[0] need some love (and royalties).
Or is learning stuff deprecated these days?
[0] https://www.amazon.com/Unix-Shell-Programming-Stephen-Kochan...
What about other people who might read the script?
Also is typing that much of a chore?
Gonna admit that I wouldn't be able to tell what grep -i, sed -e do. (I would've used `sed -e` before, but not enough to remember that). Still, some things are clear from context.
Doubt you'd be printing out a version number in a shell script though.
> Selected lines are those not matching any of the specified patterns.
For the lazy.
Edit: I don't know hackernews formatting and I'm among the lazy, so... Marked as "won't fix".
It's rare that I'd use it in a script. Perhaps I'd use it for a script whose main purpose is creating an archive to be widely distributed. Then the info might be of interest to the script's user.
But your comment brings up a good point about the difficulty of deciding what short flags truly are widely known. I assumed anyone who uses gzip regularly would know "-v", but maybe that isn't true.
But I also don't use gzip much these days. pbzip2 is a drop in replacement that is faster and compresses better.
Some command-line parsers will auto-complete partial options so you’re less likely to see ambiguity errors in future versions if you picked long-form options. (This isn’t completely foolproof, e.g. a command could have "--foo" and later add "--foobar" but it does help in most cases.)
And unfortunately, some tools with the same name will use the same letter to mean different things across Unix variants. You are asking for trouble if you aren’t being clear about what you want.
This sort of behaviour is tortuous, and should be against international laws.
Where _adding_ new, unrelated, options to the interface now changes or breaks the behaviour of existing calling scripts. Plus you never know which abbreviations are in use in the wild.
It makes it virtually impossible to maintain a stable interface without just freezing it long-hand.
I found this in Perl code; I believe it might be the default behaviour in the standard parser? Our developer really did like the philosophy that "the user may want 50 different ways to express the same thing".
1. Sadly, short options are in fact more portable than long options if you're targeting POSIX. For instance, macOS ships with versions of ls, rm, et cetera that only support short flags. This is because long options don't exist at all in POSIX (they're formalized here: https://pubs.opengroup.org/onlinepubs/009695399/functions/ge...).
2. Auto-completing partial options only happens for long flags, and it's a bit more than "some" parsers; the canonical implementation of "long" options, GNU's getopt_long, has it as a documented feature:
> Long option names may be abbreviated if the abbreviation is unique or is an exact match for some defined option.
https://linux.die.net/man/3/getopt_long
I 100% agree it's a poor feature. You basically entirely preclude yourself from adding features in a backwards-compatible way.
if (-not $(Test-Path -Path "$Path" -PathType Container)) {
Write-Error -Message "Path ($Path) does not exist or is not a directory." -Category InvalidArgument;
return;
}
When you do this properly, it feels like magic. For example, I wrote a script that does local Maven and Docker builds for a bunch of related projects. So I wrote two functions `Build-Maven` and `Build-Docker` with proper common flag support and error handling. Then, when I use them, I just do something like this: $PSDefaultParameterValues = @{
'Build-Maven:ErrorAction' = 'Stop';
'Build-Docker:ErrorAction' = 'Stop';
};
Build-Maven "$Path\A";
Build-Maven "$Path\B";
Build-Docker -Path "$Path\B" `
-Dockerfile "$Path\B\Dockerfile" `
-Tag "B:$Tag";
That first clause automatically amends `Build-Maven` and `Build-Docker` commands with `-ErrorAction Stop`. So if any of those build commands fail, the entire script halts there. And, if I pass in `-Verbose` to this script, that's forwarded to the build commands and I'll see the Maven and Docker build output.On the other hand, I've seen some people eschew `%` and `?` in favor of writing `ForEach-Object` and `Where-Object`, which in my opinion is too extreme in the other direction.
BTW, in:
if (-not $(Test-Path -Path "$Path" -PathType Container)) {
... you don't need the `$`, and if $Path is already a string you don't need the `""` either.I don't know how to feel about % and ?. The thought with not using them is that they are just aliases and could be changed. But that reasoning sort of breaks down since any command could be aliased to something dumb like `Set-Alias Get-Content Remove-Item`.
Let's be honest it's outdated, has a bad UX and is with the standards I would hold such a standard against today generally not so good.
Sure systems seen try to somewhat be POSIX compliance but from my experience not because they care about POSIX but because that happen to overlap with the idea to be somewhat compiland with other similar systems so that porting software/scripts is easier.
So if the systems you use support long options for POSIX commands go ahead and use them. Furthermore if the command you can are not in the standard anyway you again can use long options because they are as much standard as the short ones. Let's be honest this leaves very little use-cases where long options are a problem.
The most vivid example is test, aka /bin/[. POSIX specifies that, if test is invoked with one argument, you must "Exit true (0) if $1 is not null; otherwise, exit false" [1].
This means that `test --help` is forbidden from printing any help. It means that `test -d $argv` expands to `test -d` if $argv is empty, which then must "succeed" because "-d" is not null. It bites users over and over again [2].
Implementing this POSIX piece has resulted in more headaches, not fewer. I regret implementing a POSIX-conformant test, I wish I had just picked mostly-compatible but also-sane semantics.
1: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/t...
2: https://stackoverflow.com/questions/29635083/fish-shell-chec...
To be fair, while the POSIX behaviour is obviously wrong, using `$argv` on a bourne-style shell[0] is also wrong; that should be either `test -d "$argv"` or (if I recall the syntax correctly) `test -d "${argv[@]}"` if you actually intended argv to be a list rather than a single directory ("argv" suggests the former, but `test -d X` only accepts a single argument).
0: more specifically, a shell that does word splitting at the wrong place (after parameter expansion)
https://pubs.opengroup.org/onlinepubs/9699919799/
> Sure systems seen try to somewhat be POSIX compliance
The mainstream systems you know and use adhere very rigorously to POSIX. GNU Libc, Kernel, Coreutils, ... and their counterparts in BSD Unixes, proprietary Unixes, Cygwin and whatnot all take POSIX seriously.
POSIX is very helpful, and smart coders and sysadmins use it as one of their references for required behavior.
https://news.ycombinator.com/item?id=10104203
Although I would agree that there's a slight hole there: you still need to port stuff like coreutils and all the dependencies, which has been done of course. But I'd be happier if something like busybox was actually portable.
The Autoconf system is predicated on the idea that writing a portable shells script is hard, and so we hide the shell programming behind a mountain of M4 macros.
It's not necessarily easier to port a shell script today than 22 years ago, which was not written with portability in mind. There are more shells with more extensions, and then beyond the language concerns, and the environments have exploded. A shell script can easily depend on all sorts of cruft you've never heard of. Oh, just install these five things from the following github repos ...
If this is your own dotfiles or utilities, sure, but if I'm working on something collaborative, yes, I care about the standard.
It's just way too easy to run into a unexpected gotcha on some platform/configuration.
http://www.oilshell.org/blog/2018/01/28.html#limit-to-posix
Recent comment about this:
https://news.ycombinator.com/item?id=24428347
POSIX is just too limited and not what people use in practice.
If you want a portable shell script, in many cases my (biased) advice would be to make your script work on both bash and Oil. (Obviously there are short shell scripts which you may want to run on BSD, etc. This is more about big scripts, which POSIX falls down for.)
Oil already runs some of the biggest shell scripts in the world, many of them unmodified. Moreoever, when there's a patch necessary to run it, it often IMPROVES the program.
http://www.oilshell.org/blog/2020/06/release-0.8.pre6.html#p...
You'll be less tied to the vagaries of bash.
If anyone's script doesn't run under Oil, I'm interested. See https://github.com/oilshell/oil/wiki/What-Is-Expected-to-Run...
Though sometimes you don't need to go that far to break stuff, for instance switching from Fedora to Ubuntu. I've seen many scripts fail on debian derivatives because people think using #!/bin/sh as a shebang is fine since it works on their computer where sh was in fact a symlink to bash.
But on debian based distributions /bin/sh is often dash, not bash, and dash is basically the strict POSIX subset + local, all fancy stuff like [[ ]], &>, arrays, ... will fail.
Though this is less about long options here and more about general shell scripting.
But, still, that's no reason to adopt a mindset of actively wrecking the chances of such success.
(Which is what passive ignorance amounts to, effectively).
It's not difficult to do, but it is annoying. I don't blame people for ignorance, but it'd be nice if people thought about portability.
Non-portable scripts can also bite you on Linux, where Debian, for example, will swap out the shell from bash to dash for performance. but others will not. So, even within the Linux ecosystem, you can end up with scripts that behave differently across distros.
From the bash man page, "If bash is invoked with the name sh, it tries to mimic the startup behavior of historical versions of sh as closely as possible, while conforming to the POSIX standard as well." Also you should not be including bashisms if your shebang is sh and not bash. So using bash as sh shouldn't generally be a problem, if you stick to the published Shell Command Language [0], unless you hit one of the cases where the spec is a bit ambiguous and shells implement it differently [1].
[0] https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V... [1] https://stackoverflow.com/questions/16069339/different-pipel...
From what I remember, it doesn't do a particularly good job.
And, again, this still bites whenever using different linux distros.
#!/bin/bash
#!/usr/bin/env bashHomebrew comes to mind, and any other piece of software that does the installation via a shell script.
Even if the stuff will only ever run on one system, the POSIX flags (1) come from a smaller set of options, since POSIX is fairly conservative in its content this area and (2) are generally well known (often three decades old or older).
Don't make me read a man page to confirm that "grep --fixed-strings" really is the same thing as the POSIX-standard "grep -F", and not something subtly different.
You don't necessarily know what a long option means without looking it up, just because it is long. Firstly, if your native language isn't English, you may have to look it up in a dictionary. The ordinary meanings you find there may not reveal what that word means in the context of the program. So, at that point you're off to the documentation anyway.
I know what "hard" and "soft" are. Therefore, is it obvious what "git reset --soft" means? Hardly.
Once you know what it means, "--soft" jogs your memory by association a lot better than some "-X". So that would seem like it is less cognitive work. But when we have long options, it encourages the option vocabulary of a program to keep growing, which adds to the cognitive load.
Combined with use of shellcheck and shfmt working in shell scripting has come a long way in the last couple years. I still feel like BASH is a dead end--the quoting and whitespace issues combined with the bifurcation caused by Apple not shipping Bash 4.0+ make moving to Python probably the better course in 2020.
Bash probably deserves to die out, but Apple has really pulled a bait-and-switch on OSX's *nix/bsd underpinnings over the years.
I'd rather have my tonsils extracted through my ears.
I went the other way[0] decades ago and never looked back.
But I really would recommend Cygwin[1] for anyone who needs to use Windows.
Having a real shell makes a big difference.
command \
--longopt1 \
--longopt2 arg \
argument
or command | \
command 2 | \
command 3 \
--opt1 \
--opt2 \
arg cmd=(
command
-x -y
--bar OPTARG
ARG ARG
)
$cmd
This works in Zsh. In Bash that would be ${cmd[@]} I think. I use this with long qemu command lines where I often modify and comment out arguments during testing. command `# a comment` \
--longopt1 `# another comment` \
--longopt2 arg \
argument command `# a comment` \
# --longopt1 \
--longopt2 arg \
argument
AFAIK there's no hack for this.EDIT: Oh, `# a comment`. Didn't notice that. I'm going to explore it, thanks.
Otherwise, you'll get word splitting. E.g.,
cmd=(
test
"!= !="
!=
""
)
"${cmd[@]}" && echo Success
${cmd[@]} || echo FailureI'll add a comment or something explaining it but now I feel guilty and should probably go back and rewrite a few lines...
I mean—for long scripts I'll break them up but I haven't paid much mind to arguments/flags.
foo \
--group1 \
\
--group 2 \
--more-stuff 2 \
| \
bar \
--bar-stuff
using a line with a single backlash or breaking out the pipe can make grouping stuff clear. well, I mean as clear as having to use backslashes which to me is a little inelegant. Personally python and its indenting strategy has grown on me especially since editor support makes it painless to use and visually excellent.Long form also allows me to find things faster in man pages.
Do you know, without referring to the manpage, exactly what `cp --no-dereference --preserve=links --recursive --preserve=all` does? I don't. I can take some guesses based on the names of the options, but those guesses could very easily miss an important corner case.
Do you know what `cp -a` does? I do.
When live scripting, feel free to use short.
Like kubectl -f. In apply, it means file, in logs it means follow. Drive me nuts how they overload short flags with different means but no long version.
Or gsutil/gcloud.
I get google folks, don’t really care much about always having a long readable version of short flag.
Some may disagree, but unexpected flags passing through seems dangerous to me.
That would make finding and typing the long arguments easier. As it is, it’s easier to type the long arguments in a shell than in an editor.
I say just learn the flags or look them up, it shouldn't be hard or time consuming. When reading scripts, many times you can infer what the flags mean by understanding the inputs and desired outputs.
`tar -cvf` is much more recognizable than `tar --create --verbose --file`.
Moreover, why are you writing a shell script if it's NOT quick and dirty?
--r<TAB>ecursive