No script is too simple
nicolasbouliane.com
nicolasbouliane.com
Since 3.5 added `subprocess.run` (https://docs.python.org/3/library/subprocess.html#subprocess...) it's really easy to write CLI-style scripts in Python.
In my experience most engineers don't have deep fluency with Unix tools, so as soon as you start doing things like `if` branches in shell, it gets hard for many to follow.
The equivalent Python for a script is seldom harder to understand, and as soon as you start doing any nontrivial logic it is (in my experience) always easier to understand.
For example:
subprocess.run("exit 1", shell=True, check=True)
Traceback (most recent call last):
...
subprocess.CalledProcessError: Command 'exit 1' returned non-zero exit status
Combine this with `docopt` and you can very quickly and easily write helptext/arg parsing wrappers for your scripts; e.g. """Create a backup from a Google Cloud SQL instance, storing the backup in a
Storage bucket.
Usage: db_backup.py INSTANCE
"""
if __name__ == '__main__':
args = docopt.docopt(__doc__)
make_backup(instance=args['INSTANCE'])
Which to my eyes is much easier to grok than the equivalent bash for providing help text and requiring args.There's an argument to be made that "shell is more universal", but my claim here is that this is actually false, and simple Python is going to be more widely-understood and less error-prone these days.
if __name__ == '__main__': args = docopt.docopt(__doc__)
means nowt to me.
In fact I don't really understand your post and I've written a lot of python.
No language has a simpler api to executing a cli command than bash itself. By definition.
__name___ == '__main__'
simply tests whether this Python file is being executed from command line.It’s one of the ubiquitous Python idioms that doesn’t actually follow Python’s zen of “readability”.
Perhaps I shouldn't have tried to make two points in one post; I wouldn't advocate using docopt or a main function for a simple script, I was more making the case for how easy it is to add proper parameter parsing when you need it; that is easier to remember than what you'd end up writing in bash, which is something like:
PARAMS=""
while (( "$#" )); do
case "$1" in
-a|--my-boolean-flag)
MY_FLAG=0
shift
;;
-b|--my-flag-with-argument)
if [ -n "$2" ] && [ ${2:0:1} != "-" ]; then
MY_FLAG_ARG=$2
shift 2
else
echo "Error: Argument for $1 is missing" >&2
exit 1
fi
;;
-*|--*=) # unsupported flags
echo "Error: Unsupported flag $1" >&2
exit 1
;;
*) # preserve positional arguments
PARAMS="$PARAMS $1"
shift
;;
esac
done
# set positional arguments in their proper place
eval set -- "$PARAMS"
I think you've got to write a lot of bash before you can remember how to write `while (( "$#" )); do` off the top of your head; the double-brackets and [ vs ( are particularly error-prone pieces of syntax.I agree that a cli framework is often easier to use than bash, but it also is a dependency. I think everyone should use what he’s familiar with.
while getopts "ab:" OPTC; do
case "$OPTC" in
a)
MY_FLAG=0
;;
b)
MY_FLAG_ARG="$OPTARG"
;;
*)
# shell already printed diagnostic to stderr
exit 1
;;
esac
done
shift $((OPTIND - 1))
Yes, I realize it's more difficult to support long options (though not that difficult), but the best tool for the job will rarely check all the boxes. Anyhow, the arguments for long options are weakest in the case of simple shell scripts. (Above code doesn't handle mixed arguments either but GNU-style permuted arguments are evil. But unlike the Bash example the above code does support bundling, which is something I and many Unix users would expect to always be supported.)Also, I realize there's a ton of Bash code that looks exactly like you wrote. But that's on Google--in promoting Bash to the exclusion of learning good shell programming style they've created an army of Bash zombies who try to write Bash code like they would Python or JavaScript code, with predictable results.
Things that are true "by definition" are not usually useful.
"No language has a simpler api [...] [b]y definition" only if by "executing a cli command" we mean literally interpreting bash (or passing the string through to a bash process). But that's never the terminal[0] goal.
The useful question is whether, for the task you might want to achieve, for which you might usually reach for the shell, is there in fact a simpler way.
It may very well be that the answer is "no", but support for that is not "by definition".
[0]: Edited to add: ugh. Believe it or not, this was not intended.
Usage: db_backup.py INSTANCE
then just go with it.Edit: The equivalent of above would be:
#!/usr/bin/perl
my $shell=(getpwuid($>))[8];
print "$shell\n";
Though it does have simple syntax for pipes as well. # Option 1 (no pipes)
import os
for line in open("/etc/passwd").read().splitlines():
fields = line.split(":")
if fields[0] == os.getlogin():
print(fields[-1])
# Option 2 (Cheating, but not really. For most problems worth solving, there exists a library to do the work with little code.)
import os
print(os.environ['SHELL'])
# Option 3 (pipes, not sure if the utf-8 thing can be done nicer somehow)
import subprocess
username = subprocess.check_output("whoami", encoding="utf-8").rstrip()
p1 = subprocess.Popen(["grep", "^" + username, "/etc/passwd"], stdout=subprocess.PIPE)
p2 = subprocess.Popen(["awk", "-F", ":", "{print $NF}"], stdin=p1.stdout, stdout=subprocess.PIPE)
print(p2.communicate()[0].decode("utf-8")) subprocess.check_output("grep ^$(whoami) /etc/passwd | awk -F : '{print $NF}'", shell=True, encoding="utf-8").strip()
I didn't test that particular line, but in general this is how I execute shell pipelines in Python. pwd.getpwnam(getpass.getuser()).pw_shellIn bash, you have to use commands for most things. In Python, you don't.
For example, `curl ... | jq ...` shell pipeline would be converted into requests/json API in Python.
The exception being if it's a Python app in the first place, then it's fine since it can be assumed you need to be familiar with Python to hack on it anyway. For example I write a lot of Ruby scripts to include with Sinatra and Rails apps.
For a general purpose approach however, everybody should learn basic shell. It's not that hard and it is universal.
It's not a matter of familiarity with linux shell. It's a matter of wasting time debugging/implementing stuff in shell that is trivial to implement in any modern scripting language.
But even JSON and (X|HT)ML you'd be amazed how often a simple grep can snag what you want, and regular expressions are also mostly universal so nearly everybody can read them.
Python wasn't our main language (the app was in C++ and QML which is basically java script). But we both knew Python and it's the easiest language for these sort of things that we both knew and that comes with the system.
My main conclussion is - I'm starting with python the next time.
Life's too short to check for the number of arguments in each function.
Any complex or cross-platform C++ projects will need a scripting language in addition to shell and a build system.
Ha! The Oil shell does that for you :)
https://www.oilshell.org/release/0.8.0/doc/idioms.html#use-p...
(It's our upgrade path from bash)
[0]: I'm sure there's a jq style library for python somewhere but it's definitely not the norm.
and yes, parsing XML or JSON in "linux shell" is very easy. For JSON use jq, for XML use xml2 or xmlstarlet.
Fortunately for me a 10 whole lines of sed, awk and some trimming and a for loop turned my json into a CSV to pass back to a legacy system. No curl or json modules necessary. That being said, if I needed anything more complex Python would be my default.
I really like Ruby as a replacement shell language because its syntax for shelling out is very succinct and elegant. Python is a bit more verbose but a lot more explicit. Either one is a step function better than shell. You get a real programming language.
Let's extend your analogy to a logical conclusion. Should you write something in a shell because it's "not that hard and universal" or should you insist on using a programming language that lends itself to writing maintainably? If not, we should have no problems with PHP and COBOL, no? But we do.
Use the right tool for the right job. If you're not sure what that is, don't hesitate to pull out a glue language. Python, Ruby, JS -- whatever you need to get the job done. Your shell should be just that -- a shell, not the core.
And after all that it broke once it moved to a AIX machine since it was not POSIX compliant.
Maybe it's obvious to you, but when I cannot even figure out if an if statement will work, something is horribly wrong.
I don't even use python and I can just about guarantee it's easier for someone who knows neither.
listing = `ls -oh`
puts listing
puts $?.to_i # Prints status of last system command.
Some equivalent Python 3.7+ would be
import subprocess
listing = subprocess.run(("ls", "-oh"), capture_output=True, encoding="utf-8")
print(listing.stdout)
print(listing.returncode)
As an experienced Ruby programmer I imagine the code you provided looks like a no-brainer. To a Ruby neophyte the backticks, $? sigil, and .to_i method don't strike me as intuitive (but maybe they would be do someone else). We may just have to disagree about the nicety of syntax.As cle mentioned[1], the strongest criteria is
> what would the majority of people on the team prefer to use?
and I'd agree that ultimately that's what makes the most sense.
Is it mostly the Make-style task decorators for managing the project task list that you consider a win vs. native subprocess.run? Are there other features that you find valuable too?
(Back when the original fabric was written subprocess was much less friendly to work with, of course).
just put this on the first line #!/usr/local/bin/node instead of
#!/usr/local/bin/python
Fabric is the best for this: you put a `fabfile.py` (name inspired by Makefile) in the project root and add @task commands in there. Specifically the `fab-classic` fork of the Fabric project https://github.com/ploxiln/fab-classic#usage-introduction (the mainline Fabric changed it's API significantly, see `invoke` package).
Fabric can run local or remote commands (via ssh). Each task is self-documenting and the best part it that it's super easy to read... almost like literate programming (as opposed to more magical tools like ansible).
Shell is simple and straight forward, and can bring you quite far for most stuff. Only if you start using complex datastructures it makes more sense to use a mature language like python.
My last gig, I stumbled onto shebang, which allows you to invoke shell scripts in any scripting language from the command line.
https://en.wikipedia.org/wiki/Shebang_(Unix)
Major face-palm-slap. I'm certain I've read about shebang many times. But never realized that it's general purpose.
One nice thing about it is the overloading of pipe operator to allow easy creation of pipelines and/or redirection (plus it is cross platform, so you don't depend on the system having a certain shell).
Okay, but what about software engineers?
make vs. npm start make thing vs. npm run thing
But also more importantly, the Makefile itself feels much cleaner, rather than stuffing scripts in package.json files. The thing I'm confused about, why do npm scripts seem to dominate? I don't see many people using Makefiles.
[1] https://github.com/umaar/video-everyday/blob/master/Makefile
Sure, you can run Make and even shells on windows now, but that doesn't mean you have the power or even capability to do so in a corporate environment where IT dictates only reluctantly provides you with a machine in the first place.
If the Makefile needs to be cross-platform, that's easy to do, too:
```
ifeq ($(OS),Windows_NT)
# Windows
else
# *nix
endif
```
So many projects I've worked on have had nontrivial build-logic or implicit assumptions in their CI configuration which leads to distrust of local reproducibility. Eliminate as many variables and steps to reproducibility as possible. At the very least, create a `ci.sh` script in your project and make all your CI invocations go through that.
If you avoid putting "raw" npm or whatever invocations in your CI config in the first place you won't be tempted to add "just one more option" to your CI config and forget to add that option when running locally.
Such a script also becomes a de-facto API for how to use your project. This can help immensely if all of a sudden, for example, you switch programming languages (or major versions) environment assumptions, etc.
And yes there shouldn't be too much of a difference between what developers do locally and what the CI does.
It's sort of an anti-pattern if you have a bunch of automation that can only the CI can run. The developer should be able to run it on their machine too, without CI, and without YAML ...
[1] http://www.oilshell.org/why.html#what-you-can-use-right-now
The easiest way I found to debug it, is to put set -x into the code to print the command with substituted variables.
Wait, what? If your CI isn't testing the builds your bugs are (possibly) in then what's the point of CI?
OTOH, the monster I'm currently implementing as a set of ci.sh-like scripts is already way too big. It all started as a couple `kubectl apply`, but now I wish it was an ansible playbook.
Some other people on the same project abused groovy to the fullest to create some fancypants Jenkins pipeline. Local reproducibility? Zero.
I use that same concept extensively for all my projects, whatever the language/stack, though I prefer to use Makefiles since I feel like it creates a more cohesive result. Very easy to chain, depend and read in the CI too.
https://gist.github.com/cgarvis/61d70eeb1288bfee540c59cad095...
https://git.eeqj.de/sneak/mothership/src/branch/master/Makef...
As a vim user, my recommendation is to skip spacemacs, and go for straight emacs (with evil mode if you like modal editing, i do)
I liked Uncle Dave's Tutorial series on youtube: https://www.youtube.com/channel/UCDEtZ7AKmwS0_GNJog01D2g/vid...
Start with his emacs tutorial, it is mostly learn emacs and org mode as you learn to setup emacs.
Having some base level understanding makes understanding other Emacs configurations (like Spacemacs, Doom, Prelude, etc) much easier in my opinion.
Hearing the love so many have for spacemacs, I started there first instead.
Quite early on, I ran into problems. Every time I reached out on various forums I was told either: you're doing it wrong, that's a non-issue, RTFM (which isn't helpful when you don't know what you're looking for), or my favorite you have an XY problem (I didn't). So I'd go back to vim and put emacs on the back burner for a while longer, waiting for spacemacs to mature.
After the third attempt at spacemacs, I gave up and started looking for a good emacs tutorial.
Again I ran into some issues, but I found the regular emacs people very welcoming and helpful. Pretty soon I was able to diagnose my own issues, and figure out what settings I needed to change to meet my needs.
In the end, that early advice was true. You need to have some understanding for emacs to help diagnose spacemacs issues.
Will I give spacemacs another shot? Maybe one day, probably around the time the update their main release. It's been what, 2.5 years since they updated the main branch?
I've been using Emacs since early 2010s as well, so I'm biased - I had my own convoluted elisp modules before Spacemacs came around :).
This is not to preach of emacs or vim, really. I'm just saying vim and emacs are by no means mutually exclusive. I personally never got used to vim stuff, so I use Spacemacs with emacs keybindings, and my custom elisp scripts. Emacs really is more of a programming environment/mini operating system than an editor. Enjoy!
Emacs has a great Vi implementation, Spacemacs. Neovim is also a good Vi implementation. Vim, I think, is a bit outdated. For example, VimL scripting is full of quirks.
I love org-mode. It's the killer feature of Emacs in my mind. I also feel like (note, this is entirely anecdotal and not based on hard facts) Emacs has better LSP integration than Vim. I mainly use Go, so it could also be that gopls has become more stable than it was a year ago when I was first trying to get Vim working with it.
It's a small, but noticeable improvement over the way I was working before, either up-arrowing until I found the last time I ran tests, or typing, `pytest...` and letting autocomplete figure out what I was doing previously.
edit: So yes, I am also a big fan of helping to enforce consistency by scripting even the small things.
t() {
bazel test //some/specific/thing/...
# or pytest, maven, sbt, etc etc
}
Add in whatever option you want to the test command there. Then you just press “t” and it runs your tests.just as a substitute for up-arrowing, have you tried ctrl-r/reverse-i-search? (if that's a thing in fishshell)
edit: I see wodenokoto already mentioned this workflow in another comment
k: history-search-backward j: history-search-forward
There is an equivalent for emacs mode.
Here's the script, if you're interested. It's not super complex. https://github.com/wheybags/stuff/blob/master/build-dir-comm...
function retest
set --local cmd (history | grep -v history | grep -Em1 "((bundle exec)?rspec|go test|jest)")
echo "Rerunning `$cmd`"
eval $cmd
endFor instance, deploying to a specific target could be "./scripts/deploy.sh stage", backporting a patch could be "./scripts/patch_release.sh 1.1.0 dbcde45", and creating a database migration script could be "./scripts/db_changed.sh 'add new field for model'"
IMO thinking about the verbs is the first important step, but one should also always specify the nouns explicitly.
http://www.oilshell.org/blog/2020/02/good-parts-sketch.html#...
http://www.oilshell.org/blog/2020/07/blog-roadmap.html#more-...
As shown there, a lot of people are doing this under different names... I hope Oil can provide something consistent.
While the idea of having Make's dependency engine is nice in theory, it falls down for one-off automation in my experience. For a couple reasons:
(1) Make's dependency model has some well known deficiencies. It doesn't play well with tools that produce two files. It doesn't play well with tools that produce a directory tree of "dynamic" filenames (not known when you write the Makefile)
(2) Most makefiles have bugs, especially when you do make -j (parallel builds). It's basically like writing a C program with a bunch of threads racing on global variables -- your Make targets will often be racing on the same file, leading to non-deterministic bugs.
-----
So IMO it's better to mostly stick with the sequential model of shell for this kind of project automation.
But shell can invoke make! When you know your dependencies, invoke them from run.sh! And when you're REALLY sure it's correct, invoke make -j :)
In other words shell is my default, and make is only for when I want to spend the effort to write dependencies -- which is quite difficult, because Make provides you virtually no help with that. Bugs in dependencies are common and hard to find. If you're trained to run "make clean", then that's a symptom of a bug in the build specification.
Shell "gets shit done" without these types of bugs. Debugging a shell script is very easy compared with debugging a makefile. The remaining problems with shell will hopefully be fixed by https://oilshell.org/ :)
I do want to add some dependency support, but I didn't get to it:
Shell, Awk, and Make Should Be Combined http://www.oilshell.org/blog/2016/11/13.html
Make is just a small elaboration on the shell model (concurrent processes and files), but unfortunately it's implemented as a completely separate tool that shells out to shell! facepalm
These are covered by pattern matching and prerequisites, which were backported from mk(1) into GNU Make (Albeit they changed the syntax to make it incomprehensible).
> (2) Most makefiles have bugs, especially when you do make -j (parallel builds). It's basically like writing a C program with a bunch of threads racing on global variables -- your Make targets will often be racing on the same file, leading to non-deterministic bugs.
mk(1) makes it easier to avoid bugs and easier to comprehend what the makefile is doing, it also allows you to invoke programming languages for specific targets as a feature of the Makefile, and other goodies :)
https://9fans.github.io/plan9port/man/man1/mk.html
The GNU make pattern rules are OK, but they don't handle some pretty simple scenarios I've encountered in practice.
One was building the Oil shell blog, where the names of the posts are dynamic, and not hard-coded in the make file.
The other one is build variants. Say I have a custom tool with a flag. I want to build:
- two binaries: oil and opy
- two variants, opt and dbg
I want them to look like:
_build/oil.opt
_build/oil.dbg
_build/opy.opt
_build/opy.dbg
The % pattern rules apparently can't handle this. You need some kind of Turing complete "metaprogramming". (CMake and autoconf offer that, but mostly for different reasons.)This is a real problem that came up in: https://github.com/oilshell/oil/blob/master/Makefile
I plan to switch to a generated Ninja script to solve some of these problems.
mymove() {
cp --verbose $1 $2
rm --verbose $1
}
Shell has good abstraction capacities! It even has some that other languages don't have:Shell Has a Forth-like Quality http://www.oilshell.org/blog/2017/01/13.html
Pipelines Support Vectorized, Point-Free, and Imperative Style http://www.oilshell.org/blog/2017/01/15.html
http://www.oilshell.org/blog/tags.html?tag=shell-the-good-pa...
head () {
if [[ $# -eq 0 ]]
then
/usr/bin/head -$[(LINES-1)/2]
elif [[ -f "$1" ]]
then
case $# in
(1) /usr/bin/head -$[(LINES-1)/2] $* ;;
(2) /usr/bin/head -$[LINES*5/12] $* ;;
(3) /usr/bin/head -$[(LINES-1)/3] $* ;;
(*) /usr/bin/head $* ;;
esac
else
/usr/bin/head $*
fi
}By the way:
- Oil has set -euo pipefail on by default. (And also nullglob)
- Oil doesn't split words by default, so you don't need the quotes: https://www.oilshell.org/release/0.8.0/doc/simple-word-eval....
Those are all "opt in" -- OSH still runs POSIX shell scripts, if you want all the extra ceremony :)
That would truly be the opposite of this advice.
Maybe I need to think about this a bit more :)
isn't that what .aliases is for?
just give them some memorable names, and add them to your .bashrc. Or, if they are very context sensitive (that's not great), there is a way to source a file every time you enter a directory, I just don't remember it.
alias ls='ls --color'
Better: ls() {
command ls --color "$@"
}
(The "command" prefix avoids recursion)This is a short guide that is an eye opener as to just how much can be done with bash history: https://www.digitalocean.com/community/tutorials/how-to-use-...
And a Makefile sounds like it'd be pretty helpful too!
I used to have a scripts directory for each project, but I really wanted basic argument/options parsing and subcommands. That's mostly why I made mask. Now I use it daily as a project-based task runner, as well as a global utility command.
Just a small nitpick. I’d like it if we collectively moved away from including file extensions for scripts. You never know when you want to rewrite it in python or do something else. Nothing more confusing than opening up “backup.sh” only to find it’s actually a ruby script and must be executed.
Though it looks a bit too young for my taste: it's not available in most distro's base repositories yet, so it's going to be a tiny bit painful to deploy on every developer laptop, CI, and etc. I tend to prefer readily available tools like make, with 90% of the same features, but 100% distro coverage and previous developer knowledge.
Additionally, you automatically get shell completion too! https://github.com/NixOS/nixpkgs/blob/3bb54189b0c8132752fff3...
As described in the README, avoiding the 'build' part from Makefiles cut unnecessary complexity like .PHONY targets, which improves clarity. In smaller teams/companies it makes sense IMO.
#!/bin/bash
while sleep 0.1; do pacmd set-source-volume bluez_source.70_BF_92_CD_77_32.headset_head_unit 60000; done
Also, lots of scripts related to ffmpeg and other tools where the command line arguments are too hard to remember. For example, "ffmpeg-extract-sound-from-all-files-in-dir": #!/bin/bash
find *.mov | sed -e s/.mov// | xargs --replace=qq --verbose ffmpeg -i qq.mov -acodec pcm_s16le qq.wavAs such, people dedicated to the task of effort avoidance will find a way to avoid increased simplicity. The justifications and qualifiers are creative and elaborate often themselves taking great effort, often greater effort than that they originally sought to avoid. I see this behavior repeated so frequently in software and with such profound conviction.
A brief example is a single method call from a language or platform supplied API used to solve a problem instead of the entirety of a large framework. The heresy of such a disgusting travesty.
Not only that, but each script automatically runs itself in our docker container so it's fully repeatable. They all start with
. ./scripts/support/assert-in-container "$0" "$@"
which is just this: if [[ ! -v IN_DEV_CONTAINER ]]; then
scripts/run-in-docker "${@}"
exit $?
fi
And this is run-in-docker: https://github.com/darklang/dark/blob/main/scripts/run-in-do... bin/run
Then when I want to run it I just type the letter r, since I've aliased r to run ./bin/run. It's super fast. To keep source control clean I add it to: .git/info/exclude
Which allows my .gitignore to be the same as everyone elses. I used to used the global gitignore file, but I had issues at times.Same advantages as above, a more powerful cross platform language, cross-task references can be linted
def run(params):
"Run a built binary."
subprocess.run("./" + params.path)
.
.
.
if __name__ == __main__:
"Take CLI options and run selected function."""
# Parse CLI arguments
# Call functions with optionsbecause not all the tools you use in your system can import your python/similar file. Your deploy pipeline could involve 10 languages running under multiple os's / versions
or to put it another way... because the whole world doesn't run in your favorite programming language.
I usually come down on: things where the output is shared should probably be scripts, things where it's just for me probably shouldn't be (so that I can be more flexible, and there will be less bit-rot of those scripts).
Any other thoughts on this?
A key point is that when you want to "abstract", then the right way is usually to write a command line tool invoked from shell scripts in multiple projects. That is the natural way to get flexibility and reuse.
But otherwise it's a bunch of commands dumped in one place.
Examples:
http://www.oilshell.org/blog/2020/02/good-parts-sketch.html#...
Here's a random example -- some scripts to generate source code as part of the build process:
https://github.com/oilshell/oil/blob/master/build/codegen.sh
The paths are all hard-coded, which is a good thing.
In my mind, the goal is to save time and reduce mistakes. And having a consistent dev environment between everybody on a project is almost a prerequisite for that, and shell can actually enforce that consistency! (i.e. the shell scripts don't work if people have quirks on their machine. They can check the environment too.)
Setting up a pc went from 1-3 days to a 3 hour "run this setup.sh" script and call me if it crashes.
It will become obvious when you hire someone that this needs to be done. As a project lead, schedule time for it!
start.cmd
build.cmd
run.cmd