The temptation of writing shell scripts, illustrated
utcc.utoronto.ca
utcc.utoronto.ca
The -R switch reads in arbitrary text instead of json and presents it as a sequence of json strings internally.
The capture() filter applies regexes to the string and extracts subgroups (which will be represented as json objects).
The gsub() filter performs search-and-replace on a string.
Finally, the -r switch writes the result back to stdout as a raw value, removing any json quoting and escaping.
By default, processing is linewise, but there are options to process the whole input at once instead. There is also supposed to be a way how you can have per-line processing with state tracking across lines (like sed's "holding space"), but I haven't tried that out yet.
I found this approach often more flexible than sed and easier to reason about. Also, I believe it may be faster, as more things are happening in a single process.
The one exception: grep on files is still orders of magnitude faster than anything else. I believe there also was an article about this on HN a while ago.
So if you have to find some data in a larger collection of files and then have to turn the data into something structured, a good approach is to make a "coarse" query with grep or grep -R first, to narrow down your set of candidates, then postprocess your remaining candidates with jq.
Also, apropos single process: Another questionable advice is that for/while loops in bash are subshells and you can use them as part of a pipe chain. This often lets you reduce the number of processes you need.
E.g. instead of
for f in $files; do foo "$f" | sed ...; done
You can often write: for f in $files; do foo "$f"; done | sed ...;
Which will only launch a single sed process for all files instead of one per file.But yeah, shell scripts do have many shortcomings and in particular the error handling situation seems extremely hard to fix. So actually sensible advice is still: Use python.
I'd like to see that article! Faster than jq for sure, but my experience is that ack/ag/rg are all orders of magnitude faster than grep for typical uses.
Also I think there is a difference between grep on files, i.e. "grep ... file1 file2" or "grep ... -R somedir" and grep on stdin, i.e. "something | grep ...".
My understanding was that in the first case, grep can make use of random access APIs to control the size of the input chunks it processes or skip over parts that are known in advance not to match. That's not possible if the input is fed through stdin.
I haven't seen any actual comparisons though apart from "subjective" speed. So things might be wrong.
Shell is the native automation facility of the UNIX-like operating systems, and with 40+ years behind it, it is well understood, easy to write, easy to debug, and has no artificial dependencies (so, completely the opposite of Python).
Combine shell with AWK and you have an unbeatable combination for most computer automation and even very large scale data (pre)processing and number crunching.
Python is not really very slow overall unless you’re on windows. When you get to serious computing then breaking out of python is possible, but systems administration is not the place to do heavy computation.
As for debugability, I find python to be easier, set -/+x is great, but that’s not doing much more than printing out every line of execution, there’s no interactive debugger like python has afaik.
The killer thing for me is that python has modules for basically everything, which is awesome.
The thing that kills python for me is that I can’t be 100% sure if what version may exist on a system; and worse: if I actually use any modules then I have to somehow get them on the system.
For systems administration, this is pretty close to a non-starter. Prevailing sysadmin knowledge is that the admin tools should not horn in excessive dependencies, since sysadmin tools should work when the system isn’t fully set up.
This is why go is so great. But go is harder to debug than python in my experience.
pip handles that automatically - you set which version numbers of Python you can handle in your project specification.
> and worse: if I actually use any modules then I have to somehow get them on the system.
pip handles that too!
----
I'm not quite sure what the problems are, but between virtualenv/venv, pip, and modern package managers like poetry, this is really a solved problem - something you spend an hour setting up when you start a project and just never think of again.
I do this so much I have a shell function, `nenv`, which creates a new virtualenv with a specific Python and then loads its dependencies into the virtualenv.
It gives me tremendous freedom. For example, I can heavily instrument other modules' Python code inside the virtualenv (to find their bugs, e.g.) and then throw it all away and recreate it fresh in a few keystrokes.
---
There are OK, well-known solutions to most of the common packaging and distribution problems with Python programs you might have. If you go back to it, you should spend a bit of time with this, it isn't that bad.
(Poetry seems really cool - when I start something new it'll be with that.)
That said: This is very much a developers take on the packaging situation in python.
Running virtualenvs is "fine" until you're not connected to the internet, a less-and-less common scenario thankfully.
pip itself (without virtualenvs) will "dirty" the system, or you use the --user flag and make it work only for one user.
The situation for systems administrators is to do things that do not mess with the developers, if I install a package, especially with a version lock, and it directly conflicts with a developers (less and less of a problem with containers!) then they're going to be very cross with me.
I used to see python the same way, and coding on my megalithic single file would always take a couple hours to set up on a fresh computer. Then I discovered requirements.txt, where you simply list each package you need and the version (range) desired and then run:
pip install -r requirements.txt
this will install everything unless it's fucky like for example, pip install kivy-garden.matplotlib will not work with pip but everything else does.
Through the native software management subsystem: RPM, SVR4 packaging, pkgsrc, MSI... always use the native software management subsystem - that is what it is there for, for developers to deliver their software with.
Any high security environment (read: the financial industry) will have most of the servers purposely not connected to the InterNet, to prevent people from doing exactly the above and from attackers hacking in. Energy and pharmaceutical industries - same thing. That's the norm, not the exception, so pretty much any industry which is critical for society at large won't allow access to the InterNet.
One is not supposed to have more than one software management subsystem on a server because it leads to a hodge-podge mess: imagine you have 100'000 servers and it's RPM this, then pip that, then pear other, then npm whatever, then cargo something else... it would become a system administration nightmare in 0.1, and a system engineering nightmare in 0.01 seconds... you have to troubleshoot production, and do, what? Start firing different private packaging format commands in the hopes you find the correct one to find out what's been done to the system, what's available, then pray that particular private packaging system supports some sort of verification feature in order to find out who hacked which file with what tool? And that's okay by you? It certainly isn't okay by me and would have never worked in an environment that large (it doesn't even work in very small environments!)
...One is always supposed to use the native software management system of the operating system: cattle, not pets. And if there isn't a native OS package of the software one needs, or there isn't one of sufficient quality and standards compliance (like LSB FHS), then one makes one's own packages, because that is the high quality, professional way of doing things.
Most projects suitable for Python should take less than an hour end-to-end, not just for setting up package management.
(Why is Python hard to debug compared with shell? Python has a debugger - does Bash?)
"Bash"? Is that all you know, "bash"? I wasn't writing about that GNU crap of a shell, but of real, AT&T Bourne or Korn shell!
Debugging shell is trivial with set -x. I can find a problem in a shell program in seconds.
https://news.ycombinator.com/item?id=20633546
You are the embodiment of the Dunning–Kruger effect.
Don't you have that Kubernetes-Docker trash fire to babysit, instead of trolling around here?
Granted, Bash is quirky. So I should say I program in bash+shellcheck not just Bash.
And Jq. Jq is a functional programming language in itself! I do regex or string mangling inside jq filters.
Why should I use Python over this? The moment you want a clean and cool feature that Python has over Bash+jq, you should incorporate third party libraries. And adding just one non-std lib, means you have to resort to `virtual env`. And why not adding also `click` or `requests` or .. and now you have ~10 moving dependencies. Let's not talk about Python's subtle differences between versions.
Compare it to Bash-as-glue-of-unix-tools approach. You want image processing? Just use `imagemagick`. You want web scraping? `pup` is your friend.
ProTip: do `my_command || exit 1` if my_command is bound to fail. Or `if ! my_command; then printf >&2 "[ERROR] reasons"; exit 4 fi` for more elaborate exception-handling.
> Python, Go and so on need more structure and often don't have quite as simple and convenient methods of doing various shell like things.
This is Perl's niche, or at least it was Perl's first niche before it started getting used for CGI. The things you'd do in a shell script are generally convenient in Perl too, the main inconvenient part is selecting from the many different ways to invoke a program in Perl, like system(), backticks, qx//, open(), etc.
Perl is also more commonly installed and has faster startup time than Python, in my experience.
$ perl -E 'my $f=[[{X=>{Y=>["FOO"]}}]]; say $f->[0][0]{X}{Y}[0];my $g=$f->[0][0]{X}{Y}; say $g->[0]'
FOO
FOO
Seems pretty trivial.If it’s something that needs to be part of a complex system, sure higher level langs can be powerful. If it’s a standalone thing, shell scripts are generally ok.
In fact higher level languages add complexity: their build process, library updates, breaking changes of dependencies all make it a pain to maintain. Shell scripts? “ls” is never gonna change.
IMO the biggest reason shell scripts are avoided is simply because most software developers aren’t very comfortable with the syntax (tbh it does feel archaic in certain places) or know how unixy stuff like signals, pipes etc really work. If you have to do system level tasks though it’s probably worth the effort to understand these things.
Grepping through output No uniform handling of spaces, tabs. No data structures. Multiple dialects.
Shell scripts mainly work with files and process management. So many scripts don’t handle spaces correctly. This is not a problem of the creator, but because of the language.
There are 10 tiny buttons next to each other. You press the wrong one. Was that your fault or the fault of the designer?
Stop blaming "the user". In this case, not even a software engineer, but usually a sysop or power user.
If you ask me personally, I would not start a new project in any of the languages or frameworks you mentioned.
I still do not agree though, that the user is completely not at fault, when there are easily avoided traps (if that is what you are proposing?). Of course those should ne even exist. I agree with that obviously. Blaming every single mistake on the language design, which might be gnarly, is not what I think is the correct thing to do though. We do not have the perfect language yet (or might never have it). So some gnarliness is to be expected and that means, that a realistic perspective is to educate oneself about those gnarly things and avoid mistakes. Mistakes are human, so there will be mistakes, but we should try to avoid cheap ones out of laziness or being uninformed.
Also it is totally fine to criticize a language for its design flaws. I like to do that often even, with mainstream languages, because they can learn loads of stuff from less mainstream ones. This is the way to create awareness of those flaws and thus a way to improvement.
Hard disagree on that one. People aren't comfortable with the syntax because it's rather awkward and the subset of portable syntax is minuscule. The syntax itself has all sorts of variations (e.g. bash 3 vs 5 vs not-bash) and when you have to leverage external programs (e.g. sed or awk) things can get hairy quickly (especially if you're trying to support both BSD and GNU userland components).
From a security POV I'd much rather use not-sh style stuff. Error handling is much more accessible and robust in higher level scripting languages. Sure you can run "set -e" and hope for the best... until you have to use something that tries to stuff meaning into the exit code like grep. Handling filenames with spaces is significantly easier once you leave the sh baggage behind.
As for ls(1), are you talking about BSD or GNU ls?
Yeah which is the same reason why sql is avoided, regex is avoided, javascript is avoided etc. People will go through enormous trouble to avoid learning anything other than the multi threaded c-style programming model from the 1970s, that they learned in their computer science degree.
Yep, most folks learn programming in high level languages, lower level and “glue” scripts feels “not right”. I can attest to this… after using one shitty orm after another, learning sql and bash made me a much better programmer.
It has changed in the past and it will change in the future.
"ls" will always list files in some way, but it is the only thing you can take for granted. You don't really know what kind of decoration you will get, how spaces and non-ascii characters will appear, the language, the order, etc... You may also find yourself with a weird environment variable your user has set that makes ls look better in his interactive shell and breaks your script.
Generally, I avoid "ls" in scripts for that reason, but many commands have the same kind of problem. Your GNU script may break on busybox or BSD (happened to me many times). And good luck knowing what /bin/sh means, it may be bash, it is supposed to be POSIX, that's all we know.
0: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V...
But it is a minefield, it is way to easy to sneak-in a bashism because you are so sure it is POSIX that you didn't check, made worse by the fact that if you are writing shell scripts, you are probably using an much more liberal interactive shell all day long. And your script will work and no one will complain until you try to run it on another system.
It is not entirely unlike client-side web technologies, where you have to check for browser compatibility. For the web, it is a recognized problem and there are plenty of solutions to deal with it: jquery, babel, etc... I am not aware of such tools for shell scripts, or maybe just analyzers that check for posix compliance.
Many of my scripts for personal use were basically just cleaned up and slightly extended versions of the command sequences I typed in anyway. This is different from writing a script in a proper language where you usually start from scratch.
I believe, this was Perl's approach originally: Let you use your shell commands as a starting point and then gradually extend the code in a safer language. However, this still leaves you with many of bash's syntax issues.
On the opposite side, many modern languages do have REPLs - but I have not seen anyone so far use them as actual day-to-day shells.
So I think to really get rid of shell scripting, you'd have to develop some syntax that would work well both for scripts and for shell use.
Unix style shells start simple and get as complex as desired, rather than a set of typed objects which must all be kept in memory at the same time.
I think what PowerShell showed me is the next step above something the complexity of a Unix shell is a "real programming language", no matter if it's interpreted live (scripting) or compiled. There might be reason to revisit the exact syntax, but I'll wait to hold my breath until we've replaced keyboard layouts designed to slow down typing to a point which doesn't break a mechanical typewriter.
Windows apps that have PowerShell cmdlets are a boon also to administration also and I'm grateful when the option is there.
My biggest gripe with PowerShell though is that as soon as you get past the "I don't want to click around windows/.net apps", Powershell shows how shallow it is and its quirks affect performance, design, etc. The PowerShell community encourages a lot of practices that end up trapping you in bad design which makes maintenancen and readability of scripts a challenge. (the handling of variables and their scope can be confusing as PowerShell allows for a lot of "cheating" on which data is accessible and the script will work, but it's very hard to tell sometimes the state of the data being worked on and where it came from. Use of += for adding to arrays is fastest to write at first blush, but you end up with awful performance because it copies the array, but your average PowerShell script writer wouldn't know it and ends up wondering why your script struggles outside of simple lab scale tests)
The answer to this is getting into reflection and using .net calls, but to get the right performance and behavior from PowerShell you end up with a good portion of the script being just straight up .net code and at that point I wonder why bother with powershell at all and not just use .net and some faster API to get data from the app or going through wmi to work with windows instead.
You can get complex PowerShell programs for sure. But they're PowerShell only in name, the force powering the program is usually something else that PowerShell only hinders as you need to now deal with PowerShell-isms to make it execute.
So in a sense it's very successful in that I think this is exactly what PowerShell was supposed to do. But I find as I got past short scripts that PowerShell just can't really do anything it does better than another language would. It can't work anywhere near as fast as bash or any other *nix shell can, you can get a lot of depth into windows but mostly through .net calls, and for applications, a good rest API returns the same data, but faster and in a format more useful for other applications. (both program wise and just in terms of concept)
It feels like PowerShelk was (is?) supposed to be a gateway drug for .net. But it's just too easy to get hooked on it and suffer the side effects and never really take that next step to a more mature language, so you get less of high from using it but way more nasty side effects.
(Also the less said about PowerShell ISE the better. I wish MS would just remove it entirely as its difference in behavior from the actual shell wastes so much time in debugging scripts)
- Elvish (non-POSIX)
- Oil (Bash superset)
- murex (non-POSIX)
Oil is probably the most interesting on this list because it aims to be 100% Bash compatible while supporting additional syntax to make it less ugly. Thus an upgrade path from Bash.
Elvish and Murex are similar in they they’re typed shells (like Powershell but less ugly and works with existing Linux / UNIX CLI tools) but have slightly different approaches in solving that similar domain.
Like xonsh, but even less bash-y.
I think we tend to underestimate the importance of this kind of uniformity (aka homogeneity, consistency, sameness, equivalence etc) that allows us to freely shift back and forth and work across different environments (aka contexts, mindsets, etc) without any essential changes and translations. i.e. in a sense, minimizing the boundary such that the difference effectively disappears and it feels as if it's all just one same. In this case the environments being shell <-> script. But I think I see similar patterns in many other places, including outside of tech.
So, similar to as you also mentioned, as long as there is no language that beats the current shell as both a shell language and a scripting language, and one that sufficiently matures for real practical uses as both, I have a hard time imagining the most perfect scripting language alone still being able to make shell scripts go away. At least, I have a hard time imagining myself abandoning shell scripts otherwise.
- I often encounter machines without Python or Go. I never encounter machines without Bash.
- Python and Go will probably change faster than Bash. So the maintenance cost of scripts in Python or Go will be higher.
- When you figured out something tricky on the command line, you can paste it into a Bash script verbatim.
- Bash is the perfect glue language. If the task at hand is to glue together some commands, nothing beats Bash.
Backward compatibility is king.
For most of our big scripts at work, I've actually resorted to just running them inside a docker container, mounting my current folder. So tired of trying to get my env to match someone else's.
Python3 was a big stain because of those issues.
Python 4 could literally just be a platform shift to sort out paths, libraries, dependencies, documentation in order to try to solve all those problems.
It's just amazing how nutty things can be for tech we all depend on.
We would have never designed something that way on purpose.
Bash - as a non-expert I also feel is a bit wierd.
I almost what Computer Science God to come along and make a shell script, a clean light language like python/js, maybe something like Swift/Java, and then something like Rust but a bit easier - and they all feel a bit similar and integrate and the build tooling is fairly clean and the whole world can get along.
So true. You have to be very careful when writing python, because cool language idioms that you use (or function calls), may not have been introduced in say for example Python 3.5. So when someone runs it with a different interpreter it will not work.
Writing python scripts that not-python-knowing people could use, is an art.
Instead, I work for a company who's core was written in bash and perl, and those scripts still work today after 15 years.
That or I don’t have interesting enough dependencies :)
Package management and a virtual env manager in one.
What is the dependency management solution for bash that it so much better than anything in Python?
> So the maintenance cost of scripts in Python or Go will be higher
For Go, yes. For Python it's very rare for an old codebase to break unless you are talking about Python2/3
This is something I'd like to see addressed by programming languages. Wouldn't it be nice if you could say "pythonversion 2.7" at the top of your program, and it would use that version of the language for the remainder of the file?
It allows you to use shell stuff in a python script. Basically it allows you to use the best of both worlds, i. e. you can use python for the logic flow while retaining all your quick shell oneliners when neede. And it makes this interaction of python and shell super convenient, i. e. run your oneliner but process its output with python. Much faster than coding everything in python or in shell.
[1]: https://xon.sh/
It was enjoyable to use for a project recently.
What I see as a benefit is that it allows you to reuse your existing python and bash knowledge without having to learn (almost[1]) anything new.
[1]: Xonsh does add one or two new things to python. But they are pretty straightforward and aren't required. But are hella nice for more advanced interplay between shell and python.
Writing logic - functions, for loops, or if statements- is still very annoying in Bash. Likewise writing a piece of code that glues calls to external commands is extremely clunky in almost all other languages (except Perl an Ruby), notably Python which is often treated as a Bash replacement.
Finding the right balance can be tricky, but choosing only Shell or only Python is not the right way to go.
Piping structured objects between apps without having to parse output is amazing. I hope Linux will come up with something similar. Of course PowerShell is open source but I don't trust MS enough to adopt it privately. I do use it at work though which makes sense as we're a MS shop.
There are much more shell users than Go, Python, ... programmers. Also those who happen to be programmers normally use certain languages, typically more than one, typically far more diverse than just most used shells, for a certain kind of activity, like programming in a team a not-that-small project. As a result their mindset is not "quickly do something/quickly automate something for me, myself and I only, on my own desktop" but to be part of a larger game with need to cooperate, produce code easy to read by others, create a set of entry points for integration etc. Those programmers are not all, but mostly also shell users and they tend to use the shell not just for specific large projects but also for quick, simple&dirty personal stuff. So as non-programmers who use the shell they know and use it daily.
In the past we have had perhaps better tools, like Smalltalk on Xerox desktops or Emacs/zmacs lisp on LispM BUT their actual public is next to 0 and modern Emacs-ers are on modern desktops not LispM and they are typically born on modern desktop so they know the shell and in Emacs they'll wrap anyway countless shell-centric tools so end the end a script is often quicker and simpler...
N=10
find /sys/fs/cgroup/ -name memory.current | while read f; do
test -e $f || continue
echo `cat $f` $f
done | sort -n | tail -n $N | awk '
match($2, /system.slice.([a-z-]*).service/, arr) {
$2 = "/system/" arr[1];
}
match($2, /user.slice.user-([0-9]*).slice/, arr) {
cmd = "id -un " arr[1];
cmd | getline uid;
close(cmd);
$2 = "/u/" uid;
}
match($2, /cgroup.([a-z]*).slice/, arr) {
$2 = "/system/" arr[1];
}
{ print }' ps -e -o size,user,comm | sort -n | tailThe reality is that most Unix DSLs are similar to subsets of Python. Yes, Python is not the global optimum of programming languages forever more amen, but most Unix DSLs are effectively just ad-hoc procedural languages, and therefore special cases of Python. And if you're going to use Python aggressively for other things, then there's a benefit to using it as a shell as well, and thereby get some consistency.
Actually, I'm not sure I've ever done a loop in bash at all.
The few times I've tried to write even trivial programs in bash resulted in me eventually encountering something ugly and unpleasant, then quitting and starting over in Python, and having a working solution extremely fast.
I've started using that instead of bash script and js/node/deno scripts. it's kind of a nice combination of things.
the only pain with this is writing to files. > doesn't just work.
posix shells with `set -euf`: I give one care.
perl, erlang: I care about arity.
python: I give one care about data type consistency.
java: I care more about types than command line interoperability.
ruby: I care about some things but not speed or memory or scalar/vector predictability.
lua: I care about everything but indexing.
c/c++: I care too much.
go: I care about concurrency.
rust: I care about making fun of Stroustrup.
haskell: I care about... whatever. I'll find out some day when I feel more eager.
javascript: Burn the world! Burn it to the ground!
bf: I care about obfuscation.
altJVM/altJS: I care about obfuscation--with data types.
.NET: I care about taking a paycheck.
cobol: I care about indentation even more than Python does.
random DSL's: I care about doing it all over again.
d: I care so much but nobody ever listens to me.
PowerShell: I care about COMSPEC environments.
non-POSIX shell: I care about the good old days.
assembler: No you don't.
HDL's: Yes you do.
php: I care about pre-Rails projects.
lisp: I cdar care less.
regex: You'll care when I break prod in a week.
why?
- old saying
Seriously, though, treating shell script as a prototyping tool can be really powerful. Don't be afraid to rewrite it again as a 'proper' program if you have concerns about maintainability, etc.