What was the point of [ “x$var” = “xval” ]? (2021)
vidarholen.net
vidarholen.net
One key point there is that it’s still easy to trigger cases where you need this hack, because the cleverness is only stable for up to four arguments; so and, or, and parentheses all exceed the point of safety.
So x serves as a form of quoting. If I were to potentially invoke a binary to evaluate an expression, I would quote any string-like variable (using the x hack) and I would do my best to avoid numeric comparisons.
(Or use built-in shell features where the shell knows which values are quoted.)
Regular programs can make that distinction with no problem, the mistake is thinking that the outer quotes are part of the argument. In the context of the POSIX style shell, they're not, they signify that what is enclosed is one argument. So you'll have to use one of the various mechanisms the shell provides to indicate that you want to pass the quotes as an argument.
$ echo "(" '"("'
( "("
"(" the argument is (
'"("' the argument is "("
When you inspect argv in your program, you will see the argument including the quotes in the second example.I thought you were saying that the statement 'In both cases, they get the string “(“ in argv.' wasn't necessarily true. My apologies!
> you want to pass the quotes as an argument.
Yes, that's the point - in the context of /bin/[, 'x' is the quote; that's how you quote string arguments to [. (Yes, it's ad hoc as hell; that's fairly par for the course for unix.)
In shell the preferred form for "and" and "or" is outside the test, and you always quote anything with a variable reference, like this:
if [ "$x" = "foo" ] && [ -z "$empty" ] ; then
echo "Worked"
fi
I think the "&&" is easier to read, as well as completely eliminating the problem.shell has "case" as part of the language, and case is way better at string evaluation that the "test" built in
case "$x" in (foo) echo Worked ;; esac
"test" or it's other name "[" is a program or a built in. It takes a parameter to tell it what to do. It can evaluate strings, but it can also look at the file system or do a little numeric evaluation. Newer versions in gnuland may even be able to tell you if a unicode glyph is left to right or right to left. It does all this using - or -- prefixed options.
test -f /dev/sda will fail (it's a block device not a file)
test -f /dev will fail (it's a directory)
test -h /usr/bin/[ will succeed on some unices (it's a symbolic link to /usr/bin/test)
So -- what does [ "$x" = foo ] (or, more clearly, test "$x" = foo) do if the variable x happens to have "-f" in it?
The shell will interpret the " and the $x and hand to the builtin this:
[ -x = foo ] (or test -x = foo)
Now, the human is waving their hands "that's not what I meant no fair!" but a contract's a contract -- test might look around for a directory entry in the current directory and not find a file named =, or hey, it might, because = is a fair directory entry name. It may then poke at that associated inode and find that = is in fact a file, and return a 0 in the exit status. Or it may say "hey, file existence testing has one paramater and you gave me a filename, =, and another file, foo, what gives?
Regardless -- you've got no control over the contents of variable x; that's why you're asking test to look into it for you. You can blorp a thing in front of it to protect yourself from this case, but that's just silly. Test can't do any regular expressions and it does all sorts of other things, so it's taking parameters, it's just a mess. Just use case $x in ... esac.
Test is great for other stuff, but string evaluation? It wasn't ever good for it.
Anyone running a shell that still requires the x-hack is deliberately maintaining a legacy system, at which point the onus is on them to get modern tooling working on their machine.
Windows 95 is still very much alive.
Embedded-like systems are a thing.
Did not say it has to be done.
Yes. If there is one thing I wish I could drill into every programmer's head, it's that not breaking stuff for your users is the #1 most important thing. There is one of you, and there may be many thousands of users. You putting in the extra effort means saving thousands of times of effort from your downstream users. Sometimes breakage unavoidable, but those instances should be rare, extremely well communicated, and felt as a failure to be learned from and avoided in the future. I don't care how ugly it makes things for you, your users are paramount.
I would say that's #2. #1 is to never cause data loss.
Can you share any examples/use-cases? I'm not doubting you, I'm just very, very curious.
In 2011 a customer I worked with was using Windows 95 still in a test bed for their legacy equipment. The test software they had was written for it, and it still worked, so no need to go through the massive pain of updating. But even that was 12 years ago and we expected it to fall apart any day...
The alternative is to bundle all the dependencies, including the shell interpreter into some maximally portable setup.
I'm not saying that everyone needs to keep legacy support in mind. I'm saying that just taking the attitude of "it's old, too bad" is not good blanket practice. It all depends on what you're doing.
> it's not particularly callous to want to feed your family
The callous part is taking a disdainful attitude toward those customers.
I guess I thought it was obvious that most of the the time someone doesn't want to support legacy it's for this reason, not because of disdain for them.
Your disdain of legacy clients might be profitable or understandable or even morally just, but it's still disdain.
OED. You can respect someone but know their business isn’t worth your effort.
Callous is callous regardless of your motive. It might be justifiable to be callous if that's what your business circumstances demand, but that doesn't change the effect on those customers.
Jesus Christ, you just put x in front of the variable!
Nonetheless, having read the fine article, I will stop using he x construction, which I find myself typing from time to time just out of habit.
There are other things like that out there. I think I recall seeing something on HN fairly recently abut a controller for some CNC machine that was running an early version of MSDOS. Again, replacing with something more modern was very expensive. These things exist and the users aren't necessarily being wilfully malign in not upgrading
I won't support Windows 95 for a pittance
This is not true at all. There are plenty of such systems, run by major companies, in daily use. It isn't the most common case, but it certainly matters in important parts of the real world.
But like I said, I'm sympathetic, so I'd love to hear a case for why I should do it anyway
Because our product has to run on a wide variety of hardware, it was essential to maintain compatibility both for the crufty old shell implementations and the shiny new -- so best practice there was to develop the habit of writing shell code that could run across the eons, including the x-hack. Multimillion dollar contracts depended on it.
If your situation isn't like that, then I understand why this is a nonissue. But there are plenty of situations where such backward compatibility is essential. Not the most common thing, perhaps, but not as rare as some think.
I do think there is a tendency amongst some programmers to think that everyone is running stuff that is at least moderately modern. In the real world, this is simply not true.
"It probably doesn't matter, but you will build habits and at some point you may wish you already had the habit in place"
?
If so, that's a reasonable point, thank you
A much smarter thing is for the script to detect, at the top, that it's running on a crappy /bin/sh, find the better shell, and re-execute itself with that shell (setting some environment variable to prevent recursion).
Then the rest of the script doesn't have to use unusually stupefying coding techniques. Just the regular stupefying techniques.
Scripts and programs will always get run in unsupported (broken) environments. If your program can't exit with a clear error in such a situation, it's not very portable.
Personally I tend to use the multi-line nix-shell shebang:
#! /usr/bin/env nix-shell #! nix-shell -i SHELL -p DEPENDENCIES #! nix-shell -I nixpkgs=NIXPKGS_TARBALL_URL
If you want `bash` as your shell, then `bash` goes in the SHELL location and as one of the DEPENDENCIES. If you want GNU coreutils that goes in DEPENDENCIES. Etc. The last line is optional, it pins the versions of all the dependencies you want. This also works for Python scripts.
$ foo=""
$ [ $foo = "" ]
-sh: [: =: unary operator expected
$ [ x$foo = "x" ]
$ echo $?
0
I'm not arguing that this is a good idea, merely observing that this case remains.Empty strings had always been the case that I'd been taught required the x-hack.
[[ $var == "" ]]
works as you describe and as expected. What the parent is talking about is when you have an invocation like my_program $opt
my_program "$opt"
Which when empty will evaluate to my_program
my_program ""
respectively. % bash -c 'a=xy; b="x*"; [[ $a == $b ]] && echo "$a == $b"'
xy == x*
My hot take is that you should avoid using [[, preferring to use [, which at least pretends to work like a normal program that can be invoked using normal quoting rules—it just interprets a funky mini-language on its command-line arguments. [[ is some exotic thing that seems reasonable at first glance but follows no rules but its own and varies its behavior depending on syntactic nits that literally cannot possibly matter and aren't even distinguishable for normal programs.And then just quote consistently, which is something you need know how to do with almost every other line of a shell script anyway!
I tend to agree, and have taken this one step further, preferring the "test" command:
if test "$a" = "foo" ; then
I do this for a few reasons:- "Subsetting" your language can be beneficial: https://wiki.c2.com/?LanguageSubset
- "test" reinforces the notion that this is just the name of a program, not a syntax construct, which encourages using other programs as predicates, e.g. 'if which -s foo'.
- "man test" doesn't force you to wade through a 5,000 line man page.
Recent example: https://leopard.sh/leopard.sh
So [[ $a == "$b" ]] works.
$ shellcheck oops.sh
In oops.sh line 4:
[[ $a == $b ]]
^-- SC2053 (warning): Quote the right-hand side of == in [[ ]] to prevent glob matching.
For more information:
https://www.shellcheck.net/wiki/SC2053 -- Quote the right-hand side of == i...
Unlike the bugs that the x$var works around this is all in the manual!https://www.gnu.org/software/bash/manual/bash.html#index-_00...
There are lots of ways to hack around lots of ways to break code.
It's probably the biggest drawback to self-learning.
This advice is highly dependent on specific choice of shell. This kind of quoting is unnecessary and discouraged in zsh, for example.
If you use nonstandard mechanisms, then whatever.
When I started off in the UNIX world, the shop had a collection of AIX, SunOS, Solaris, HP/UX and Digital UNIX. the x-hack was standard form and I spent a few years fixing errors with the x-hack (or converting the scripts to Perl).
It is just second nature to use the x-hack in if's and case statements on this point. I have to try really hard to NOT use it. Of course, the folks with less gray in their hair and often confused by it, referring to it a "Olde Unix."
I have approved plenty of PRs with the comment, "Updating mek's Olde Unix syntax."
If it wasn't abundantly clear, this is not an issue with any of the shells today.
The least-common-denominator aka /bin/sh and its syntax will stick around for decades to come. You can't rely on any of the "modern" shells to be around, particularly if you write scripts that target a multitude of unix-like systems (macOS, xBSD, Linux).
If you omit those, a fairly intuitive unambiguous parse is very easy: if we’re [, check last argument is ] and throw it away; if one argument, return whether it’s empty; otherwise, if two, the first is a unary operator, evaluate and return that; otherwise, if three, the second is a binary operator, evaluate and return that; otherwise fail. That’s it.
I really doubt it. If the comparison parameters are undistinguishable from the operators, it's only a matter of enough creativity and people will to find new ways to break it.
* which I abhor, by the way—just like GNU abhors man pages—but that doesn't make your argument any less absurd
imagine someone pitching shell programming today. the response would be: LOL so you mean each language will have a common subset of syntax that supposdely works the same across all of them? how can anyone ever program that way?
* ironically, a big part of why embedded products are so flaky is just because of their shell scripts
NB. Almquist shell does not necessarily use readline. Neither NetBSD's sh nor Linux's dash use it. (Not to imply editline isn't bloated, too, but it's not nearly as bad.)
Further, I did not mean "the terminal". I had to look up "LARP". FWIW, I do not play video games. I do not use readline. I do not use a "terminal emulator". I use textmode. These choices are not because I think highly of ash, editline, textmode, non-graphical computing or minimalism in general. They are because I find the alternatives (large scripting languages, readline, video games, terminal emulators) are really annoying.
Why not share some of the programs H4ZB7 has written, so we can see how to do things the right way.
In theory, the shell could be eliminated and people could boot into something else. Onine commenters love to bring this idea up. But in practice, this is rarely implemented. The UNIX shell is the overwhelmingly choice. This is not an "argument". It's a fact. The shell and shell scripting are everywhere.
The "proper scripting language", OTOH, is a highly opinionated concept. It will keep changing. A never-ending popularity contest.
Shell scripting is here to stay whether language wonks like it or not.
PowerShell is just another Perl / Python / Ruby. Except made by Microsoft, so it's got a ton of useless features, lacking some essential stuff and has bombastic syntax. It's not suited for the role of UNIX Shell. Not by a long shot.
----
On a larger note, the whole idea of needing a shell is bad. It comes from a defective (or rather overly simplistic and un-insightful) design of UNIX. A better system wouldn't need a distinction between a system programming language and a shell. You would just use one and the same thing for both purposes.
If you are on Linux, Guile is supposed to be that, but... maybe it was supposed to be that?.. I'd still want it to take the role of Shell, even though it doesn't tick all my boxes.
Or worse, the output is stable (because peopleare scraping it with aws and sed) and reflects things as how they existed 20 years ago versus now. For example `ifconfig` doesn't really match how Linux does networking.
BTW, I don't know of any distro wit LTS releases that didn't yet expire that doesn't ship with iproute2. In my opinion, iproute2 has done a very good job, given the circumstances to organize and systematize information about networking. So, if you don't like ifconfig, today you most likely have a better alternative.
people don’t update tools because it breaks scraping.
If fact if they did ip addr wouldn’t exist: we’d be able to update ifconfig.
What people are you talking about?
Anyone who's on a payroll is liable if they don't upgrade systems when they go out of support. It doesn't happen because of someone's whims... If some individual user chooses to stick with an outdated system: it's on them, and they have no ground to stand on if they complain about it...
> we’d be able to update ifconfig.
No we wouldn't. "ip a" is a tiny fraction of what this command does. The reason it came to be was not that ifconfig needed replacement, but because Linux utilities are a zoo of halfbaked projects, most of which have huge flaws in their design, especially when it comes to extensibility.
Did ifconfig foresee bridge interfaces, bond interfaces, vLANs etc? What about IB? What about sr-iov etc.? -- Nope, and since there wasn't a single structured and systematic tool, that could've been extended by the authors of eg. bridge drivers, they were forced to come up with their own tools, eg. brctl. iproute2 wasn't an improvement on top of ifconfig. It was an aggregation and systematization of various tools and configurations that dealt in a very partial way with various aspects of networking.
There would've been no way to keep any kind of backwards compatibility with ifconfig, so there was no reason to keep the name. It's not an evolution of ifconfig, it's an entirely different thing.
People that maintain software.
> Anyone who's on a payroll is liable if they don't upgrade systems when they go out of support.
Yes. Let's make that process easier.
> Did ifconfig foresee bridge interfaces, bond interfaces, vLANs etc?
No, that's precisely why I wrote the comment you're replying to. Yet people still use ifconfig, and wonder why the results don't make sense with their vlanned bonded interface.
> It's not an evolution of ifconfig, it's an entirely different thing.
Yes, agreed completely.
But if we got away from scraping we could have added them a lot more gently.
> There would've been no way to keep any kind of backwards compatibility with ifconfig
Disagree. We can come up with a design here but I don't feel like you're actually responding to anything I write so I don't wish to do so.
> there was no reason to keep the name
To this day people are still scraping ifconfig and wondering why the results are bad.
I don't believe I've ever parsed output from ls. Definitely not with sed or awk. I don't think that any Linux admin worth their salt would do that.
Needless to mention that neither ls nor sed nor awk have anything to do with Shell.
> What essential stuff is missing?
Since Microsoft tried to replace not just the shell, but the entire set of utilities with PowerShell, they forgot about xargs, for example.
But, if we are talking about the shell proper, then things that are missing would be a serialization format that would allow one to pretend that remote shell sends "objects" rather than strings.
Most importantly, however, PowerShell is an attempt to fix the problem at the wrong level. The reality of the OS it's running on is that processes take command line and environment variables, and they don't take objects. They return exit codes and two or more output streams, they don't return objects. PowerShell is trying to pretend that communication between processes happens in objects, but that's a lie, and it shows. In any non-trivial use of such a shell, you will have to escape the objects, and face the reality of what's happening underneath.
That's why earlier in the days, when objects were a novelty programmers thought that a programming language that is object-oriented from the bottom up is preferable to the one which implements objects as a library.
PowerShell is even worse in this sense than Perl or Common Lisp. Objects in PowerShell aren't an afterthought. The authors wanted them from the start, but couldn't have them. Unfortunately, we live in the world where most popular projects we use today were still-born in engineering terms (eg. Unix, and later Linux, any Fortran-like language created since 80's etc.) So, it may as well happen that PowerShell will succeed in terms of popularity, but it's broken by design unless we completely replace the whole way we work with processes... and I don't see it happening unless there's some cataclysmic event that wipes most of us out.
You can just use foreach with an array or a multi-line string:
$newFiles | foreach { git add $_ }
> Most importantly, however, PowerShell is an attempt to fix the problem at the wrong level. The reality of the OS it's running on is that processes take command line and environment variables, and they don't take objects. They return exit codes and two or more output streams, they don't return objects. PowerShell is trying to pretend that communication between processes happens in objects, but that's a lie, and it shows. In any non-trivial use of such a shell, you will have to escape the objects, and face the reality of what's happening underneath.
If you’re executing processes, then yes, things might get hairy. But if you’re using PowerShell cmdlets (which either do the thing natively, or call a subprocess and parse it to an object), the object-oriented stuff is really convenient and powerful.
Here’s a random example of fancy PowerShell things. How can you do this (filter by a keyword, group by parent directory, and sort the results by count) with UNIX tools? What magic combination of ls, find, sed, grep, awk, sort, uniq would solve this?
Get-ChildItem -Recurse C:\Stuff | Where-Object {$_.FullName -notmatch "xyz"} | Group-Object Directory | Sort-Object -Descending Count
That's not what xargs does though...
> If you’re executing processes,
This is why shells exists. They don't exist to execute cmdlets. Nobody cares about that. Cmdlets are about internal functionality of the shell. Executing processes is the service shell provides to it's users.
> What magic ...?
I don't know what that code does. I don't have Windows, and wouldn't install Microsoft's products on my own or company's computer... But, my guess would be that it would take about half the effort that it took you to write this.
What magic to you is not magic to me, because I know that stuff. You must know PowerShell better than I do. I simply know enough not to use it. I don't care about the gory details.
I have a fair amount of shell scripts in CI pipelines and things usually break when we hit unhandled edge cases which is exactly what I would expect.
That said, if your script is for a single platform you won't need any of these hacks which only exist for compatibility with different/older shells. My scripts won't even run on bash 3 and I accept that.
While shells can and do have security issues, much of that is historical and given the smaller surface area they just have fewer security issues than say, Node. I don't need a shell version manager à la NVM, virtualenv, or rbenv. I don't need to worry about a massive dependency graph. A shell of some sort comes pre-installed with every *nix system I've ever used so I don't need to worry about figuring out how to install a new runtime in an unprivileged environment.
I'm sure some people love Bash and love shell scripting. For many of us it's just a matter of practicality.
OIL and ABS are close but who uses those.
I used tcl 15 years ago and it was pretty great, but the SDKs were already getting old and many falling unmaintained. I worried about the future of tcl
`ls | where createtime < now - hour`
That's pseudocode, I don't use pwsh every day.
pwsh never really took off as it was good tech during bad leadership (Steve Ballmer), but I hope someone (maybe Rust coreutils people) steal the approach.
Regarding you last sentence: yes that’s why I mentioned it didn’t take off in the comment you replied to.
But long term scraping will die.
But well, most languages support piping program together. You will need a 5 to 10 lines wrapper on the most used ones, but it's always well supported.
And as a bonus, you get to explicitly say how the pipe should behave, instead of searching the docs to learn how to handle failure or early stream termination.
What admin scripting languages bring to the table:
1) Easy access to executable binaries - no creating a Process object to describe it, just run it
2) Easy piping and filtering
3) A standard library designed for terse and easy-to-use access to administrative tasks.
4) Pleasant in REPL form.
That's the good side. The bad side is that they're all such a collection of hideous hacks and edge-cases for basic concepts like variable scope and typing that even the simplest operations that would be hilariously trivial in a real language are utterly agonizing.
$ str="-e"
$ [ \( ! "$str" \) ]
bash: [: `)' expected, found ] # bash
> POSIX fixes all these ambiguities for up to 4 parametersSo... It's not fixed, and the hack is still needed?
Oldest bash after that I have available is cygwin bash 4.4.12 (2017-10-13) which doesn't have the issue.
Just install a newer bash using ports or homebrew and bob is your shell uncle.
There it arguably makes sense to be so robust against old/broken shells, since the whole point is to successfully configure the build on a broad spectrum of environments.
worse than brainfuck? I find that hard to believe...
on a serious note though, I don't agree. the syntax is not that different from the awk -> perl -> ruby heritage. Just a few constraints/conventions due to whitespace significance that other langs don't have.
Malbolge perhaps. Brainfuck is tedious, but it's a language of 8 symbols; there's nothing to forget. But doing comparisons in bash calls for a whole quickref of its own.
Was it just single [ (test) or [[ as well?
if test x$var = val ; ...
instead of using the [ command?Some years ago, I checked: the [ alias for test was already in the Bourne shell in 1979.
Assume the year is 1999. What real portability problem do you solve by avoiding [ in favor of test?
Real meaning that the rest of the script, and the software stack the script is related to, runs on a system where [ wouldn't work.
- it looks like every other thing in sh (`[` is the outlier)
- it prevents confusion like putting `&&` inside the `[ ... ]` brackets (understandable that `[` works that way if you think about it, but it’s another way it’s an outlier)
- mildly discourages use of one particular bash-ism, `[[`, as it works slightly differently, and I prefer POSIX sh unless there is a good justification
- it doesn’t require a closing `]`
I don’t use it for what I think you mean by portability, that is, because `test` exists whereas `[` does not, or anything like that.
>>> foo = b'foo'
>>> foo[0] == b'f'
False
Don't get me wrong. I love Python, it has been my daily driver for 20+ years (and it pays my salary), but it does have its own set of cognitive burdens.Also, shell scripts I wrote in the 90:s work today. Python scripts - not so much.
"Want" isn't always the whole equation. Sysadmins are often called to support ancient OSes, in particular in environments which keep systems in operation long after their "best before" dates (e.g. internal IT in banking systems).
There's literally no disadvantage, beyond aesthetics, to _not_ using the x-hack. For many of us older folks, using it is built into our muscle memory.
bashisms are forbidden in every corner of the world, sorry
https://askubuntu.com/questions/1277922/what-are-syntax-diff...
I'll also note that BusyBox supports [[.
There is, however, Inspur K-UX, a Linux distro which is a Unix.