Posix.1-2024 is published
ieeexplore.ieee.org
ieeexplore.ieee.org
As a specific example, the seemingly simple matter of when the shell decides to split a string based on $IFS and when it does not were quite confusing to me until I went through the specification here: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V...
For example, if
a="foo bar"
then ls $a
will split the value into two fields (thus two arguments to ls). Of course we should surround $a with double-quotes to avoid the field splitting. However the following is fine: case $a in
No field splitting occurs here. However, to be kind to your code reviewer, you might want to double-quote this anyway for the sake of simplicity and consistency. Behaviour like this is specified in sections "Field Splitting" and "Case Conditional Construct" of the aforementioned link. Specification documents like this were formative in in my journey toward learning to write shell scripts confidently.> June 14, 2024: IEEE Std 1003.1-2024 has been published by IEEE. The Open Group Base Specifications, Issue 8 has been published by The Open Group. At this stage only PDF is available. The HTML edition to follow soon.
This kind of bullshit is how I made a career rewriting people's buggy shell scripts in Python
I also rewrite my stuff in Python as soon as it becomes nontrivial.
x=$a # not split! It means the same thing as x="$a"
They were taught that you have to quote everything, which is a reasonable rule to follow, but it's not true.---
I never wrote about this on the Oils blog (https://www.oilshell.org/ ), but the post would be titled:
Shell Has Context Sensitive Evaluation
Basically the two contexts you should think of are:
(1) EVAL WORD SEQUENCE
This occurs in 2 places in POSIX shell:
ls $x$y # simple command is a sequence of words
for i in $x$y; do echo $i; done # for loop
And 1 place in bash: a=( $x$y ) # array literal
In these cases, the shell "wants" a sequence of strings, not a single one. So it does splitting.---
(2) EVAL WORD TO STRING
But there are many other contexts where the shell does not "want" a sequence of strings.
It wants a SINGLE string. And conversely, it actually JOINS arrays of strings, rather than splitting.
Usually "$@" is an array / sequence of strings, while $@ or $* is a string, roughly speaking.
But the shell doesn't want sequences of strings in MANY cases, e.g.
a=$@ # I only want 1 string here, so I JOIN rather than splitting
echo hi > "$@" # redirect arg (not all shells agree though!)
case "$@" in ... esac # as you point out
So the bottom line is that variables aren't really strings OR arrays of strings. Whatever the shell wants, it converts it to.And shells also DISAGREE on the specifics of those rules. POSIX shell has the array "$@", but arrays in general are not in POSIX.
---
And even worse, think about this case:
local x=$a
Does it behave like an assignment, which wants a single string?Or does it behave like a simple command, which wants a sequence?
You can look at it both ways. The bottom line is that assignment builtins are special and they don't follow the normal rules of simple commands. Shells have differed, but POSIX decided on this awhile ago.
---
This is all of course mind numbing trivia that has no real reason for existing ... YSH fixes it, and it's now pure native C++, no more Python.
YSH Doesn't Require Quoting Everywhere - https://www.oilshell.org/blog/2021/04/simple-word-eval.html (Oil was renamed to YSH since this blog post was written)
Simple Word Evaluation in Unix Shell - https://www.oilshell.org/release/latest/doc/simple-word-eval...
In YSH you can tell just by looking it's a single string or an array.
ls $a # identical to ls "$a"
ls @myarray # splice an array
It never "molests" your variables. There's no auto-conversion, and you can upgrade to those rules with shopt --set ysh:upgrade $ readarray -d '/' <<<${PWD:1:-1}
$ echo ${MAPFILE[@]}
and you get a nice list of folders you can push/pop as you wish...In general YSH is pretty different than zsh though -- it's more of a Python- JS-like language with structured data, e.g.
ysh$ var a = ['list', 'of' strings']
ysh$ write -- @a
list
of
strings
In zsh I still think that's the pretty obscure "${a[@]}" rather than @a.Arrays are also "flat" in zsh -- you can't have an array of arrays, because there's no garbage collector. But YSH has arbitrarily nested JSON-like data structures, and JSON serialization built in.
I need to put some code examples on the home page, but for now - https://www.oilshell.org/release/latest/doc/ysh-tour.html
While incompatible, the zsh behaviour makes a lot more sense.
You can use "setopt sh_word_split" to get the POSIX behaviour.
Or split explicitly with ${(s: :)a}, ${(s:SPLIT-ON-THIS:)a}, etc. (not compatible with anything but zsh).
set -euxo pipefail
The "u" has basically the same effect as the question mark, but for every variable usage.I am not sure how it could break anything though, unless you are parsing stderr of your script in a subsequent step, which would seem unusual anyway.
Maybe I can start using it again. (I think I noticed that issue while I was at Google, and they used an older version of Bash.)
* readlink/realpath (https://austingroupbugs.net/view.php?id=1457)
* find -print0, xargs -0 and read -d (https://austingroupbugs.net/view.php?id=243)
* find -iname (https://austingroupbugs.net/view.php?id=1031
* sed -E (https://austingroupbugs.net/view.php?id=528)
* set -o pipefail (https://austingroupbugs.net/view.php?id=789)
Finally! It always seemed very strange to me that posix said that shared objects were a thing and provided a rtld API for using them, but never specified how to create them.
The sed -E option makes it easy to portably use extended regular expressions.
The find -print0, xargs -0, and read -d provide portable ways to securely process lists of files. They were already widely implemented, but now they're officially part of the spec and can be counted on being present in many other places.
Thanks for the improvements on POSIX, I've read many issues and discussions raised by you in the past couple of years.
If fact, I think it was one of yoir comments on make(1)'s dynamic dependency graph that reassured me I had a correct grasp on its execution model!
It wasn't standardized before, so it didn't "always work" on any implementation.
Is it mostly for shell scripts? Aren't people targetting bash or basic bourne shell features intead of posix? Is shellcheck checking for best practices instead of POSIX compliance?
And for other applications (GUI, servers, etc) strict POSIX compliance might be too restrictive?
And with many things being Linux (or Linux-like like WSL) the need for this might be less?
Are Android and/or iOS fully POSIX compliant?
Any good blog or presentation describing the current state of POSIX?
I know many banks still have AIX systems with shells like ksh89, ksh93, etc. as the default shell. So if a shell script is written to work with a POSIX shell (instead of a particular shell), it has a better chance of running on such systems.
Also, on Debian, the default non-interactive shell is dash [1]. This is the Debian Almquist Shell (dash). It is a POSIX-compliant shell derived from ash. So again, if we write system scripts for Debian and want it to run on Debian without any hassle, it makes sense to write the system scripts to conform to POSIX shell. Although shellcheck cannot perform full POSIX compliance check at this time, it is still a pretty good tool that can help with checking compliance with dash in particular.
Or explicitly use bash in your shebang.
One of the problems with Bash is that it insists on doing bash-y things even when you tell it to act like sh.
People ask why you should write (or at least test) code to be multi-platform (even the basics of running it on BSD or macOS): it's because it forces you to be honest. Things change and initial assumptions may not be the same forever.
* https://wiki.debian.org/Shell
Its behaviour is a common behaviour of all sh, not just Bash.
> Since it executes scripts faster than bash, and has fewer library dependencies (making it more robust against software or hardware failures), it is used as the default system shell on Debian systems.
I’m not talking about shops that ship software that customers receive and install on prem on their HPUX or whatever. That’s still a thing and people have to take that into account. I’m grateful I’m no longer among them.
https://lwn.net/Articles/343924/
One big factor is for performance reasons in shell scripts. At the time, the switch decreased boot times for Debian by 7.5%. Bourne shell features add a lot of overhead and that's not always an acceptable tradeoff.
Also, if you're using bash features in a script, you can always just add #!/bin/bash to the top of your file instead of #!/bin/sh to force a bash compatible shell.
I seem to recall it was much smaller than that, something like 4% on a 2008 EEE-PC or something like that, but I can't find any numbers on that right now.
The Debian startup scripts were already POSIX; it's not hard to get better performance out of zsh or bash by avoiding expensive processes lookups.
Overall, I consider this to be mostly a myth, or at least extremely simplistic.
And that should no longer be a relevant factor, since most of the boot process is now implemented directly in C (within systemd), instead of a bunch of shell scripts.
https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
POSIX omits from standardisation what is not necessary for an application to work, so that a wide range of systems can be supported. As for Shebang, it can be rewritten in the installation script and is therefore considered an area of system administration.
EDIT: I was thinking about Linux, but I suppose macOS users are stuck with needing this for Homebrew-supplied bash?
Defining a stable API to code against?
> And with many things being Linux (or Linux-like like WSL) the need for this might be less?
Define "being Linux". RHEL? Ubuntu? Other? Is /bin/sh linked to Bash or something?
* https://mywiki.wooledge.org/Bashism
* https://linux.die.net/man/1/checkbashisms
> Are Android and/or iOS fully POSIX compliant?
UNIX® Certified Products include macOS:
* https://www.opengroup.org/openbrand/register/
POSIX:
As for the current state of POSIX, well, you're looking at it. Might find a blog or two of someone on the POSIX committees, but the organizations aren't the kind that keep blogs. Probably best to just dive into the Wikipedia article on POSIX and start following the references on the bottom. You'll probably want to look into SUS, the Single Unix Specification, as well: it's identical to POSIX (plus curses for some reason) but it's the label that OS vendors may use rather than POSIX. macOS and some Linux distributions claim to be fully SUS-compliant; Linux as a whole does not, because its official scope is limited to the kernel which only implements a subset of POSIX.
Fun fact: the name "POSIX" was coined by Richard Stallman.
Meaning: i don’t want to know all the 6000 functions glibc support, i want to know what will (for example) net/inet.h will bring to my code (ideally with documentation).
For sonme reason that doesn’t seem to be a thing. Not for glibc for sure. But the SUS does that. And i like it.
Every now and again, I get annoyed by those cases where Linux has some API/utility/etc but macOS doesn't. Getting that API/utility into POSIX greatly increases the odds that Apple will end up implementing it. (Whether just by copying it from FreeBSD, or by writing it themselves.)
As an implementer I'm often more interested in the exact changes than in the current wording. My product is already supporting the old spec, what do I need to change to support the new one? A redlined version is more valuable than the full PDF. Bonus points if it actually comes with the reasoning behind it so I don't have to guess why some seemingly-arbitrary change was made.
My dream documentation is a simple Markdown file (or similar) stored in a git repository. It allows me to see the current version, the old version, the diff, and the commit messages can even store the reasoning.
https://lore.kernel.org/linux-man/04801FEA-3560-4BA5-93EF-76...
i guess the parts of the ecosystem low enough to care about things like POSIX compliance are mostly attached to some foundation or other, so maybe those foundations will purchase copies for their core maintainers? but that's a pretty counter-intuitive thing. i wonder if there are large closed-source POSIX implementors out there that this is aimed at, but are there really enough closed-source implementations out there for any of them to care about compatibility with eachother?
if command -v command >/dev/null 2>&1
then command -v local >/dev/null 2>&1 || alias local=typeset
fi
eval "__fn=;__fn(){ local __fn=leak;};__fn || :;"
if test -n "$__fn"
then echo local leaks here!
fi
It will not crash on posh because I'm being tricky with eval and command. posh supports local variable scope.The only shell partially missing local support is ksh, but it has a gotcha. It works if the function is declared with the `function` keyword. All you have to do is use ksh's own tools to redeclare all functions automatically:
__list=$(typeset +f)
IFS="$__eol" # __eol should have a line break
for __decl in $__list
do
__name=${__decl%" #"*}
__name=${__name%"()"}
__body="$(typeset -f "$__name" || :)"
eval "function $__name ${__body#"$__decl"}"
done
IFS=" "
Of course, for this to work, all functions must be loaded before running and any declared after the fix will not be local, which is a good idea anyway. I always put the alias polyfill on the header of my library and the eval/for polyfill just before invoking my main function.There is still an inconsistency with default local values. To get the same behavior everywhere, always initialize local variables
THIS IS FINE:
local foo=;
local foo=bar;
THIS CAN INHERIT WEIRD STUFF: local foo;
Done, you have portable bourne sh scope everywhere imaginable.I'll ask again, what would it mean for POSIX to "support" Wasm?