Something about functions in Bash (2017)
catonmat.net
catonmat.net
In BASH, as well as any POSIX shell, you can define a function with function_name () and just remove the "function" from the front.
one of the many things I learned by using Shellcheck for static analysis https://github.com/koalaman/shellcheck
I finally started appreciating Perl after having to write some nontrivial Bash scripts.
(My hierarchy is: default to Bash when my task is mostly running commands. If there's a bit of text-processing or regex required, Perl. If I need any hierarchical datastructures, Python)
Ha. Talk about 'damned with faint praise'.
I'm not really in the Unix admin world, but can't you just bite the bullet and use a 'real scripting language' like Python? Other than for things like configure scripts, is the universality of BASH/POSIX really non-negotiable?
There is also the cultural difference. In Python, it's so easy use pip and modules, but if you want something to replace Bash, you need to resist the temptation and only use what's in the standard library. The thing with Bash scripts is that they are often self contained and you can grab a huge snippet of it and it will work pretty much everywhere. (Exceptions of course, the script may do some fancy include or use a program not available, but you get my drift with this.)
Bash has been stable a long time, and the Bourne shell compatible part of it for even longer.
I don't think it's fair to say that bash programs are not without dependencies, just that the developers (tended to be?) are more thoughtful about portability. For example I've seen scripts that use wget or fall back to curl when it's not available. I've never seen a python script fall back to urllib3 when requests wasn't available on the system.
Hm... I wonder if that is a good idea or a horrible one... :)
On that note this article (https://iridakos.com/tutorials/2018/03/01/bash-programmable-...) on bash completion scripts has been a big boon for my effectiveness recently. I've been creating simulations of our complex interactions with third party systems and the completions scripts are doing things like pulling the correct values from the database rather than having to type them in manually.
I knew I'd saved a bunch of time but it really hit home yesterday when I had to explain to QA how to repeat the process and it's "open this file, change these values, save it here, open other file, set these values from previous file, save here". For me the whole thing is "cmd<tab><tab>".
Does anyone have any tips on taking scripts like that and generalizing them for whole teams? At the moment they're very much built for my personal needs.
Of course it doesn’t work for everything but if you can try it out.
At some point, using bash isn't effective anymore. It's great and all, but an important skill is knowing when to use it and when not to.
[1]: http://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3...
That said, I'm curious what are the nifty tricks that one can do by using functions as "aliases" for compound commands -- I cannot see much of a difference between `function sleep1 () { while true; do "$@"; sleep 1; done; }` and the similar definition without the curly braces. Only once in my life I have had to use the `function foo () ( ... )` syntax (i.e., using subshell instead of normal `{ ... }` grouping).
It also should really help that Debian, IIRC, actually sets /bin/sh to dash by default?
...which is a great jumping off point to casually read a bit of them now and again.
( cd some_directory
do_a_thing_here )
... instead of all the pushd popd nonsense, and so much easier to follow.You can also do this:
( set -e
bail_from_this
if_anything_fails ) < redirect || echo 'Something went wrong here, carrying on...'
And that won't set -e in your outer shell. You can also set -v to see what's being run just for those commands.As far as writing functions without braces, I don't. It doesn't do anything useful, it's magical and weird to bash n00bs, and when you have to add a command you'll wind up adding the braces then.
The reason bourne shells have that is because the shell's raison d'etre is composing processes, so it makes sense that compound statements are simply processes themselves. (Maybe process is the wrong term here...) To expose that in the syntax, though, is just being overly clever, and it's more obvious now that people find it confusing.
When scripting, I find general-purpose languages are better when you have structured data (i.e. some JSON coming in) or you want to use libraries, and shell scripting better when you're wiring up multiple programs to work together, or starting up a program after setting up the environment for it.
IIRC, the FreeBSD sh(1) man page was a good entry to writing POSIX compliant, thus quite portable shell scripts.
[1] Compare the ones here: https://github.com/cadadr/configuration/tree/master/bin
And now that JigSaw landed with JDK9 and with single file running coming in JDK11(?) and with precompiling bytecode to native, it becomes easier and easier to deploy good quality code as quick and fast scripts, tools, utilities, etc.
Yes, of course, I still use bash, because it's there. But habits are easy to form and very hard to break. (Because why form re-form an old habit into a new?)
That said, my brain does still think in POSIX scripting, so it always takes me a little while to get back in the swing of Powershell.
You can work around it, but you really shouldn't have to.
Normally you just use Write-Host when you need it and otherwise just collect return values inside the function by assigning them to variables.
It's basically solved if you never call a command without storing the return value somewhere. Which is not optimal in a scripting language where you just want to get things done, but it's not the end of the world.
Dictionaries are often used for return values, as they can be cast to PSCustomObject for a quick way to make objects for the pipeline. But in that use case they aren't passed in, and that's not anything to do with the `return` keyword or all output from all commands becoming function output.
get-content file.txt
and the result was the file content visible on screen. Then you were happy with your code and you put it in a function for reuse: function test {
get-content file.txt
}
and now there's no output. PowerShell pipeline is not stdout, get-content doesn't display anything on screen, only the output formatters at the end of the pipeline do. If the function has no output by default then you would see nothing. That would be an annoying behaviour the other way. In PS, braces {} make anonymous codeblocks and they're used often, e.g. in code like # square the first 5 numbers
1..5 | foreach { $_ * $_ }
# rename some files
Get-ChildItem | where { $_.Name -notmatch '[0-9]' } | Rename-Item -NewName { $_.Name + "2" }
That would be really annoying if you always had to 'return' results out of scriptblocks. And it would be annoying if scriptblocks in functions behaved differently to scriptblocks in filterscripts or calculated properties.Looking more, it's kind of bad, because they break the sections in the reference up by the modules, and it's utterly mysterious what the modules are for. I mean, they all seem to be for doing incredibly obscure stuff in Windows, which I'm sure is useful, but I can't figure out obvious things like the syntax or types.
This may be because the intended audience is liable to just copy and paste stuff and hammer away it until it works...
Different versions of Windows came with different versions of PowerShell. Some can be patched to have the newer language features (Windows Management Framework 5.1 can be installed on Windows 7), but that won't bring in all the same cmdlets as Windows 10 has, because there aren't the required internals in Windows 7.
If you have different things installed (Hyper-V, ActiveDirectory, any big Windows role/feature) you'll have different modules available, and if you install things like RSAT (Remote Server Administration Tools) then you'll get server management cmdlets on workstation operating systems.
It's not so much like you download Python 3.6 and get one standard library everywhere, its origins are more in "companies managing their own Windows server estate", so your environment is not standard, it's whatever environment you have.
The standard library, though, is mostly .Net, so depending on your Windows and PS version, it's ".Net Framework 4.5" (or 3.5 or etc.) for .Net framework library features accessed directly from PS.
https://github.com/powershell/powershell is PowerShell 6 / Core, which is a lot more like downloading Python 3.6 - there is everything a PSv6 environment will have, open source on Github and documentation in https://github.com/PowerShell/PowerShell-Docs
rm foo bar
rm -rf *
must be invocations of the rm command with the appropriate arguments.You need bare strings, you need variable interpolation, you need redirects.
I think there's some room to break a few idioms and make something nicer, but it really is a hard problem. The choice to break things has to be based on how common an idiom is, how painful it is to not use it, and how much benefit you can get.
So where I decided to deviate was areas I felt could be enhanced without breaking the POSIX syntax too significantly:
* shell configuration,
* error handling (something traditional shells are appallingly bad at)
* and support for complex data structures (eg so you can grep through items in a minified JSON array as smartly as you can with a traditional stream of lines.
I always respected the power of POSIX shells even before embarking on my pet project. However I never quite appreciated just how sane it's ugly syntax was until I attempted to replace it with something more readable. I mean sure there are still some specific areas I really don't agree with but that can always be argued as personal preference.
As an aside note, one thing I didn't quite appreciate until writing my shell is just how much shells have in common with functional programming. Yes I know it's a far cry from LISP machines; but if you take away the variable interpolation and environmental variables then you're left with a functional pipeline that take something from STDIN and write it to STDOUT and STDERR and does so in a multi-threaded, concurrent, workflow by default. Weirdly this still makes POSIX shells more efficient for some types of data processing than writing a monolithic routine in a more powerful (or should that be "more descriptive"?) programming language.
So there is some surprising elegance amongst all that ugliness.
I'm not suggesting the work I'm doing is any better that jq though - there's a lot of areas where jq will run circles around my shell. I see it more as different design goals but with a fairly large area of overlap.
Maybe something based on Color Forth's ideas?
For what it's worth though, most shells (including my own one) do support switching between different text entry / hotkey modes - such as vi - even if the syntax is consistent.
It is already open source but I'm the only contributer currently. Which is fine as it's a personal project anyway. So anything beyond that is a bonus.
The readline API I wrote for it is pretty nifty though. I'm thinking of spinning that off into its own repo since I've been a little disappointed with the existing realine APIs for Go so I think other people might genuinely benefit from that even if they're not particularly interested in a semi-Bash compatible alternative $SHELL.
The source can be found at https://github.com/lmorg/murex
Happy to discuss any questions or comments you might have on it.
http://www.oilshell.org/blog/2017/12/17.html
I have a few more blog posts on that topic I haven't published, but hopefully that gives the idea.
I probably used to agree with you (if I'm understanding correctly; it's not entirely clear what the conflict you see is). But after writing the parser in this style, I don't think it's too big a deal.
I plan to use the same technique to parse the Oil language, which is a new language without the legacy. The command syntax is largely the same, but compound constructs like function, if, while/for, etc. are different.
The parser is regularly tested on over a million lines of shell:
http://www.oilshell.org/blog/2017/11/10.html
Note that Python, JavaScript, and Swift all need something like this too, now that they have arbitrary code inside string literals, just like shell:
print(f'three = {1 + 2}') # Python f-strings
console.log(`three = ${1 + 2}`) # JavaScript
println("three = \(1 + 2)") # SwiftReally, the only complaint I could make is why chose something silly like `fi` or `esac` to end compound statements, but such an issue is too superficial to abandon everything else that comes with these shells.
Substitution via `=()`, `<()`, and `>()`. `=()` save's a process output in a temporary file that I don't have to mess with explicitly and substitutes itself with the path. An example of use is `viewnior =(maim -s)` which shows in an image viewer a window that I select; `diff -u <(...) <(...)` or `cmp`, or `comm` instead of `diff` compares the output of multiple processes without having to save them to files; `tee >(...) >(...) | ...` feeds the output of one command to multiple commands without needing to save files
Subshells. I've used them interactively, and I prefer to not have troubles with quoting doing `fish -c '...'` in my commands.
Extended globbing. It's far more concise and less error prone than using a combination of `find` and text processing tools.
Dynamic directories. ~some_project/ for me expands to ~/work-for/client/some_client/some_project/ and works great with completion.
What attracted me the most to fish was the ability to work with multi-line commands as a single command in the history. However, I've seen that zsh has better support for this.
What killed fish for me back then was that their pipes were fake back then. This is something that they've since fixed, but back then, instead of running all commands at the same time with their inputs and outputs linked, I think fish would run the first command, save it to a file and then ran the second command with that file as input. I mean, the behaviour that I saw was that I wouldn't see any output until the first command finished, and some of the commands that I wanted to run took very long to finish or never were supposed to finish without me seeing some of the output first. I'm talking about watching file changes under a directory and manipulating the presentation of the output of the file watching command (e.g. `inotifywait -m ... | awk ...`) or looking for specific events in the system log as they happened (e.g. `journalctl -f | grep ...`) or looking for specific system calls of a never-ending running process (e.g. `strace -fe trace=file -p $pid | grep ...`). That made fish pipelines useless for me.
Anyway I wrote a more lengthy post once describing the differences between bash and zsh that I liked:
https://news.ycombinator.com/item?id=16963856
However, there must be more differences between fish and zsh. I can think of history syntax, right now. I can get an argument from any command I've ever typed based on a substring by doing !?substring?%. If I find myself calling two commands in succession multiple times, I can combine them in a single command without retyping them by doing `!-2; !!`. Then I only have 1 command to re-run next time instead of 2. If I typed a command and realize I want to include the last argument of the previous command, I can type `!$` to include it.
Stephen Bourne was a fan of Algol.
The source to the shell is… interesting. Via a bunch of preprocessor macros, the (C) code looks like:
IF !letter(*cp)
THEN return(FALSE);
ELSE WHILE *++cp
DO IF !alphanum(*cp)
THEN return(FALSE);
FI
OD
FI
return(TRUE);The language wasn't designed; it was accreted over decades.
Looking at the difficulty of the Python 2 -> 3 transition might give you an idea of how hard it is to make incompatible changes, so that's what we're stuck with.
It's a little bit like asking why C has so many warts -- e.g. the insecure standard library, "holes" in the type system like silently converting pointers to bools.
But it's easier to "clean up" compiled languages than interpreted languages, and it was somewhat done with C89, C99, etc. You can have different compilation modes, and the user is more likely to fix things that a compile error flags.
With something like shell, you don't have a good chance to surface errors and clean up corner cases. It just evolved over time until it was out of control.
I'm working on fixing that with my Oil project: http://www.oilshell.org/blog/2018/01/28.html
cat filename | grep -i "hello"
Someone else will probably point out a better way. It is also only a couple of lines of Python.
grep -i "hello" filename # ;)
> I always can't help but notice the lack of expressivity. I mean it takes 1/2 page of code to read a text file and print all lines that say "hello"
I think you hit the nail on the head there. I've used just about every build system known to man and we're currently using cake at work, which is uses c# as a scripting language. Everything is 10 times more verbose than the equivalent shell/makefile would be. Things like concatenating sql files are 30 lines instead of code instead of one liners.
grep -i hello filenameWhat's so bad about the syntax? I find it rather beautiful. Can you give a (simple) example where you find the syntax ugly, and how would you prefer it to be?
Really? Those are your best examples? The answer is Python in any case.
For most things no. They could just as well have proper hashmap and array structures for example, instead of the clusterfuck they have.
Or how about exceptions and try/catch (instead of traps and the like).
Or the side effects of the bizarro [ command (it's a command, not just the syntax for an opening bracket).
And tons of other things besides -- e.g. better argument parsing.
Better ways to add auto-completion instead of the monstrocities of bash completion and the like.
Then I noticed, tucked away at the end:
...
kill --all humans
touch $children
fork yourself
} | log
Because, you see, components of a pipeline have to be run in subshells...Thus most people will not notice it unless you point it out very explicitly, and it has the potential to cause great confusion.
https://www.gnu.org/software/coreutils/manual/html_node/inde...
declare -f function_name
to print out the contents of a function, in a format that can be directly evaluated with command substitution. This is useful if you want to send a function (or a number of functions) into another BASH instance (such as via an SSH connection).A google search for "RPC in Bash" will take you to a GIST I put up on Github a while back that gives details of how this works.
You can do this in ksh93 also, I believe you use typeset -f instead of declare -f.
get-content function:function_name
`function:` being a filesystem-like PS Provider.Huh.
He also self-published "Awk One-Liners Explained" and "Sed One-Liners Explained" and has a "Bash One-Liners Explained" in the pipeline, see http://www.catonmat.net/books
Informally, a compound command is a list of commands (`( ... )` or `{ ... }`) or a looping or conditional construct. See [1] for a precise definition.
[1]: https://tiswww.case.edu/php/chet/bash/bashref.html#Compound-...