Pure Bash Bible
github.com
github.com
https://github.com/taviso/ctypes.sh/wiki
There are some little demos here:
https://github.com/taviso/ctypes.sh/tree/master/test
I even ported the GTK+3 Hello World to bash as a demo:
E.g. flock a lock file and set CLOEXEC so that subprocesses don't hold the lock open after the shell exits.
E.g. Use memfd_create to create a temp file and write a key. Then pass /proc/$$/fd/$FD to programs that need the key as a file. When the shell exits, the file can no longer be opened.
You can do similar things with traps, but they aren't guaranteed to execute, whereas these OS primitives will always be cleaned up.
Here's an example of what bash is capable of: https://github.com/dylanaraps/fff/ (a TUI file manager written in bash)!
Did you just decide one day that you have to write a distribution from scratch? What was the thought process, and how complicated is it actually. Also, I'd like to contribute if there's a chance.
Thanks, I appreciate it! :)
> Did you just decide one day that you have to write a distribution from scratch?
Pretty much. I'd been distro hopping for some time and wasn't happy with any of the choices in front of me.
I wanted something that could run without `dbus`, `glibc`, `systemd`, `wayland`, `polkit`, `elogind`, etc etc and none of the other distributions could provide this.
Even Gentoo through their arms in the air when Firefox 69 broke the `--disable-dbus` configure flag (and added a mandatory dependency on `dbus`).
I instead spent the hours patching `dbus` out of Firefox 69 and that's how I ship it in KISS. https://github.com/kisslinux/repo/blob/master/extra/firefox/...
> What was the thought process
Start from zero and build piece by piece questioning each step along the way. Is this needed? Are there alternatives? Can we do this in a "simpler" way? Step away, come back to it later and ask "was this right?", "can we trim back the fat?".
This repeated until things were effectively "done".
> and how complicated is it actually
No piece of software seems to list its (mandatory) dependencies properly so it was a trial and error of figuring out _exactly_ what each piece of software needs.
Looking at other distributions themselves wasn't much help as they list a lot of "optional" dependencies as "required".
There's also no (or very little) documentation online for how to write a package manager or Linux distribution from scratch.
It's been a tedious but rewarding process thus far. I'm talking to you from KISS right now! It feels good to turn on my laptop and be running a distribution I created from scratch. :)
> Also, I'd like to contribute if there's a chance.
Go for it! In terms of contribution there's bug reporting, fixing documentation, adding missing packages, fixing bugs in existing packages etc.
Hop on IRC (#kisslinux @ freenode.net) if you'd like to chat. :)
Upon checkout you need to confirm the sale via a email sent to the email address provided by you.
I may be the exception and granted - it's basically due to a very crappy phone on which email ceased to work - but I was not able to finalize the sale since I'm not able to access my private email remotely.
I sent myself the link and will probably give it another shot from home. But you may want to take it up with the seller that there are folks out there for which this is not really a convenient way to close a sale. Especially not after entering valid credit card information.
Thanks for letting me know and apologies for the inconvenience.
I'm probably anyway the exception nowadays not being able to access private email remotely. But what you may point out to them is that it may be worthwhile to think through their checkout process. Also for their own benefit and the benefit of other authors
As I said I sent the link to myself and if I don't forget will buy it from home.
I really like to support you and your efforts and it looks like an awesome resource for somebody using bash a lot.
I can see now that there's a clear interest in a release in physical form. I'll start seriously looking into it. :)
The hacks (as a 30+ year sh / ksh / bash user) are indeed Very Cool.
The first example is not human readable, if the name of the function is a lie, I have no idea what this piece of code do :
trim_string() {
: "${1#"${1%%[![:space:]]*}"}"
: "${_%"${_##*[![:space:]]}"}"
printf '%s\n' "$_"
}Just yesterday I was struck by the difference between if [[ ]]; and if [ ];
I didn't even bother grokking the difference in the end. I simply found something that worked and moved on with my day.
This is my understanding. Possibly not 100% correct but essentially correct enough for me to understand the reason/rationale for the difference.
It's certainly less readable than, say, my_str.strip()
The quotes don't function like the parens. If this were a two argument function, you wouldn't put one pair of quotes around the whole thing. They're clearly transforming the variable somehow, but I'm not sure how and/or why they're necessary.
I generally tell people that, once a script is longer than ~100 lines and/or you start adding functions, you're probably better off with something like Python.
I know that's not a popular opinion with shell enthusiasts, but it's saved me so much frustration both in writing new scripts and coming back to them later for refactoring.
In the past I almost exclusively used Python, but I'm starting to like Go.
The practice of programming in shell languages was well established when Bash was designed and Bash was definitely designed with that use in mind. So by your definition Bash is a "real" programming language.
Lisp on the other hand, was much more designed as a system and formal notation to reason about certain classes of logic problems.. So, Lisp is not a "real" language?
No, it can not.
You can definitely tell there's a different "feel" to bash and Tcl, Python, Perl, Go, etc, yes? Shell languages basically evolved out of batch processing languages that were meant to only run programs in sequence and it shows.
Agreed. My threshold is usually "when you want to start using arrays or dictionaries". I find bash best for file manipulation and running programs in other languages
I never even bothered to learn Bash properly because even that's difficult, and I figured it'd be "good mental money after bad". And, now that Python ships with all distros, I feel even less need to.
I do like Bash's range syntax though:
for i in {0..5} ; do echo $i ; done
Like Ruby's. It's so nice! D: I lament Python's lack of it. echo {1..5}
is even nicer. Or echo {1..5}{1..5}{1..5}I am curious if people will uptake PowerShell Core (though not much of a fan there either).
If you need >100 lines of code your code is too complex in the script world, and you should split it into several scripts where each does one thing, and that one thing well. Usually, bash scripts are applied with pipes. A sequence like cat x | tr a b | sort >output is much more likely and easy to handle then a single script which does all these things.
KISS with Bash is a very different approach compared to other scripting languages like Python and Perl. There you can write long code easily and conveniently. However, things can get tough when larger scripts need to be maintained.
I consider the examples in the "Bash Bible" a collection of useful black boxes. It's fine if they just work. Regarding the "unreadable" trim_string for instance, if you have problems to understand that code, and you have to change something then you can simply write your own new trim_string script, even in Python or Perl if you like. Pipes work also well with them.
Update: Another advantage of bash and pipes over Python/Perl is that the Unix system can assign each script in a pipe to a separate thread. That means, simple bash scripts with pipes can work _much_ faster than single scripts in Python or Perl.
String munging in Python is so much objectively easier to read than bash. I would say that it's because it is similar to a lot of the more popular languages than the sort of cryptic parameter substitution bash provides. And if that makes it easier to maintain some automation, then its worth the effort just to use Python IMO.
If people can't figure that out... I don't know how you're going to expect them to read a cryptic bash script.
Don't get me wrong there is totally some good use cases for bash. Init-scripts come to mind when you don't want to lug around a Python VM in a lightweight container, for example.
Seems disingenuous to me. Java and Python are in a different class of readability than Bash or Perl.
Example: "abc" + "def" = "abcdef" vs "abc"."def" = "abcdef"
Would be better comparison if you would use whitespace in the other case as well. Which then makes it just as readable.
Also. "0" + "42" would that be "042" or 42? It may be better readable, but the semantics are unclear.
>>> 0 + 42
42
>>> "0" + "42"
'042'
>>> "0" + 42
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: cannot concatenate 'str' and 'int' objects
>>> int("0") + 42
42"How do we close a statement.... errr. Let's spell it backwards".
That's on the original Bourne shell and its author's evident love for Algol 68. Wikipedia has a fine summary.
Here's how I'd write that function: trim_string() { python3 -c 'import sys; sys.stdout.write(sys.argv[1].strip())' "$1" }
If you're building something that multiple folks will be reading and using just call out to external processes and make it easier to follow what's going on. It's neat to know how to do these things, but they're honestly a bit of a security hole because a large portion of folks using them will never comprehend the why and how and just assume it's doing the proper what.
Also readability: While some of the examples may be somewhat confusing to someone who doesn't have a lot of Bash experience, to someone who does it can be a lot more readable then feeding a string in yet another language to an external program and processing the output..
Your definition of "reasonably complete" might not work for everyone else's use case.
There are resource constraints even in large-ish embedded. Your main file system may live in a decent amount of flash space, but when the kernel is booting, it uses a tiny file system in RAM, which is pulled out of an initramfs image. There is a shell there with scripts.
The partition for storing the kernel image (with initramfs, device tree blobs and whatever else) might be pretty tight.
However, because a lot of these snippets depend on Bash extensions, they preclude the use of a smaller, lighter shell.
Every explanation of why to use bash here seems like it's got a whole lot of constraints on when it's a good to use - I've finished some tasks just in bash scripts but when I'm writing something for anything other than one time passes it seems like more of a maintenance liability than anything else.
IMHO, spending the time to actually learn some bash can really improve one's CLI life. Bashing on bash seems mostly like a tired ol' trope, "If it's in bash, it's bad."
The author of the original article really has written some pretty shell code!
I also think that a discussion about readability makes sense in a post about Bash. I don't personally program very often in Lua, Go or Ruby, but for the most part when I encounter code in these languages I don't find it very difficult to understand and modify. In contrast to this, every encounter with a significant amount of Bash seems to lead to a great deal of googling.
It is neat to discover functionality of bash I was unaware of by see folks push it to the limit though.
This is all from somebody who has just written a linux user space. Perhaps the author avoids shell as much as possible. It just isn't possible. I doubt it, though. Bash is awesome.
#!/bin/sh
FOO=" some long string here "
FOO="$( echo "$FOO" | sed -e ' s/^[[:space:]]//g; s/[[:space:]]$//g ' )"
This isn't "pure bash", but most of what I write in shell scripts isn't "pure bash". It's shell scripting: dirty, slow, easy, effective.Like any 'language', it takes on the complexity you put into it. English is really complicated, but you can also use a subset of it with only 850, 1200, or 2400 words, and suddenly it's very simple and clear.
Your example fails when $FOO is "-n", for instance. Also the g modifiers are redundant in your example since there's only one beginning of line per line and one end of line per line.
I would instead write this:
sed -r 's/^\s+|\s+$//g' <<< "$FOO"
EDIT: I think I see now what you probably thought would happen by using g, but no it wouldn't remove multiple spaces. So, you example also fails when $FOO is " x" (using 2 or more spaces at the ends).https://www.gnu.org/software/bash/manual/html_node/Shell-Par...
My inner pedant has no comment regarding readability of Bash parameter expansions.
If we didn't have regexes we would have 100s of if/else's scattered all over program logic. That would be more hard to handle than the regex itself.
These scripts/snippets for me are just an unofficial extended standard lib
Then stop maintaining code from languages you don't understand. This is frustrating. I've seen solutions on the comments that call PYTHON! Are you effing kidding me? PYTHON ??
Granted, BASH docs aren't particularly succinct, but shell scripts are an absolute necessity in the OS world.
It's pretty simple: If you don't understand shellcode, don't maintain an OS, or rather, don't expect to be accommodated for lack of knowledge for something that's been standard for 30+ years.
Let's use bash instead of a real programming language! It is super convenient! To do super fundamental things like trimming strings you just have to implement your own function with 40+ non-ASCII characters in a row! That example makes regexps seem like human readable and god knows if it is even correct or works portably across different versions of bash.
To be fair the article assumes you are first in bash and want to avoid launching subprocesses but for any production scripting that is the wrong hypothesis to begin with. A better approach would be to not start inside bash at all. Just do your scripting in a higher level language where basic ABC stuff like string, list, number and error handling are already there for you. Even if that trim_string function and everything else from the article would be provided in some bash-bible-std-lib it wouldn't even come close to what's available in say Python for example, and you still have to wrestle the syntax and other obscurities like -a (or -e) meaning file exists.
While there are some good examples in there if you are stuck in bash, this article was more of a 100 reasons not to use bash to me.
My focus for the past few months has been writing a Linux distribution (and its package manager/tooling) in POSIX sh.
I've learned a lot of tricks and I'm very tempted to write a second "bible" with snippets that are supported in all POSIX shells.
(I created the bash bible).
That could be potentially even more interesting than a bash-specific one (as it is harder to get it right -- bash can be figured out out of the single reference, anything "portable" has many dependencies).
- Safely working with "string lists" (list="el el el el").
- Filtering out duplicate items.
- Reversing the list.
- etc.
- Using `case` to do sub-string matching (using globbing).- Using `set -- el el el` to create an "array" (only one array at a time!).
- `read -r` is still powerful in POSIX `sh` for getting data out of files. `while read -r` even more so.
- POSIX `sh` has `set -e` and friends so you can exit on errors etc.
- Ternary operators still exist for arithmetic (`$(($# > 0 ? 1 : 0))`).
- Each POSIX `sh` shell has a set of quirks you need to account for. What works in one POSIX `sh` shell may not in another. (I found a set of differences between `ash`/`dash` in my testing).
POSIX `sh` is a very simple language compared to `bash` and all of its extensions so there won't be as many snippets but there's some gold to be found.
I've started working on it here: https://github.com/dylanaraps/pure-sh-bible
> are there any bashisms that are truly essential and you don't want to live without?
The only thing I'd say I miss when writing POSIX `sh` is arrays.
I work around this by using 'set -- 1 2 3 4' to mimic an array using the argument list. The limitation here though is that you're limited to one "array" at a time.
The other alternative I make use of is to use "string lists" (list="1 2 3 4") with word splitting.
This can be made safe if the following is correct:
- Globbing is disabled.
- You control the input data and can safely make assumptions (no spaces or new lines in elements).
While it's something that'd be nice to have, there are ways to work around it.
EDIT: One more thing would be "${var:0:1}" to grab individual characters from strings (or ranges of characters from strings).
set -o pipefailhallelujah.sh
yes, it looks like a device node in /dev, but it's really a pure bashism for opening tcp connections to arbitrary hosts and ports
I've since implemented a very bare-bones and very featureless IRC client using /dev/tcp and bash.
https://github.com/dylanaraps/birch
I will get around to it eventually. The one hurdle I want to get over before writing a piece about it is the handling of binary data using bash.
This is something a little tricky to do with bash but it'd allow for a 'wget'/'curl' like program without the use of anything external to the shell (no HTTPS of course).
I want to really understand the feature before I write about it though in the meantime I could just write a reference to the syntax/basic usage. :)
I came here to say the exact same thing. This is my favourite thing most people don't know exists in Bash.
In particular the "obsolete syntax" section, I wasn't aware of it.
https://github.com/dylanaraps/pure-bash-bible#obsolete-synta...
Bash in RHEL is in /usr/bin and /bin as /bin in symlinked to /usr/bin. I think it is equally unlikely that RHEL (Debian, SLES..) will will move either /bin/bash or /usr/bin/env as it would break a million scripts out there.
If we should migrate to FreeBSD while, for some reason, reusing linux oriented bash scripts, changing the path to /usr/local/bin/ would be the least of my headaches.
I agree that 'env' can make good sense if you don't know who/where/when your script is used. For internal projects, I don't really see the advantage.
$ uname -a
Linux localhost 3.10.49-5975984 #1 SMP PREEMPT Thu Oct 8 17:25:20 KST 2015 armv7l Android
$ which env
/data/data/com.termux/files/usr/bin/env
https://xkcd.com/927/ trim_string() {
# Usage: trim_string " example string "
: "${1#"${1%%[![:space:]]*}"}"
: "${_%"${_##*[![:space:]]}"}"
printf '%s\n' "$_"
}
Ok, so the : is somehow a temporary variable...
Then there is a variable starting at $ and you lost me :D
Can someone break down that line for me? What the hell is going on here? : "${1#"${1%%[![:space:]]*}"}"Because $_ is used in the expansion of itself, it is the same value, because the command has not yet completed, which would (re) set $_. So use of a thing ($_) in expanding the same thing ($_) is perfectly fine until after that (null) command (:) runs. You see that the final $_ is used standalone. Hope this helps.
: "${1#"${1%%[![:space:]]*}"}"
The $_ temporary variable contains the result of removing the leading spaces. In the next line, the spaces at the end are removed from the temporary variable with the "${_%..." syntax.You can test this in your own shell by e.g. doing:
: $PATH
echo $_: is a "do nothing" command -- but the line is still evaluated
%% means to replace leading chars that match pattern
## means replace trailing chars
I don't know why they're using $_; thats the variable containing the interpreter name, i.e. "/bin/bash" [edit - also the name of the previous command!]
I can't be bothered analyzing it any further :-)
There are some neat tricks in here but they don’t seem very readable compared to perl/awk/sed.
``` trim_string() { # Usage: trim_string " example string " : "${1#"${1%%[![:space:]]}"}" : "${_%"${_##[![:space:]]}"}" printf '%s\n' "$_" } ```
and think: "I should use bash more".
Bash is nice for making simple things simple but for complicated things it's just shitty. I used to think that this is due to the complicated quoting rules which make the simple things simple but tcl does a much better job at that.
In either case I prefer the clean rules of a Python or Perl for anything larger.
If you find yourself frustrated by the lack of bash extensions, your program is probably complex enough that you probably shouldn't be writing a shell script.
I had the same issue with reading other people's perl as well.
I think the great and terrible thing about both languages is that there is literally a million different ways to skin the cat / write a regex.
trim_string(){
# Usage: trim_string " example string "
: "${1#"${1%%[![:space:]]*}"}"
: "${_%"${_##*[![:space:]]}"}"
printf '%s\n' "$_"
}
This reads horrible. I see no reason to prefer this over programs like sed, bash is after all a shell, intended firstly for running external programs/commands.Are there systems that don't come with sed installed? (Some docker containers I have logged into don't seem to have less).
[1]: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/s...
Just use the shell function by its descriptive name...
The reasons are explained in the foreword:
Calling an external process in bash is expensive and excessive use will cause a noticeable slowdown. Scripts and programs written using built-in methods (where applicable) will be faster, require fewer dependencies and afford a better understanding of the language itself.
I'll take the grep/sed/awk version.
I guess my bash uses cases are more just around internal tooling and are never performance critical, so my personal preferences are for readability under those circumstances.
> decades ago
Frankly, I can only imagine how the environment then would be. Thinking back with your current experience, what do you think you would have done if you had to fix it again?
Sometimes even bash isn't an option, you have to deal with older shells like ksh.
I use to make lists first, then process the lists, I still do this sometimes since its faster. If you have to run a query every time, your probably doing it wrong, but for small stuff, everything is so far, I can chain gnu apps and be done. I'm not a programmer, I'm a sysadmin so mostly deal with the fixing things like auditing or fixing data on a file system. (or maybe db)
lower() {
# Usage: lower "string"
printf '%s\n' "${1,,}"
}
And with uppercase, "${1^^}" instead.Why I think people gravitate towards it is because languages such as python add too much pomp to launching a shell process. A language like perl is usually easier to use but everyone hates it now.
I would heartily agree. If your shell script grows beyond half a dozen commands or so, you're probably better off rewriting it in just about anything. Python, ruby, go (gorun), rust (cargo-script), or whatever else, doesn't really matter as long as it's not shell.
It also tends to be very portable.
This is perhaps one big benefit but not one that is exclusive to a sh-like language. Instead I would like to see a language with strong flow control or metaprogramming capabilities take on processes as a first class citizen. Perl is probably the closest but still has some warts related to redirection.
The best pattern I have seen is encapsulating business logic into Python or Go and then if really necessary piping it to another script. But, often, if you do this you can just keep the piping internal using data structures.
Unix-pattern facilities work very well for interactive use, which tends to be simple, exploratory, and trial-and-error. But a project's build script may not be simple.
In fact this is how Perl 1 looks: https://st.aticpan.org/source/RCLAMP/perl-1.0_16/
https://github.com/Perl/perl5/commit/8d063cd8450e59ea1c611a2...
I quite like Powershell on Windows, but I'm not sure I dare try it on Linux. I think it might make my head explode.
I wholehartedly agree on perl but can you expand on the redirection warts? I seldom had problems with perls' FHs, while, on the contrary I seem to be unable to wrap my head around the contorted syntax involved in bash's handling of descriptors - especially when more than 2 handles are involved.
The wart is that you need IPC::Open3 or equivalent because Perl's intrinsics can not synthesize the pipe operator (though you will think that they can) if you need to insert yourself into the middle of a chain of commands.
Nowadays there are decent wrappers for calling Open3 but more commonly you just find people running one half of the command, buffering the output, and passing it to the second half.
https://github.com/lmorg/murex
It currently has:
* Proper error handling (eg try and catch blocks)
* unit testing and debugging frameworks to help with development and maintainability
* data-type aware, including complex types like how CSV, JSON and YAML are all handled as memory structures and thus the same tools can query any structured data format without understanding it's contents
* while still ostensibly working the same way as a traditional POSIX shell
There's also some work on improving the REPL experience too where I've included:
* automatic man page parsing for flags
* a "tool tip text" like hint line which tells you where commands reside on the fs, what it does, etc. Which is handy if you're trying to debug an existing commend.
* if you paste multiline text into the console you get offered a chance to preview the text before executing it (handy if, like me, you're pretty useless at copy/pasting content reliably)
* a package management system so you know exactly which functions and imported scripts are loaded from which sources (no more "where did that autocomplete suggestion / alias / etc get loaded from?"
There's a few other features like support for events and such like, but they're not yet documented.
The shell is currently beta but I've been using it as my daily driver for about 18 months now. There are still quite a few bugs, plenty of places where code needs to be rewritten for performance and lots of stuff isn't yet documented (though the documentation is pretty good already considering it's only me working on it). So don't expect a finished product. However I do think I'm at the stage where I'm ready for more users to have a play and I welcome PRs, issues raised, general comments and feedback, etc.
# end of shameless self promotion :D
I'm aware of Elvish and really impressed with what's been built and the traction that has gained. There is definitely some overlap between murex and elvish but also a lot of area's where our shells differ.
I think there is sufficient difference between the two shells to justify their existence.
For these, being able to avoid the various $(echo... |sed) can be refreshing. Beside, the book makes for a nice repository of techniques.
It's definitely possible to write shell code which properly passes shellcheck's linter though it's an uphill battle to learn the ins and outs and _why_ X is wrong when Y is right.
I even managed to write a full TUI file manager in bash!
https://github.com/dylanaraps/fff
I full understand that there are times when the shell should not be used and when other languages are a better way to solve a specific problem, however I love pushing the shell beyond its supposed limits! :)
I've read pretty much everything I could get my hands on regarding the shell (including the mentioned link) and I still love it.
> Do you write truly correct Bash/POSIX code?
If we define correct as passing shellcheck, avoiding all pitfalls and maintaining compatibility (POSIX sh not bash), then yes, I like to think so. :)
> Do you still love it?
Oh yeah! I've been writing a ton of POSIX sh as of late. My latest project being a Linux distribution: https://getkiss.org/
(hello from Firefox in KISS!)
Edit: follow up question is Do you believe bash / POSIX shells actually follow KISS principles?
Not questioning whether your OS is KISS, but I don't think that necessarily reflects the KISS-ness of the underlying language.
My questions clearly reflect my current impression that in the long term, shell pitfalls largely undermine the benefits of its apparent simplicity. The gist would be for you to provide some way to change my mind. I guess the codebase you provide is a strong counter example; but you'll agree it doesn't reflect general usage of shell in the wild.
Edit 2: You know what, I just read your original comment again. I kind of retract my question since you do concede that it's an uphill battle and you love it in spite of its flaws. I guess that's cool (and I agree it's fun trying to write correct bash as a challenge) as long as you're in control of the code being produced, but my main impression remains that it's a bad language to publicize and its presence in most codebases inherently bears a strong cost.
POSIX `sh` yes. `bash` less so but I'd still lean more towards a yes.
Ultimately though, it depends on how we define "simple". Both `bash` (2.6MB) and POSIX `sh` shells (`dash` (232KB), `ash` (1.2MB (busybox)), etc) are tiny in size if we compare them to Python (137MB) or Perl (44MB).
(Numbers taken from my system using `du` on each file which belongs to each shell/language.)
If we define "simple" to language features then I think the shells come out on top again (especially POSIX `sh`).
If we define "simple" as ease of use (without shooting yourself in the foot) then I'd agree with you and say that the shell loses here.
There's a time and place for using any tool (in production) but I find it fun to push the shell beyond what is thought possible in my personal projects. :)
f = open("ls|", "r)
f.read()
f.close() f = Popen('ls', stdout=PIPE).stdout
f.read()
f.close()
alternatively given this exact behaviour: run('ls', stdout=PIPE).stdout cat something | grep "this" | cut -f 1 | sed -e 's/.../.../'
You end up writing too much code, it's very verbose. Some times symbols are what you want. In fact the biggest progress in the growth of Math happened when they tossed out doing math with words and bought in symbols.> cat something | grep "this" | cut -f 1 | sed -e 's/.../.../'
Literally none of this is actually useful if you're already in Python:
(
line.split('\t')[0].replace(…, …)
for line in open('something')
if 'this' in line
)
> You end up writing too much codeI can believe that if you're calling to external processes to perform operations which are pretty much trivial in the language.
This question is specific to pipes.
That's missing what I'm noting though, which is that you don't need pipes anywhere near a shell script if you can simply do more of your work in-language. Case in point being that the entire pipeline you cited as an issue has no reason to exist outside of a shell or shell script.
It doesn't feel like the language was designed for these tasks.
But oh, if you do set -o pipefail the grep will stop the whole pipeline when none of the lines matches "this". So you have to keep fiddling with ${PIPESTATUS[0]}. And none of pipefail or PIPESTATUS are really portable.
Not so simple after all.
The python code to do piping ends up as longwinded and the plumbing of pipes ends up a massive headache so I wrote tidycmd to overcome that issue https://github.com/laurieodgers/tidycmd
In Perl you literally took the part of your shell pipeline, quoted it and opened it like any other file.
./output_generator | wc -l
becomes
open(SRC, "output_generator|");
open(WC, "|wc -l);
And you do whatever you want with those filehandles. It's been ages since I wrote any perl. I miss it.It's deprecated and does something quite different.
> If you ever dealt with much perl you know how much more pleasant and easy launching processes was in that language - like shell.
I'm sure it is, but you're missing my point.
> If you want to use that as evidence that I'm a garbage developer go ahead
Well that escalated quickly.
> open(SRC, "output_generator|");
> open(WC, "|wc -l);
>
> And you do whatever you want with those filehandles.
Again python does roughly the same, just with more overhead: a trailing pipe is an "stdout=PIPE", an input is a "input=<whatever>", and you access / forward stdout explicitly:
src = Popen('output_generator', stdout=PIPE)
wc = Popen(['wc', '-l'], input=src.stdout)
And as the sibling notes, for shell replacements you can use the sh library to lower the syntactic overhead of popen.Programmers take great pride in freeing accountants and ware house workers from drudgery. But seldom do we look at our work in the same way.
In the real world, most software work is done very similar to digging coal mines with shovels. Laborious manual hand typing jobs.
There are things you can do in line of perl that will take you 5 -10 min in python, but once program goes above one liners, difference is not that big.
And of course, if you learned vim instead of emacs, you would be even faster :))
Anyway bash is still faster to use, and a lot of Unix services are based on it. The book is well written and have a great added value. Thank you for sharing!!
Ah, the luxury the modern breed of programmers enjoy today makes me feel jealous. In these days of Splunk and Document databases(returning jsons and xmls) its hard to understand why so many things in the past were the way they were.
Apart from DBMS interaction(Perl had DBI/x modules for that), pretty much every thing in the decade of 80's even upto late 2000's(tons of legacy systems) was so non standardized that people were literally parsing through log files and non standard data formats to store/exchange a lot of things, this was when the internet was growing crazy year over year, systems needed to be built and put in place. This means you really need to have first class facilities to manipulate text. You needed regexes baked neatly into the language. You needed qw, you needed ``, you needed while<FILEHANDLE>, you needed binary file handling features, you needed string manipulation facilities that could help you drill through any text file you could imagine, you needed powerful functional programming features, you needed OO etc etc. And you needed to get this done under tough deadlines on slow machines. Remember Java being cross platform compliant was one of the biggest selling point of the day. Perl had this before Java.
I personally worked on building a store using rcs and perl, that could store versioned config files, almost like a document data base. The parsing facilities required for the application we did demanded nothing short of a tool like Perl.
Python, and also Java growing rapidly once the data exchange formats were reduced to mark up languages and JSON. Suddenly you could with a library what most Perl programmers were doing using their language powers.
Also look at the Human Genome Project and Perl usage there.
As to why people like it: it often feels more natural when you are automating what you would type interactively. Unix pipes/coreutils/etc. also feel like a better fit when that automation is mostly about connecting other programs (it's the auxiliary stuff that you'd maybe want in pure bash). Reading the subprocess Python documentation does not exactly fill me with joy. I've heard libraries like Plumbum make it a bit neater - but then you have to ask why learn a bunch of new libraries when I already know bash? In the end it's about the best tool for the job. The danger with bash is going too far, especially if you don't actually know it very well.
if [ ! -t 1 ]; then
exec > /my/log/file 2>&1
fi
The if statements tests if your at an interactive prompt, if your not all output from the script get's redirected to /my/log/file. The above poster is instead redirecting into a subprocess " > (tee)" that will both print the output and log it.It should be noted that often the bottleneck is the terminal itself, try running your scripts with a "> /dev/null" to suppress output and verify the slow part is actually the script.
It is unstructured only in the way you allow any running command within your script to dump their output.
I run most commands inside my scripts with `> /dev/null 2>&1` and then rely on exit codes to wrap structured information to be echoed out with this function:
echo_l(){echo;echo '--------';echo "${1}";echo '--------';echo}
Or functions such as this to populate a log:
log_cat(){echo "${1}" >> ${__LOG} }
But the tee named pipe is the winner.
PS. The __DOC_LOCAL and __DIR variables start with these magic variables below. These variables are a life saver and allow easy directory and file manipulation, they kind of setup a top-level context:
# Set magic variables for current file & dir
__DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
__FILE="${__DIR}/$(basename "${BASH_SOURCE[0]}")"
__SCRIPT="$(basename ${__FILE})"
__BASE="$(basename ${__FILE} .sh)"
__ROOT="$(cd "$(dirname "${__DIR}")" && pwd)"
If anything code that i used to write in Python at the past i write it in Bash nowadays, exactly because Bash has a better record when it comes to not breaking stuff. Though it helps that my Python use is also mostly scripts meant to run from the shell.
And honestly I can't see any good argument for this patchwork approach to gluing things together. I guess some ops people might argue that you'd have to have ruby everywhere but the counter argument would be that we use docker images for everything and adding ruby as a dependency isn't any worse than all the insane dependency gymnastics it takes to get our node apps working.
And all of this applies equally for any language with a reasonable standard library (python, perl). I think people have weird feelings about using bash or make or whatever to accomplish things, like they are riding closer to the metal or that they are living some deeply pragmatic zen Unix philosophy, but mostly they are making an un-testable mess until it works once and then, if they are lucky, they don't have to touch it again.