Understanding the fork system call in Unix
mohit.athwani.net
mohit.athwani.net
Edit: I really care about the "should" part too, not just what currently happens. I've been trying to think of semantics for forking that could work in these situations but I haven't been able to think of great solutions.
That solution is very obivious once you realize it is exactly the same as what happens with the terminal and network connections.
For example, after forking, both processes are still attached to the same terminal (what else can they do? It’s not like they can make new terminals - what if it was an ssh connection). They better agree who gets to read/write, or user will get confused - in shell, the parent voluntarily stops accesing terminal until the child is done.
or you simply avoid touching them. remember that forking only creates one new thread in the child process, there won't be any code running that you don't control or that would know about those file descriptors existing.
The reason fork (or vfork, or clone) and exec need to be separate steps is because you may want to do some stuff in between the two steps. The most obvious example is redirection: when you run "grep foo < bar", the shell needs to set stdin to point to "bar" before it execs "grep".
(IIRC, posix_spawn was introduced relatively recently, so most old software wouldn't use it).
FWIW, it’s not required to be able to run any setup code in the child to do arbitrary setup. For instance, WinNT makes that very hard, but has a larger syscall interface to compensate, so it’s actually more general. An alternate design of running all code in the parent can be made work if the kernel supports passing a pid handle to every syscall: https://lwn.net/Articles/360747/
On the other hand, this is perhaps a bit more complex in some ways? or at least just different.
http://man7.org/linux/man-pages/man3/posix_spawn.3.html#CONF... says POSIX.1-2001, so I'm guessing between 1998 and 2002 (according to https://en.wikipedia.org/wiki/Single_UNIX_Specification#2001... it started development in 1998 and was released in 2002). That man page also says "since glibc 2.2", which according to Wikipedia was released November 2000.
> The vfork() function shall be equivalent to fork(), except that the behavior is undefined if the process created by vfork() either modifies any data other than a variable of type pid_t used to store the return value from vfork(), or returns from the function in which vfork() was called, or calls any other function before successfully calling _exit() or one of the exec family of functions.
This roughly simplifies to "the only safe thing to do after a vfork is exec", and in this decade you can use posix_spawnp for that instead.
In particular, you can't do safely do anything that you'd normally want to after forking, like closing fds or moving them around (that'd be "calls any other function"). This means that unless the executable you're execing into is also under your control and knows what it's getting into, you have to mangle your fds &c. in the parent, vfork-and-exec, then undo the mess you've just made.
It is in principle _possible_ to navigate this error-prone process without observable side-effects. (The same could be said about implementing speculative execution, and look where we are now.) But I would hesitate to call it correct; the approach is clearly wrong, even if the result looks okay..
The situation is only slightly better if you're targetting some particular OS, which can provide stronger guarantees.
I don't understand how you can claim that vfork can't be used correctly when there are thousands of existence proofs showing that it works fine.
As libc implementer, you're inherently OS- and probably compiler-bound, and you can either fudge it to work on the system you're writing for, or have enough clout to make it so. It's an implementation detail.
As the author of a program which is not strictly OS-bound, you do not have this luxury. Even if you want it, you're probably better off calling clone(2) directly so that you won't get an unintended semantic.
Without counting indirect uses via libc popen/system/&c., I do not believe there are "thousands" of correct uses of vfork. If we count distinct hand-written code paths that call vfork, I'd wager closer to a few dozen, with at best a handful outside of a libc. It's a hideous interface.
After a vfork(), the memory image is shared and the parent is prevented from running. The child is a fully separate process in all other ways. This special vfork state ends when the child does an exit or exec. This is how it works, no matter if POSIX wants to add weasel words or not.
Even with the lame POSIX specification, the "executable you're execing into" is perfectly safe messing with file descriptors. Once that exec happens, the child becomes like any other. Systems with a proper vfork() are also safe before the exec.
It certainly would be nice to have an interface similar to vfork that would take a function pointer and allocate a stack, optionally letting the parent continue on. Linux itself actually can support this, but glibc is uncooperative in exposing it.
Note that these are different things: std::cin is a std::istream, std::cout and std::cerr are std::ostreams, and stdin, stdout, and stderr are FILE *s.
if (pid < 0) { /* TODO: handle error */ }
Reading the actual code I doubt it's just an oversight, the coding style used seems to lend strongly towards not thinking an error is possible: if ((pid = fork()) == 0)
{ //Child
longTask();
} else {
children[i] = pid;
}So does every tutorial which doesn't check the return from printf or check for errors from cout. Did you notice that the parent and child are both writing to the same file descriptor? Maybe the writer should add some words and tests in his code about the inherent race condition there. A true pedant would push this even further and teach their audience how to eliminate this problem with cross process mutexes or semaphores. Better be sure to check the error codes on those calls too.
I know arguing with a pedant is like wrestling with a pig, but honestly I'm more offended by the "#define N 2" than by the lack of error checking.
Please stop wasting electricity and time by writing shell scripts that fork processes to do ridiculously simple things like calling "bc" to add two numbers or "sed" to perform simple string manipulation, since your computer has a "mul" instruction that doesn't require millions of instructions and bytes of memory and disk accesses to perform.
It's like using Bitcoin to pay for a stick of gum.
https://www.quora.com/How-much-energy-does-it-take-to-make-a...
>According to Digiconmist, each bitcoin transaction took 215 kWh. Whereas an average Indian household consume : 90 kWh per month. It sums up that, each bitcoin consumption can power a average indian household for 2.5 months probably. The average electricity cost in India is 5 rupees per kwh.
It's like rolling coal because you want to trigger the libs.
https://www.youtube.com/watch?v=fkmLf5OXxz0
Stop and think before you write that shell script. It's not going to take you any more time to write it in Python, and then you won't be so shocked and surprised and totally fucked when you realize later that you need to add more features and make it more complicated than you originally though it would need to be in the first place.
It just boggles my mind that some people have such a hard time understanding how spectacularly inefficient shell scripts are, on top of how horrible to write, read, understand and maintain they are.
Have you actually measured the amount of energy each takes? I would not be surprised if the Python VM's startup makes it more inefficient.
That is because Python actually has a built-in way to multiply numbers and perform string manipulation, right out of the box, without forking another process! The overhead of the Python interpreter required to get to that "mul" instruction is nothing compared to the overhead of forking another process to run "dc". This isn't rocket science, it's not even close. You really should check it out some time. So do a lot of other languages. Just not bash.
I can't believe I need to explain this. This isn't about Python being amazingly well designed or efficient. It's about bash being extremely badly designed and inefficient, and people having absolutely no clue how expensive it is to fork processes like "bc" to multiply numbers and "sed" to split strings.
$ echo $((1*2))
2
$ string="hello, world"
$ echo ${string/hello/goodbye}
goodbye, world bash-3.2$ echo $((1.5+0.5))
bash: 1.5+0.5: syntax error: invalid arithmetic operator (error token is ".5+0.5")
https://unix.stackexchange.com/questions/40786/how-to-do-int...I gave up INTEGER BASIC for APPLESOFT several decades ago, and never looked back.
(LOADING INTEGER INTO LANGUAGE CARD)
]INT
>PRINT 1.5 + 0.5
*** SYNTAX ERROR
>FP
]PRINT 1.5 + 0.5
2
It's 2019, and you're struggling to do in bash what you could have done on an Apple ][ in 1979, but the syntax you need to use is even WORSE than BASIC, so your code looks like line noise from when your mom picks up the phone while you're using a 300 baud modem. Is that really progress?https://en.wikipedia.org/wiki/Integer_BASIC
>Integer BASIC was phased out in favor of Applesoft BASIC starting with the Apple II Plus in 1979. This was a licensed but modified version of Microsoft BASIC, which included the floating point support missing in Integer BASIC.
https://en.wikipedia.org/wiki/Applesoft_BASIC
>Applesoft BASIC is a dialect of Microsoft BASIC, developed by Marc McDonald and Ric Weiland, supplied with the Apple II series of computers. It supersedes Integer BASIC and is the BASIC in ROM in all Apple II series computers after the original Apple II model. It is also referred to as FP BASIC (from "floating point") because of the Apple DOS command used to invoke it, instead of INT for Integer BASIC.
Notice that there are even machine language instructions for cos, atan, tan, sin, ln, etc. How do you call those instructions in bash?
https://docs.oracle.com/cd/E18752_01/html/817-5477/eoizy.htm...
>x86 Assembly Language Reference Manual: Floating-Point Instructions
>The floating point instructions operate on floating-point, integer, and binary coded decimal (BCD) operands.
https://en.wikipedia.org/wiki/Turing_tarpit
>54. Beware of the Turing tar-pit in which everything is possible but nothing of interest is easy.
But I think that most people who use bash simply have no clue how expensive forking another process is, and not enough experience to realize that software always ends up growing up to be much more complex than they originally expected, and think that "1+1" is as much arithmetic as they will ever need.
https://en.wikipedia.org/wiki/Greenspun%27s_tenth_rule
>Any sufficiently complicated C or Fortran program contains an ad-hoc, informally-specified, bug-ridden, slow implementation of half of Common Lisp.
https://en.wikipedia.org/wiki/Jamie_Zawinski#Principles
>Every program attempts to expand until it can read mail. Those programs which cannot so expand are replaced by ones which can.
You're talking to someone who ran a web farm in the 90s so yeah, I saw the evolution from CGI to mod-perl and on. I mean I know the cost of forking. But come on, aren't you taking your criticism a bit far?
I also had an Apple ][+. Actually, one wasn't available so we Frankensteined one together from an Apple ][ and an Applesoft language card by swapping the motherboard ROMs and language card ROMs.
We both have some perspective from the Apple ][ days, and can appreciate how astronomically fast computers are compared to how slow they used to be.
But doesn't it give you pause (literally and figuratively) when you see people thoughtlessly and casually write code in 2019 that runs slower on a high-end MacBook Pro than the equivalent code would run on an Apple ][ in 1979?
People should take a momemt to reflect on what they're trying to accomplish and if there are better ways of going about it, before writing Rube-Goldbergesqe bash scripts heavily punctuated with line noise, haphazardly stringing together Unix processes and temporary files with cryptically abbreviated commands using mysteriously curt flags and parameters (like "dc -l", or is it "bc -l", and is that dash-ell or dash-one? I can't remember, and we haven't even gotten to the part about how I have to quote and escape the arithmetic expression parameter)!
It's not as if they're making thoughtful trade-offs, and that as a result of choosing to use bash despite all its weaknesses, they're actually developing more robust, easier to read and maintain code than they would have had they used Python.
You're making a false equivalence like "there is slow arithmetic on both sides". No: one side is astronomically slow, requiring millions of instructions, the other side is tolerably slow, requiring hundreds of instructions. There is no comparison.
Again, this isn't about Python. That's just an example. It's about forking or not forking. You know, the topic of this thread.
If you like, I could translate my Python argument to JavaScript, and you could explain how a bash JIT compiler could optimize out all the instructions necessary to fork "bc" to call the cos instruction.
Let me head you off at the pass before you ask "Why would I ever want to calculate a cosine in a bash script?" when you should be asking "Why would I ever want to use a language that doesn't let me calculate a cosine? Or even draw a circle?"
To illustrate my point: Perhaps you want to write a script that draws a pie chart to illustrate your disk space. That's a typical sysadminy scripty thing to do, right?
With bash, you'd have to call out to many other processes to do that. Why is that? Tell me why hasn't anyone come up with a way to draw like the canvas API in bash, without forking off dozens if not thousands of processes?
Don't you think floating point would be useful for calculating disk space usage more accurately than integer percentages, and the canvas api would be useful for drawing pie charts?
If you think you'd never need to do something like that in bash, then you have an enormous blind spot from bash tragically lowering your expectations, imagination, and productivity as a programmer.
Because bash can't even perform the simplest APPLESOFT BASIC floating point math operations or HGR hires graphics HPLOT drawing commands.
So you would first have to fork off to one process to multiply and add a few numbers and calculate a cosine, and then you'd have to fork off to another process to multiply and add a few numbers and calculate a sin (because "bc" can only return one result at a time, thank you). Then you would have to pass those numbers to another program to draw them, probably with a here-file, some loops, a bunch of variable substitutions, temporary variables, intermediate files, lots and lots of string concatenation, converting back and forth between binary floating point numbers and strings, all meticulously decorated with a phalanx of quotes, escapes, double and triple backslashes, back-ticks, colons and exclamation marks before single character codes, and double secret parenthesis.
You're being feisty, whatever.
> We are discussing the overhead of forking processes to do math, which Python doesn't have to do.
Python uses dynamic dispatch on boxed types for every single operation. You're right this /only/ takes hundreds of instructions, but that just shows my analogy is appropriate. That's the turtle, in case you didn't understand it. I think it's silly for you to say this is "tolerably slow" for anyone but yourself. For other people, a fork per operation is completely tolerable. (I'm not one of those people...)
> [...] explain how a bash JIT compiler could optimize out all the instructions necessary to fork "bc" to call the cos instruction.
If anyone cared, bash could implement "bc" as a builtin the same way it does "echo". That'd be ugly, but it seems like the kind of thing the GNU folks would do if it was a common enough use case.
So I'll make an analogy too: A turtle -vs- a snail are the same order of magnitude slow. The Nazis -vs- the protesters are not of the same order of magnitude evil in Charlottesville, just like forking a new process -vs- dynamically dispatching to a function in the same process are not of the same order of magnitude slow in Unix.
And I'll reiterate:
With Python (and JavaScript) it's POSSIBLE and even NORMAL for a JIT compiler to optimize the code to call mul instructions directly. PyPy is a thing, it's free, and it works just fine, thank you:
But it's IMPOSSIBLE to implement a bash JIT compiler (or even AOT compiler) that optimizes out the millions of instructions that are required to make a system call to fork "bc" to call that same "mul" instruction. In Linux, you simply can't optimize out system calls. (Unless Alexia Massalin's Synthesis kernel is giving you a free piggy back ride ;).
https://en.wikipedia.org/wiki/Alexia_Massalin
http://infolab.stanford.edu/~manku/quals/summaries/gribble-s...
And no, integrating "bc" into bash isn't a valid solution, or they would have done it ages ago. How about integrating the canvas API into bash too, while you're at it? Tell me when you convince them to accept your pull request!
Are you intentionally invoking Godwin's law because you know your argument is untenable?
> PyPy is a thing, and it works just fine, you know
I like PyPy. It's not perfect, but it's pretty good. Maybe you should recommend that in the first place next time.
> But it's IMPOSSIBLE to implement a bash JIT compiler [...] to fork "bc"
Never say never. Before V8, everyone thought JavaScript had to be as slow as Python too. Now we have PyPy and LuaJIT, which I would've thought were impossible.
> And no, integrating "bc" into bash isn't a valid solution, or they would have done it ages ago
You seem really hung up on "bc". You do realize that "echo" wasn't always a builtin, and that they added it for performance reasons, right? If anyone cared enough, they could add "bc" as a builtin. Hell, the expression evaluator for integers is well on it's way. You're right that probably nobody cares enough to build in "bc", but I wouldn't be surprised to see a triple paren or square bracket math interpreter with floating point support in bash someday. Heh, maybe they'll even add complex numbers. And a bash JIT is exactly the kind of thing some grad student would do for the absurdity of it all.
> [...] you ask "Why would I ever want to calculate a cosine in a bash script?"
I wouldn't ask that. This kind of argument strategy is silly too. You should choose your own words, and I'll choose mine. Leave the fake dialogs to Aristotle.
> To illustrate my point: Perhaps you want to write a script that draws a pie chart to illustrate your disk space. With bash, you'd have to call out to another language to do that. Why is that? Tell me why hasn't anyone come up with a way to draw like the canvas API in bash, without forking off dozens if not thousands of processes? Don't you think that would be useful?
I routinely use gnuplot as plotter in a child process. This works with bash, python, C++, or any other language which supports popen (that's fork, dup, and exec).
> Because bash can't even perform the simplest APPLESOFT BASIC floating point math operations [...]
I'm not defending bash. I'm saying that putting Python much ahead of it is silly. Why would you want to run a race with either a snail or a turtle? Silly parables about the tortoise and the hare aside, pick a cheetah or something.
> And of course needlessly converting all those numbers back and forth between strings and floating point so many times.
Python tediously checks if the vtable contains the necessary operator, evaluates the function pointer it finds, which checks the other operand's type, unpacks the 8 bytes of doubles from each 32 byte struct, finally performs the operation, then allocates a new 32 byte structure for the return result, which is checked to see if an exception occurred, and then pushed onto the stack. I probably missed some steps, and I certainly ignored the sophomore level interpreter loop overhead.
I mean, yeah that's faster than fork, but I'm not sure it's a huge win over anything else.
This is not "astronomically slow" on any hardware from this millennium.
> Let me head you off at the pass before you ask "Why would I ever want to calculate a cosine in a bash script?" when you should be asking "Why would I ever want to use a language that doesn't let me calculate a cosine? Or even draw a circle?"
Because there are places for languages that can't do everything you'd ever want. Sometimes they can be extra specialized for a common task you do often: for bash, this happens to be "how can I glue together things quickly?"
> Tell me why hasn't anyone come up with a way to draw like the canvas API in bash, without forking off dozens if not thousands of processes?
Because this isn't something that bash would be good at, as you've mentioned. It's a task that would be a lot better suited to some other tool, which bash could then interact with.
How do you call those in Python ;)
Fundamentally, you're making a big deal about something that isn't really that horrible. Not everyone has Python on their system, and a fork() is not a terribly costly operation for a shell script. Sure, it might take an extra microwatt, but arguing that this is the same as Bitcoin mining or rolling coal is an argument in the extreme.
Simple! You shell out to your math-compiler:
https://github.com/skx/math-compiler/
That compiles your math down to assembly language, which has got to make sure they're super-fast, right? ;)
https://en.wikipedia.org/wiki/Incompatible_Timesharing_Syste...
https://dspace.mit.edu/handle/1721.1/6153
https://github.com/PDP-10/its/blob/master/doc/info/ddt.33
https://github.com/PDP-10/its/blob/master/doc/_info_/ddtord....
Sure, I love python, but if I want to find and filter some files nothing is easier to write than bash, so it's possible to have a quite big bash script even doing a good job and being well readable.
I'm honestly a little shocked that this doesn't seem to be common knowledge anymore and your post got upvoted so high.
My main language is Python btw.
> but if I want to find and filter some files nothing is easier to write than bash
That's like 10 lines in Python and much easier to read and test.
It's a single command in bash, invoking a program far more efficient than anything you can put together in python. It is literally a tool written for this exact purpose. I think it is simply ridiculous to not use a pre-existing tool that frankly you cannot out-do with a scripting language.
Even if you are writing a python script, you should _still_ use find to search for files if there is _any_ complexity in the search requirements.
Look. We are working an professional field with a strong scientific foundation. There is no "this bridge is harder to drive over for me therefore I think it's bad architecture". The easiness and expressiveness depend on the problem you try to solve, not on the people, their tastes or their experience.
Shell tools are highly optimized to work in pipes, to work on files and to work on line separated entries like lines of code or lines of CSV. So if you have a problem where you need to pipe file handling into some filtering and into some line based transformations then you will have a hard time to find anything as easy or expressive as a bash script. And most projects on a unix-like operating system at least start like this: Understand the problem based on database dumps, logs, csv data measurements, previously written code etc.
If your regular expressions are not able to cover all the edge cases anymore and/or become unreadably complex then of course you should start an actual program. That only happens a few weeks into most topics though.
> doubly so when you need to throw more of them at this thing written by a former developer that thought he or she was saving money for the org by not spending an extra 30 minutes writing the program in Python or Go.
Yes, there are people who lack the skill or ambition to develop good code. We all hate when we need to deal with their stuff. But honestly, would they have written better code if they had used Python or Java or C++ for that problem? Unlikely.
A good coder would have some feeling for when he outsources some functionality to a python script. And if he develops a code base for a longer time most of his code will be in a real programming language. But even then it might be wrapped into a thin layer of bash to use find or grep or something that cleans up the input a little.
Please don't tell me never to use spaces in file names.
How many human resources are required to track down and fix why a bash script suddenly starts failing mysteriously, years after it was written, when somebody mistakenly dares to use a space in a file name?
That kind of problem never happens with Python scripts (including problems with file names including other punctuation and unicode characters), because Python simply represents file names as strings.
For a scripting language that was indented to manipulate file names all the time, bash sure fucked up big time.
https://www.linuxjournal.com/article/10954
>Work the Shell - Dealing with Spaces in Filenames
>[...] I've outlined three possible solution paths herein: modifying the IFS value, ensuring that you always quote filenames where referenced, and rewriting filenames internally to replace spaces with unlikely character sequences, reversing it on your way out of the script.
>By the way, have you ever tried using the find|xargs pair with filenames that have spaces? It's sufficiently complicated that modern versions of these two commands have special arguments to denote that spaces might appear as part of the filenames: find -print and xargs -0 (and typically, they're not the same flags, but that's another story).
>During the years I've been writing this column, I've more than once tripped up on this particular problem and received e-mail messages from readers sharing how a sample script tosses up its bits when a file with a space in its name appears. They're right.
>My defensive reaction is “dude, don't use spaces in filenames”, but that's not really a good long-term solution, is it?
>What I'd like to do instead is open this up for discussion on the Linux Journal discussion boards: how do you solve this problem within your scripts? Or, do you just religiously avoid using spaces in your filenames?
>Dave Taylor has been hacking shell scripts for a really long time, 30 years. He's the author of the popular Wicked Cool Shell Scripts and can be found on Twitter as @DaveTaylor and more generally at www.DaveTaylorOnline.com.
(Dave wrote that 8 years ago, and there are still zero comments on it. Nobody has a good solution! Bash is just permanently fucked when it comes to file names.)
#!/bin/bash
# run this scri0t with 0 or more filename args
for file in "$@" ; do
do_something_with "${filename}"
done
> Nobody has a good solution! Bash is just permanently fucked when it comes to file nameFrom the man page bash(1):
Special Parameters
@ Expands to the positional parameters, starting from one.
When the expansion occurs within double quotes,
each parameter expands to a separate word.
That is, "$@" is equivalent to "$1" "$2" ...Most people don't test their bash on spaces. Look at the dismissive commenter above who just things "nah why would people use spaces in filenames?". Do you think he quotes everything properly? No chance.
CMake is also really bad about this. Although it is arguable slightly better because it uses semicolons as list delimiters instead of spaces, and semicolons in filenames are going to be very uncommon. I would expect there isn't a single CMake project out there that works if you build it in a path containing a semicolon though!
The same kind of lazy care-free people who love the convenience of PHP's "register_globals", because it "saves typing".
If they're too lazy to use something like Python in the first place, they're probably also too lazy to be conscientious and anal retentive enough to write and test bash scripts that handle spaces in file names properly, or do any kind of user input sanitization.
And they don't think twice about running their scripts as root, either. Or writing blog postings telling other people to copy and paste their code snippets into a root shell, or worse yet download and run them without auditing first, by piping curl into a root shell.
I recently had an Xcode project fail to build (blender, which uses cmake), because I was using a version of Xcode that had a space in its directory name (because I had foolishly added a space followed by the version number to the name "Xcode 9.4.1.app"). Of course that was because of some of the tools and custom build steps were calling out to bash. The problem wasn't in any of the source code or configuration files, it was in the file name of the tool itself! It took hours of frustration poring over machine generated project and makefiles to diagnose the problem from the obscure unhelpful error message, which I will never get back.
In the long term, isn't it more efficient of a lazy person's time to simply use a scripting language that just doesn't have these problems in the first place, and doesn't require you to preemptively bend over backwards and dodge silent invisible bullets all the time?
It's pretty ridiculous that a scripting language intended to be used to operate on files falls flat on its face and is so difficult to use when you have something as common and innocuous as files with spaces in their names (or worse yet, file names that begin with dashes -- think of the security implications of that: not something a lazy careless bash scripter would ever pause to reflect about)!
It really is an extremely common and unsolved problem. Notice how each "better answer" is actually worse than the last and comes with its own host of problems and edge cases and work-arounds:
https://unix.stackexchange.com/questions/9496/looping-throug...
https://stackoverflow.com/questions/3967707/file-names-with-...
https://www.cyberciti.biz/tips/handling-filenames-with-space...
https://www.tecmint.com/manage-linux-filenames-with-special-...
https://www.hecticgeek.com/2014/02/spaces-file-names-command...
$ find . -type f
./foo bar
./--stupid-filename
./-annoying dir name!/--very annoying path...
$ echo "pid=$$"
pid=31281
$ while IFS= read -r -d '' file ; do
> echo "[pid=$$] do something with '${file}'"
> list[i++]="$file"
> done < <(find . -type f -print0)
[pid=31281] do something with './foo bar'
[pid=31281] do something with './--stupid-filename'
[pid=31281] do something with './-annoying dir name!/--very annoying path...'
$ declare -p list
declare -a list='([0]="./foo bar" [1]="./--stupid-filename" \
[2]="./-annoying dir name!/--very annoying path..." [3]="./iter.sh")'
Using the character literal $'\0' also works as the delimiter, but Bash doesn't support sending null bytes in args. In bash, $'\0' is the empty string: $ x=$'\0'
$ declare -p x
declare -- x=""In any case I don't say put in a Bash script and let it grow to ten million lines over 20 years.
My approach to bash scripts is this:
1. start with a bash script to understand a problem and see how a solution might be structured. PoC level is enough here.
2. When you know you solve the right problem and the customer is willing to make something bigger, then put ~80% of the code in a good programming language (Python is my personal favorite, but Ruby, Java, PHP, C#, C++, Lisp might also work). That happens about 3 months into a project.
3. If you start to hit bottlenecks and performance ceilings put the 30-50% that are relatively stable and most performance hungry into C/C++ code and really optimize the heck out of it. This should be done only with an aged project, like 2 years in, and very pessimistically, since highly optimal code usually is badly readable and badly maintainable. You certainly don't want to go into that code 6 years later and drastically change its API to your python code.
So a ten year project would have a thin layer of bash/docker at the top, a strong core of highly flexible code (of Python for me as well) in the middle, and a leg of high performance C code here and there.
I honestly have never seen a project that really needed the C-layer though. In modern IT world it seems most problems are old at least from a political/management point of view between 2 weeks and 3 months. So in fact, although I count myself as a Python coder, the last few years I've spent writing way more bash and docker related stuff, which then is either never touched again or directly thrown away after a few weeks, no matter if successful or not.
PS: Please don't use spaces in file names. Not just for scripts. It is also not clear for users where something ends, and they are even less used to putting string indicators around file names than you are about handling spaces with bash there.
And last but not least stuff like spaces in filenames, starting scripts with upper case letter etc, they might even be just a small annoyance. But you can be sure that skilled developers will put you 2 classes below their average when they see this, because it simply shows a lack of the right spirit.
This is not true, and I do not see why it should be like that. You cannot, for example, have slashes in filenames; and for good reasons. Disallowing spaces in filenames solves a lot of problems.
There used to be a bug in the Gatorbox Mac Localtalk-to-Ethernet NFS bridge that could somehow trick Unix into putting slashes into file names via NFS, and Unix would totally shit itself when that happened. That was because Macs at the time (1991 or so) allowed you to use slashes (and spaces of course, but not colons), and of course those silly Mac people, being touchy feely humans instead of hard core nerds, would dare to name files with dates like "My Spreadsheet 01/02/1991".
https://en.wikipedia.org/wiki/GatorBox
I just tried to create a file name on the Mac in Finder with a slash in it, and it actually let me! But Emacs dired says it actually ended up with a ":" in it. So then I tried to create a file name with a colon in it, and Finder said: "Try using a name with fewer characters or with no punctuation marks." Must be backwards compatibility for all those old Mac files with slashes in their name. Go figure!
If you think nobody would ever want to use a space or a slash in a file name, then you should get out more often and talk to real people in the real world. There are more of them than you seem to believe!
I think that there are more good reasons to disallow spaces than slashes. For one, it is very useful to have a natural character to separate filenames in a string. Then your bash scripts can be simpler! :)
> trick Unix into putting slashes into file names via NFS, and Unix would totally shit itself when that happened
LOL. I love this stuff. There is a similar trick that you can play today to wreak havoc on osx systems. You commit into a git repository a new file whose name only differs in case with an existing one, and ask a mac user to update the repo. Hilarity ensues.
> If you think nobody would ever want to use a space or a slash in a file name, then you should get out more often and talk to real people in the real world. There are more of them than you seem to believe!
But these users should not see the filename anyway, but the "title" of the document, which can be stored inside the file and can have whatever character they want. At the minimum, the user-facing interface can convert user-typed spaces into another character (e.g., a unicode non-breaking space) when writing the file into the filesystem. Raw spaces in filenames are an abomination that make shell scripting unnecessarily difficult. I would love a mount option that disallow the appearance of such things in my disk.
So you think all files should have a human readable title inside of them, huh? Then you're OK with MICROS~1 8.3 file names, and not ok with any file format that doesn't allow a human readable title inside? Should empty files be prohibited too?
I do agree though that filesystems are not the most user friendly abstraction of data. Maybe we can do better than that.
But also filesystems have the huge advantage that they are somewhat user friendly while also being somewhat technology friendly. Therefore I'm also not too mad about their existence. There are a lot of end user problems that can be solved quickly exactly because they just need a professional one time query and FS+bash+clis are such a good combo to craft that query.
https://en.wikipedia.org/w/index.php?title=Comparison_of_fil...
(Spoiler: Note the last column. Also: https://en.wikipedia.org/wiki/Hans_Reiser ...)
And good for them! What kind of savage puts spaces on their filenames?
https://en.wikipedia.org/wiki/Whitespace_(programming_langua...
I also use spaces in the file names of all my C++2000 source code that takes advantage of the wonderful new whitespace overloading feature -- it's conventional to use the ".cpp " and ".h " suffixes with trailing spaces.
Now that the modern compilers and interpreters (clang, python3, etc) allow unicode in variable names, the next step is certainly to accept spaces in variable names. With a slight modification, the python language can be adapted to support spaces inside variable names.
After all, people have the right to name their variables however they want!
https://en.wikipedia.org/wiki/Non-breaking_space#Width_varia...
For simpler, more expressive Python system scripts use Plumbum: https://plumbum.readthedocs.io/
$ cat foo.sh
#!/bin/bash
a="1.5"
b="0.5"
sum="$(echo "${a} + ${b}" | bc)"
echo "The sum of ${a} and ${b} is ${sum}"
$ /usr/bin/strace -q -D -w -c ./foo.sh
The sum of 1.5 and 0.5 is 2.0
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
55.53 0.001225 123 10 read
9.75 0.000215 215 1 execve
5.53 0.000122 7 18 mmap
4.03 0.000089 8 11 3 open
3.04 0.000067 7 10 mprotect
2.99 0.000066 5 13 rt_sigprocmask
2.54 0.000056 56 1 clone
2.45 0.000054 5 10 close
2.40 0.000053 5 11 rt_sigaction
1.81 0.000040 5 8 fstat
1.41 0.000031 8 4 lseek
1.22 0.000027 5 5 3 stat
0.91 0.000020 20 1 write
That was 56us to clone(2) the process and 215us to execve(2) 'bc'. All together just 12.74% of the time spent running syscalls, and an even smaller portion of the total time to launch and execute this simple example. Most of the time is sent launching bash itself at the beginning. Using a language such as python or launching a custom binary will have similar costs to to launch the program. If we're talking about a single persistent process (e.g. an interactive shell), then we're still talking about paying just 0.25ms of CPU time to perform a useful calculation. Compared to intentionally wasting time by design maximizing hash prefix guesses to "mine" Bitcoin "ain't the same fuckin' ballpark, it ain't the same league, it ain't even the same fuckin' sport"[1].However, the actual cost to launch an external program like 'bc' isn't particularly important. It could be multiple orders of magnitude slower and it would still be acceptable because we're talking ab9ut a cost that only needs to be paid a few times. Yes, compared to "full" programming languages like python, running an external program to perform a simple calculation might seem like a wasteful, half-assed solution. As Gary Bernhardt explained[2], "Half-assed is OK when you only need half of an ass."
I don't care about the cost of running 'bc' to do simple floating point calculations because I usually don't need to do floating point calculations (which are a lot harder to use safely[3] compared to simple integer math) for the type of task I would implement in a shell script. If the task required more than a few calculations (or other complex features such like a GUI), using something other language like python/ruby/C/rust would be a better tool than shell script. Learning how to choose the appropriate tools to use when solving a problem is an important skills for any engineer.
Shell scripting languages like Bourne (sh) are designed to be a glue language that you use for simple (often interactive) tasks. They are the user interface (or "IDE") of Unix. Asking why shell d9oesn't have built-in support for efficient floating point math or complex GUI features is kind of like asking why Eclipse or Visual Studio don't have a built-in spreadsheet or email client. Yes, you could probably make that happen as a half-assed hack, but it's obviously not the type of task those tools were designed for.
As for the syntax, if you find it confusing or ugly and prefer to use other tools, that's fine. There are a lot of great languages available, so use the tool that is best for you. If it had #!/bin/sh art the top, you can probably substitute any language you prefer without and problems. However, remember that for some of us Bourne shell is the easy, reasonably efficient tool we prefer to use for some problems.
[1] https://en.wikiquote.org/wiki/Pulp_Fiction#Dialogue
[2] https://www.youtube.com/watch?v=sCZJblyT_XM
[3] https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.h...
And therein lies the rub. Let's put the donkey before the cart: The reason you would not do many kinds of task in a shell script, is that shell scripts are not any good for many kinds of tasks, not because you are wrong to want to to some kinds of tasks.
The term "shell script" is an artificial construct, just like the shell scripting language itself. There is no reason you need to limit your imagination, aspirations and abilities by defining yourself as a "shell scripter" who cannot add 1.5 and 0.5 to get 2, and who cannot draw a pie chart.
Of course you should learn other languages, then you won't need to write shell scripts. Their disadvantages vastly outweigh their advantages in important ways specific to scripting and unix administration (like dealing with file names with spaces, or performing arithmetic and statistics, for examples).
Shell scripts are good for gluing together tools that do disparate things and simple automation. Most shell scripts end up becoming "real programs" once they start containing actual logic in them.
When you have a pipeline between two nontrivial programs, for example, or any kind of fd manipulation.
Python and JavaScript are also good for gluing together tools that do disparate things and simple automation. But they can do arithmetic and graphics, too! And without the enormous overhead of forking a process to perform simple arithmetic and string manipulation.
Bash scripts aren't good for many useful tasks because bash is badly designed, not because there's some universal law of computing that proclaims that scripts should never use floating point numbers or draw graphics.
Apple figured all this out in 1979 when they made a deal with Microsoft and switched from INTEGER BASIC to APPLESOFT BASIC, and added floating point and HIRES graphics. Even Apple ][ DOS 3.2 could deal with file names with spaces in them! How long will it take the Unix community to figure that out, too?