A peculiarity of the GNU Coreutils version of 'test' and '['
utcc.utoronto.ca
utcc.utoronto.ca
Oh how times have changed:) But even with the demise of most commercial unixen, I was under the impression that Darwin preserved the ancient tradition of installing the system and then immediately replacing its coreutils with GNU?
sed -i '' 's/foo/bar/g' # BSD sed
sed -i 's/foo/bar/g' # GNU sedThere's far less variance between Perls on different systems than there are for sed/awk/shell. So if you want performant portable code, Perl does better than all of those.
I'd never use Perl for a 'big' program these days but it still beats the crap out of the mess of sed/awk/bash/ksh/zsh.
I don't even have it on my Linux system any more; for better or worse, fewer and fewer things use it.
If you make GNU coreutils the first binaries in PATH, you can expect subtle, _very nasty_ issues.
The last one I encountered was `hostname` taking something like 5 seconds to run. Before that, there was some bug with GNU’s stty which I don’t remember the specifics of.
echo foo | sed -i pattern
echo foo | sed -i .orig
But you can still use "-i.orig"; that should work on both. That will leave you with a .orig file to clean up, but arguably that's not a bad thing as -i can clobber files.(Thus continuing and extremely long tradition of layering GNU over the vendor tools; ex. Sun had their own tools, but everyone liked the GNU versions better.)
Possibly the same type of people that target bash specifically in shell scripts I guess.
For example, I managed to create a surprisingly good test suite with bash and the rest of the GNU coreutils:
https://github.com/lone-lang/lone/blob/master/scripts/test.b...
It even runs in parallel. Submitted a patch to coreutils to implement the one thing it couldn't test, namely the argv[0] of programs. I should probably go check if they merged it...
Assuming you meant `shellcheck`: you know it works for POSIX compatible shell too right?
What's wrong with bash is that people invariably end up requiring features that aren't available in some deployed version, and then you've lost a lot of the benefit of writing in shell in the first place.
Similar to shellcheck, shunit2 works just fine for running unit tests in posix compatible shell.
I did. Edited my comment, thanks. I don't know why I'm so prone to that particular typo. My shell history is full of it.
> What's wrong with bash is that people invariably end up requiring features that aren't available in some deployed version, and then you've lost a lot of the benefit of writing in shell in the first place.
But I've gained quite a lot too. Bash has associative arrays. I just can't go back to a shell that doesn't have that.
Shell scripting makes it simple to manage processes and the flow of data between them. It's the best tool for the job in these cases. So there are still reasons for scripting the shell even if one is willing to sacrifice portability.
Now how long are we going to have to wait until somebody invents a way to do named parameters? That will revolutionize the computer industry! I guess it's way to much to ask for a built-in json or yaml parser, all we can hope for is maybe a stringly typed sax callback based xml parser after another 20 years from now, because dom objects in a shell would be heretical and just so unthinkably complicated.
Why are people so afraid to just use Python? Shell scripting and cobbling together ridiculously inefficient incantations of sed, awk, tr, test, expr, grep, curl, and cat with that incoherently punctuated toenail, thumbtack, and asbestos chewing gum syntax that inspired perl isn't ever any easier than using Python, especially when you actually need to use data structures, named function parameters, modules, libraries, web apis, xml, json, or yaml.
Here's the summary of a discussion I had with ChatGPT about "The Ultimate Shell Scripting Language", in which I had it consider, summarize, and draw "zen of" goals from some discussions that I feel are very important (although I forgot about and left out PowerShell, that would be a good thing to consider too -- when I get a chance I'll feed it the discussion you linked to, which made a lot of important points, and ask it to update its "zen of" with PowerShell's design in mind):
The Zen of Python.
https://peps.python.org/pep-0020/
Discussion about Guido van Rossum's point that "Language Design Is Not Just Solving Puzzles".
http://lambda-the-ultimate.org/node/1298
https://www.artima.com/weblogs/viewpost.jsp?thread=147358
Discussion of Ousterhout's dichotomy.
https://en.wikipedia.org/wiki/Ousterhout%27s_dichotomy
https://wiki.tcl-lang.org/page/Ousterhout%27s+Dichotomy
Email from The Great TCL War Part 1, started by RMS's "Why you should not use TCL".
https://news.ycombinator.com/item?id=12025218
https://vanderburg.org/old_pages/Tcl/war/
Email from The Great TCL War Part 2, started by Tom Lord's "GNU Extension Language Plans".
https://vanderburg.org/old_pages/Tcl/war2/index.html
Summarization of the important points in those discussions that apply to The Ultimate Shell Scripting Language, resynthesized into a "zen of" list.
Discussion of Support for Declarative and Procedural Paradigms and how it applies supporting standard declarative syntaxes including json, yaml, and xml (which bash still doesn't and probably never will, and PowerShell does of course).
Suggestions for some more "zen of" and design goals, specifically focused on addressing weaknesses or design flaws in popular languages like sh, bash, tcl, python, perl, etc.
Discussion of how "Simplified Debugging and Error Handling" can be balanced with "Macro Processing and Syntax Flexibility", better than debugging C++ templates, Lisp macros, TypeScript code compiled to minified JavaScript, etc.
Discussion of how JavaScript / TypeScript fall short of those goals.
Discussion of visual programming languages for interactive shell scripting as well and general purpose programming.
Discussion of layering visual programming languages on top of textual programming languages.
https://donhopkins.medium.com/the-shape-of-psiber-space-octo...
Interoperability of text and visual programming languages with LLMs for efficiently and reliably analyzing and generating code.
https://docs.google.com/document/d/1QJ98QwC2ubsTNKOFAzUAw6Zy...
Discussion of using Python to build a visual data flow node based shell scripting language on top of Blender (that just happens to support 3D, image processing, video editing, GPU programming, machine learning, Python module integration, and everything else that Blender is great at).
https://www.youtube.com/watch?v=JOeY07qKU9c
Discussion of how to make efficient use of token budgets when using LLMs with text and visual programming languages.
Here is the condensed discussion with the bulk text I had it analyze omitted, so you can more easily read the summaries and recommendations and "zen of" manifestos:
https://docs.google.com/document/d/1wKhdEoLWCZX9TNaftQxLp6ot...
Here is the entire unexpurgated discussion including all the email messages and articles I had it analyze, if you want to see what it was considering (they're interesting discussions to read in their entirety if you're interesting in that kind of stuff, but ChatGPT is excellent at summarizing them and highlighting the important and relevant points):
https://docs.google.com/document/d/1RTxeYjZ2vZsNU4xTj0ZhiXXj...
With the rise of LLMs I think now is a great time to reconsider the design of text and visual scripting and programming languages, with LLM interoperability and compatibility in mind.
ChatGPT:
Designing a text representation for a visual programming language that interoperates well with Language Models like GPT necessitates careful planning to ensure the representation is both meaningful and efficient. Here are several considerations and strategies to optimize the text representation:
Token Efficiency: The text representation should be concise to make efficient use of the token budget of the language model. Avoid verbose syntax and prefer shorthand notations where possible.
Syntax Simplicity: A simple and consistent syntax helps the language model predict and generate code more accurately. Overly complex syntax can lead to increased errors and confusion for both the model and the user.
Standardization: Using standardized syntax and well-known programming patterns can help the language model better understand and process the code since it has likely been trained on similar patterns.
Semantic Density: Each token should carry a significant amount of meaning. This can be achieved through the use of expressive keywords or operators that encapsulate complex actions.
Modular Design: Break down the text representation into modular components. This approach allows the language model to process each part in isolation, reducing the complexity and cognitive load.
Contextual Keywords: Use keywords that clearly indicate their function within the language, making it easier for the language model to provide contextually relevant suggestions.
Window Size Awareness: Design the text representation to be legible and understandable within a typical window size, avoiding lines that span too long and require horizontal scrolling.
Minimize Ambiguity: Ambiguity in text representations can lead to multiple interpretations by the language model, resulting in inaccurate predictions. Clear and unambiguous syntax helps maintain consistency.
Commenting and Documentation: Incorporate support for comments and documentation within the text representation. This not only aids human understanding but can also provide language models with additional context for generating code.
Escaping Visual Elements: For elements that are inherently visual and do not translate well into text, provide an escape mechanism that allows these elements to be described textually in a way that the language model can process.
Alignment with Model's Training Data: Align the text representation with the structures and syntax that the language model has been trained on. This can include following common coding standards and naming conventions.
Iterative Refinement: Use an iterative approach to refine the text representation based on feedback from testing with the language model, ensuring that the design evolves to better suit the model's capabilities.
By considering these aspects, you can create a text representation for a visual programming language that is both language model-friendly and conducive to an efficient programming experience.
I got burned by supposedly "portable" scripts by the whole Ubuntu dash disaster, and again when I started using Mac OS X, and then again once I used Cygwin and msys2 on Windows.
I do keep portability in mind when writing shell scripts to ease porting later, but without testing there's really no way to be sure "/bin/sh" is right. And some of the Bash features such as arrays are legitimately useful.
I was referring specifically to interactive use: ie why someone would install bash to use as their shell, when zsh is already there.
As a Unix graybeard I always place options first. Options last feels like Windows command prompt, so nothing I want to see...
I always tell younger colleagues who place options in the end, it might work with some commands, but just don't do it. I did not know that "some" includes all of GNU coreutils? A single common code style is a virtue, even in interactive use if there are onlookers. So I guess I will continue to point it out.
An interesting factoid is that FreeBSD/macos sort(1) was using GNU code until recently, since this is quite tricky to implement. Eventually it was reimplemented for GPL avoidance reasons.
We do consider macos though, and ensure all tests pass on macos for each release
test, [, and [[ (2020) - https://news.ycombinator.com/item?id=38387464 - Nov 2023 (225 comments)
A few months back I noticed '[' under /bin on my mac. I tried googling to understand what it was, but my google-fu came up short. It felt like one of those ungoogleable things. This link is an excellent starting point for me.
by the name, you'd think you could just use [ 0 -gt 1 <enter>
after all, we don't have to type grep query perg
if [[ $0 == "norg" ]]; then
gron --ungron "$@"
And so onThe invoked binary has no way of aborting script execution. All it can do is barf out on stderr and return an error code, which the shell interprets as false and `if [ "x" = "x"` (without ]) goes into the else branch.
Consider this example:
for i in $(seq 1 10); do
for f in /*; do
if [ $f = /etc ]; then
: # Do nothing.
fi
done
done
I have 23 entries in /, so this will run [ 230 times – not a crazy number. % time sh test
sh test 0.00s user 0.01s system 91% cpu 0.009 total
bash, dash, zsh: they all have roughly the same performance: "fast enough to be practically instantaneous".But if I replace [ with /bin/[:
% time sh test
sh test 0.11s user 0.38s system 96% cpu 0.509 total
Half a second! That's tons slower! And you can keep "fast enough to be practically instantaneous" for thousands of files with the built-in [, whereas it will take many seconds with /bin/[.(Aside: if I statically link [ it's about 0.4 seconds).
Of course if the script were to handle all files of a filesystem with many small files it could get disturbingly slow. I don't deny that there are cases were it matters. But in over 90% of the scripts I write or use it doesn't.
It's up to you whether half a second is "fast enough" (just an example: can easily also be 2 seconds, or 5 seconds), but it's definitely a lot slower than 9ms and not "only much faster in theory", or "unnoticeable in practice", and "number crunching" doesn't come in to play regardless.
Since the cost of this optimisation is minimal (you can just use the source of /bin/[ as a builtin) I don't see why anyone would choose half a second over 9ms.