First: text is not well defined. Is it ASCII? Is it UTF-8? Some programs can spew UTF-32 with proper locale configured, it's a mess.
Second: encoding and decoding of objects to text is not defined at all. Those problems with filenames is just one example. Using newline as a separator is a natural thing that is easy to implement, yet it is wrong.
In my opinion two things should be done:
1. Standardise on UTF-8. No other encodings allowed.
2. Standardise on JSON. It is good enough to serve as universal exchange format, tools like `jq` exist for some time now.
So any utility must read and write JSON objects with some standard env set. And shells can be developed with better syntax to deal with JSON. This way you can write something like
`ps aux | while read row; do echo ${row.user} ${row.pid}; done`
Please don't use that underdefined joke of a spec. Define "PosixJson" and use that instead. Right now it's not even clear what the result of parsing {"a": 1234678901234567890} is. Is this a parse error? A bigint? A float/double? Quiet wraparound? Something else? I've seen all these behaviors in real world JSON implementations across different languages.
See https://pubs.opengroup.org/onlinepubs/9799919799/basedefs/V1...
> 3.387 Text File
> A file that contains characters organized into zero or more lines. The lines do not contain NUL characters and none can exceed {LINE_MAX} bytes in length, including the <newline> character.
So, if you have some non-printable characters like BEL/␇/ASCII 0x07, that's still a text file.
(and I believe what bytes count as a valid character depend on your `LC_CTYPE`).
But the moment you have a line longer than {LINE_MAX} bytes (which can depend on which POSIX environment you have), suddenly your text file is now a binary file.
> 3.185 Line
> A sequence of zero or more non-<newline> characters plus a terminating <newline> character.
So a file with some characters but no trailing newline is reported by `wc -l` as having zero lines.
The definition should read "one or more lines" instead or (probably better) specify that a text file contains "zero or more characters".
You never have to worry about whether you're dealing with ASCII vs. UTF-8, but rather if you're dealing with UTF-8 vs. ISO-8859-1, or worse, Shift JIS or similar.
% java -Dfile.encoding=UTF-32 Test | hexdump -C
00000000 00 00 00 48 00 00 00 65 00 00 00 6c 00 00 00 6c |...H...e...l...l|
00000010 00 00 00 6f 00 00 00 2c 00 00 00 20 00 00 00 77 |...o...,... ...w|
00000020 00 00 00 6f 00 00 00 72 00 00 00 6c 00 00 00 64 |...o...r...l...d|
00000030 00 00 00 0a |....|
00000034
From quick googling it seems that glibc does not support it, so it should not happen.`iconv` does, and this is enough in common. Among with tons of eerie EBCDIC/whatever...
I regularly use `iconv -t utf-32be | hd` to look what a bizarre sequence is denoting yet another weird symbol like an itchy hedgehog.
And what is a real reason to disallow this?
Command line scripting is supposed to be adhoc and hack.
But then, how well would it work for ad-hoc usage, which is probably one of the biggest uses of shells?
If you can't get the immensity of the cleverness of Unix foundations, you should not talk about them.
That idea is what made it possible for you to type that sentence in the first place.
I haven't seen any other tool with so much general utility and availability.
> to loop over a string array and print its contents
Is incredibly easy in bash and bash like shells. As highlighted the issue is that tools like 'ls' don't create "a string array." They create one giant string that has to be parsed. The rules in the shell are different than in other languages but it /will/ do most of the parsing for you, or all of it, if you do it carefully.
This is a fine tradeoff. As evidenced by it's wide usage and lack of convincing replacements.
> availability
That's the real reason why we use Unix shell. It's not good, but it's available. Like a cheap hooker.
> but it /will/ do most of the parsing for you, or all of it, if you do it carefully.
"It mostly works if you're careful" doesn't sound very convincing to me.
Username checks out.
Would you rather write your own parser?
I tried both python and lua interactively, but they are a pain when it comes to handling files. You have to type much more to get the same things done.
Five people around the globe isn't wide use.
Human psychology is fascinating!
Bash-alternatives that are not completely compatible frankly just don't have a chance.
Why they force the restriction at syscall level?
Python and lua are pretty close to that.
Python maybe often installed by default but it's definitely not an essential/required package "out of the box" on every install. Also, in a thread where one topic is how POSIX shell handles whitespace in filenames, it's hilarious (not in a good way) that someone suggests a language that handles whitespace the wrong way in it's own code. Yes, significant whitespace is objectively wrong.
What OS/distro is Lua included on out of the box? That doesn't mean "available in a package". I mean literally included in every single install and cannot reasonably be omitted?
Regardless of the availability, the parent comment says
> better by every objectively measurable metric
Neither Python nor Lua are "better" than shell, at the types of things shell is commonly used for - they're objectively worse.
You should use the word "objectively" less.
Just FYI, there are UNIX-like, POSIX compatible systems that are not a Linux distro.
> rpm or pipewire depend on lua. Ubuntu and Debian ship with pipewire per default.
Pipewire? Do you mean this? https://packages.debian.org/bookworm/pipewire
That isn't even close to "installed on every system". Best I can tell from the reverse dependencies, it's required for some Gnome Remote Desktop tool, and best I can tell, it doesn't rely on Lua anyway (at least on Debian).
> You should use the word "objectively" less.
I specifically used the word objectively, because the original comment that I replied to, said this:
> better by every objectively measurable metric
Pipewire being the Pulseaudio replacement from Redhat.
Bookworm is probably the last Debian without :P
Right, so it's a desktop package that ultimately will be installed on about 1% of all Linux machines because the vast majority are servers without a desktop environment.
Also worth pointing out: liblua on Debian at least, is the shared library. It's not the binary to execute standalone Lua scripts.
Check your own installs and tell me if you find some that dont have liblua or libluajit.
For the library thing: I said "Python and lua are pretty close to that." earlier. I did not say that they have interpreters ready everywhere. But if the language core is already installed on a large fraction of machines, then adding the interpreter is not a big cost.
So far you've presented no evidence of this though, just that it's used by a new desktop-focused package.
All linux desktops over the last 30 years is not even a "large fraction" of total Linux installs, much less the ones that have already migrated to this new audio system.
> adding the interpreter is not a big cost
It's nothing to do with cost. It's about "how do I know this will absolutely 100% run on any POSIX machine I throw it on without any extra steps".
Remember the argument here is about something that is claimed to be "objectively better" than Shell. The ubiquitous nature of POSIX shell is a huge barrier for any possible competitor, and saying "well you just need to install it" just defeats the purpose. You might as well write it in fucking java and say "well you just need to install a JVM".
Edit to Add: a good number of systems I manage do have liblua installed... because HAProxy requires it, and those systems have HAProxy installed. Not because it was installed as part of the base OS or even a default group of packages.
Incidentally, HAProxy and thus liblua were installed on those systems by infrastructure management that's implemented as shell script. So what kind of chicken and egg argument do we need to have here about how exactly I can run a Lua script to install Lua?
/thread
Compatible with most bash scripts
Compared to what?
It's a lang for an interactive shell, typing literally translates to developer speed. I understand the want for clarity and maybe that's nice in large scripts, but the main goal is to be a shell. So, optimize for that. Also, you probably shouldn't be using powershell for large scripts anyway.
The only recent lang I've seen that has a handle on this is Rust. You can tell they put a lot of thought into having keywords be as short as possible while still being descriptive.
My God next you will say getopt() --longform is the bestest
$ ll /bin/dash
-rwxr-xr-x. 1 root root 113536 Nov 5 2018 /bin/dash
I happen to have an old powershell installed: $ rpm -qi powershell | grep Size
Size : 126588370
A strict POSIX shell is always going to be vastly smaller, for many reasons.I would prefer that the POSIX shell was an LR-parsed language, but you can't have everything.
Dear anal_reactor, what is a "string array"? I have used unix shells since nearly 30 years and never heard about them. And I consider myself a script-fu master!
There are two array-like constructions in the shell: list of words (separated by spaces) and list of lines (separated by newlines). Both cases are implemented as a single string, and the shell makes it trivial to iterate through its components.
I’m not saying that we can’t improve, but I’m more in favor of making the tool more apt to solve a problem than making it easier to learn. Because the latter often wants to forego the requirement of understanding the problem space.
Compare this to the newer programming languages where you explicitly call something with speaking names like .Trim(), .EndsWith(), support from compiler and IDE.
In my experience automation and general programs often are the same thing once things get more complicated. Bash scripts usually grow rapidly and are a giant PITA to maintain or refactor. Throw in build systems and helper scripts and you quickly receive a giant pile of spaghetti. Personally I just switch to one the mentioned programming languages once it goes above a simple sequence of operations.
Personally I don't see how to improve it much without becoming a full blown programming language, at which point it would probably make more sense to just release a library for common automation tasks that is also composable. Maybe I'm just not the right target audience.
And for bigger automation projects, there are lots of projects and programming languages that can help.
I would even make the case for expert tools being as unsurprising and familiar as possible unless there is a very good reason for them not to. Also they should be robust against misuse and guide the user towards good practices. There are always beginners, people that rarely need to use it, people that do programming as "just a job" and people that make mistakes because they are distracted, tired or just human. Something like "rm -r /" is a good reminder of that for many people.
Plus there are already a lot of tools required. Reading a book about every tool I have to use would be unpractical for most projects. Maybe more expert tools should just be tools. The same way I can now just use Ubuntu and get a working desktop system including drivers for most common hardware. If I compare that to the past where I installed a Linux distribution and then found out I lack a driver for my network card but I need to download it from the internet... I still can modify my system if I need to, but it's nice that I don't have to. I think we can do similar things with many parts of development and free some capacity for other tasks.