If programs used typed data, they'd still need the option to output text to present results in a format the user can understand. To do this, a negotiation protocol could be established, like skissane said. This, in my opinion, is BAD, because then there's the possibility or probability that they'll be differences in the information conveyed in the different formats.
I believe that the use of plain text as the universal format of communication between programs is one of the greatest design decisions for Unix and CLI.
You'd standardize it once via an RFC and you'd be done with it.
I don't think your problem is as big as you say it is.
The real problem is that this pile of code we already have kind of works and it's already making trillions for its users. Changing the whole ecosystems would cost millions and millions, for only a very long term and unclear benefit.
In other words, worse is better.
As oblio said, you can come up with a standard conversion of structured data to text. The other way round you need write a parser for every textual output format, and typically people come up with fragile ad-hoc parsers that don't deal with edge cases properly.
Of course, who am I kidding, in real life we have some sort of crappy text interface which is half-baked both for humans and for machines. But we've been using it for almost half a century and it's too widespread to redo, so there we are, plowing through it daily.
1) We don't need to depend on each individual program's programmer to present every control and information consistently between the 2 interfaces.
2) Automation matches normal, manual use. Just put what you normally do on the command line in a file and you're done. There's no need to look through documentation on how to do what you so frequently do, only in a manner that you rarely do.
Also there is no need really to change anything (as in, breaking existing scripts). Selecting a different output format can simply be a command line option. Many tools already offer Json, XML or CSV output. But since development of those tools is so decentralized you'd be hard pressed getting them all to agree on one. But theoretically you can pick any tool you want right now, add --json support and submit a patch.
Are TUIs like htop, tmux, vim, emacs, less, etc. going to be impossible now, or will you do the negotiation protocol? Both options suck.
When programs have both normal output and errors intermixed are the objects going to be intermixed in the output that's presented to the user? For example, if you do a `find /etc/pacman.d`, instead of:
/etc/pacman.d
/etc/pacman.d/mirrorlist.pacnew
/etc/pacman.d/gnupg
/etc/pacman.d/gnupg/trustdb.gpg
/etc/pacman.d/gnupg/crls.d
find: ‘/etc/pacman.d/gnupg/crls.d’: Permission denied
/etc/pacman.d/gnupg/.gpg-v21-migrated
/etc/pacman.d/gnupg/private-keys-v1.d
find: ‘/etc/pacman.d/gnupg/private-keys-v1.d’: Permission denied
/etc/pacman.d/gnupg/tofu.db
/etc/pacman.d/gnupg/openpgp-revocs.d
find: ‘/etc/pacman.d/gnupg/openpgp-revocs.d’: Permission denied
/etc/pacman.d/gnupg/gpg.conf
/etc/pacman.d/gnupg/secring.gpg
/etc/pacman.d/gnupg/pubring.gpg~
/etc/pacman.d/gnupg/pubring.gpg
/etc/pacman.d/mirrorlist
will you have: [
"/etc/pacman.d",
"/etc/pacman.d/mirrorlist.pacnew",
"/etc/pacman.d/gnupg",
"/etc/pacman.d/gnupg/trustdb.gpg",
"/etc/pacman.d/gnupg/crls.d",
[
{
type: "Permission denied",
path: "/etc/pacman.d/gnupg/crls.d"
},
"/etc/pacman.d/gnupg/.gpg-v21-migrated",
"/etc/pacman.d/gnupg/private-keys-v1.d",
{
type: "Permission denied",
path: "/etc/pacman.d/gnupg/private-keys-v1.d"
},
"/etc/pacman.d/gnupg/tofu.db",
"/etc/pacman.d/gnupg/openpgp-revocs.d",
{
type: "Permission denied",
path: "/etc/pacman.d/gnupg/openpgp-revocs.d"
},
"/etc/pacman.d/gnupg/gpg.conf",
"/etc/pacman.d/gnupg/secring.gpg",
"/etc/pacman.d/gnupg/pubring.gpg~",
"/etc/pacman.d/gnupg/pubring.gpg",
"/etc/pacman.d/mirrorlist",
]
]
That's sometimes going to lead to a syntax error.You could have every function incorporate their errors into their normal output, but that means giving up a standardized way of working with errors and warnings. I don't know if you know this, but when you do substitution or piping, by default, only stdout is used. That means that we you do piping, the programs in the pipelines normally do not see the errors in their inputs, and the errors of the multiple concurrently running programs are shown to you intermixed while the pipeline is working. That's a friggin' incredible effect that came from simple design, but when each one would output objects, you'll could get syntax errors or a completely different object like what happened in the above example. You could say, "well, only make stdout an object and let stderr be text," but the fact that they're both the same type means that you can work with the error or only with your errors in pipelines and other shell constructions. For example, `find /etc 2>&1 >/dev/null` will output the directories in /etc you can't read for whatever reason. You might want to pipe that to `xargs chmod` (for whatever reason) after preparing the output to only include the paths.
Right now, programs can strike a good balance in presenting its output in a format that is both readable to humans and other programs. By forcing their output to be structured as objects, and not giving them the option of presenting 2 formats (because we don't want that either), you're removing their ability to present the output in a manner that is readable to humans.
Take for example, rspec's output (a unit testing framework):
$ rspec spec/calculator_spec.rb
F
Failures:
1) Calculator#add returns the sum of its arguments
Failure/Error: expect(Calculator.new.add(1, 2)).to eq(3)
expected: 3
got: nil
(compared using ==)
# ./spec/calcalator_spec.rb:6:in `block (3 levels) in <top (required)>'
Finished in 0.00131 seconds (files took 0.10968 seconds to load)
1 example, 1 failure
Failed examples:
rspec ./spec/calcalator_spec.rb:5 # Calculator#add returns the sum of its arguments
Mind you, that's full of colors in the terminal. It's output that easy to read with the eye and parse with a bit of awk. Can you imagine that being output as a JSON with a generic pretty printer? How will it compare when reading with the eye?The main thing is, though, that, in the question of what the universal format of communication between programs written in different languages should be, text is the simpler, more natural choice over objects. Take note, I don't mean easier. The fact that it's easier is merely coincidence. Simplicity leads to good design because it means less arbitrary choices to make. Less controversial choices to make. Choosing objects leads to more questions: What should the primary types be? Should arrays/lists allow multiple types of elements? Floating types or decimals? Precision restriction on the decimals? Should integers and numbers that allow fractional parts be the same type or different? Should we have a null type? Should we have a date primary type? What about a time primary type? What about a datetime primary type? Whatever answers you give, there will always be groups of people that will dislike them. When you chose text, the only question is really, what encoding? Utf-8. done. Natural, simple design is what we want to be the foundation that myriads of programs and languages can base themselves on and depend on.
There's only one way I'd agree with you that structured output would be nice, and that's with mono-language OSes, like a lisp OS or some other OS where all code is in the same language, and there would be no concept of programs or shared / dynamically loaded libraries or such. In an OS like that, every function is a program, and your shell is the language's REPL. This is bliss when the OS is done in your favorite language. The problem with these kinds of OSes is that we don't all like the same languages and so it'd lead to ridiculous situations where we'd translate a new language into the high level language of the OS. That's what we do in the OS known as the web browser and why we're coming up with WebAssembly.
In conclusion, multi-language OSes like those that are Unix based are awesome, and text as the basis of communication in multi-language OSes is awesome. Therefore, text as the basis of communication in Unix is awesome. :)
If you think something like this is easy much less "you standardize it once and you're done", then you are only cheating yourself out of an essential life lesson.
I’d like to see something like cvs used more, where it can handle the edge cases without breaking, but still doesn’t need a translation step
I have suggested this before: https://news.ycombinator.com/item?id=14675847
(But I don't really care enough about the idea to try to implement it... it would need kernel changes plus enhancements to the user space tools to use it... but, hypothetically, if PTYs got this support as well as pipes, your CLI tool could mark its output as 'text/html', and then your terminal could embed a web browser right in the middle of your terminal window to display it.)
ssh carthoris ls /mnt/media/Movies | grep Spider
(this is just an example). Note that in this example, we have two processes running on two different machines. Indeed, the OSs and systems on these machines may be, um... different. Indeed, I routinely include "cloud" machines in pipelines. Indeed, with ssh, the -Y (or -X) option can introduce a GUI to a part of the command.
I have wished that shar was part of SUS. Also, I find that "exodus" is useful (across Linux anyway -- the systems have to be "reasonably" homogenous). https://github.com/intoli/exodus
In order for this to work over SSH, the SSH client and server would need to be enhanced to exchange this data, and also an SSH protocol extension would need to be defined to convey it across the network.
One might define IOCTLs that work on pipes and PTYs to send/receive control messages. So sshd would read control messages from the PTY and pass them over the network, and the SSH client would receive them and then pass them on to its own stdout using the same ICOTLs. (Alternatively, one might expand the existing control message support that recvmsg/sendmsg supply on sockets to work on pipes and ptys as well.)
Any program supporting such an out-of-band signalling mechanism would have to gracefully degrade when it is absent. If your SSH client or server, or some program in your pipeline, or your terminal emulator, etc, doesn't support them, just fall back on the same mechanisms used today to determine output/input formats.
(rsh is such a deprecated protocol, there would be no point in trying to extend it to support something like this.)
Ha ha, I had dreamed up something like this in one of my wilder imaginings, a while ago: A command-line shell at which you can type pipelines, involving some regular CLI commands, but also GUI commands as components, and when the pipeline is run, those GUIs will pop up in the middle of the pipeline, allow you to interact with them, and then any data output from them will go to the next component in the pipeline :) Don't actually know if the idea makes sense or would be useful.
And another apparent one is image viewer or editor.
There seems to be some precedent for this sort of thing. For example, DVTM can invoke a text editor as a "filter", where the editor UI is drawn on stderr and result saved to stdout.
I did some experimenting recently and found you can open a new /dev/tty file descriptor and tell curses to use this, the rest of the application can continue reading stdin and writing stdout as normal.
The UNIX style.
echo -n 'string' | md5
Compared to that obvious command, this is utter madness. You can't write a single thing without googling. $string = "string"
$md5 = new-object -TypeName System.Security.Cryptography.MD5CryptoServiceProvider
$utf8 = new-object -TypeName System.Text.UTF8Encoding
$hash = [System.BitConverter]::ToString($md5.ComputeHash($utf8.GetBytes($string)))
$hash = $hash.ToLower() -replace '-', ''
No shit the Bash on Windows is a godsend.You can do the same on PowerShell and write a pipeline just the same way.
The best alternative I could find is some community maintained PowerShell extension (with just 177 GitHub stars now), which is far better but the lack of interest in making PowerShell act more straightforward is weird.
"string" | Get-Hash -Algorithm MD5
(https://github.com/Pscx/Pscx)
Still feels funny to read the GGP's comment.
> 70s-era tooling is inferior in some way to something designed with 30 years of hindsight
- structured data
- ability to call any public function across dynamic libraries, COM objects and .NET frameworks
- Powershell modules can also be called as plain .NET code
Which UNIX shell provides this?
$exclude = @("main.js")
$excludeMatch = @("app")
Get-ChildItem -Path $from -Recurse -Exclude $exclude |
where { $excludeMatch -eq $null -or $_.FullName.Replace($from, "") -notmatch $excludeMatch } |
Copy-Item -Destination {
if ($_.PSIsContainer) {
Join-Path $to $_.Parent.FullName.Substring($from.length)
} else {
Join-Path $to $_.FullName.Substring($from.length)
}
} -Force -Exclude $exclude
All to do: find sourceFolder \! -path */app/* | xargs cp -t destFolder
Screwed up the formatting a bit in the first example. I'm not positive the second works perfectly but I know which one I'd rather debug. rsync -a --exclude=app/ --exclude=main.js sourceFolder/ destFolder/
Thank you, thank you unix tools.This example has targeted a very specific deficiency in the way Copy-Item works. I could similarly point out that the following Powershell command would be much more difficult in standard unix tooling:
Get-WMIObject -Computer $remote -Class Win32_Product | Where-Object { $_.name -like "7-Zip*" } | select Version
Which gets the version(s) of 7-Zip installed on a remote machine.