Unix Is Not an Acceptable Unix (2015)
mkremins.github.io
mkremins.github.io
Unix permits programs to communicate with one another, and with the user,
exclusively through character streams. You can’t write a function that
returns a list of files because the shell doesn’t know what a “list” is,
doesn’t know what “files” are, and couldn’t tell you the difference between
a “function” and a “program” if its life depended on it. Programs don’t
“take arguments” and “return values”, they read characters from stdin and
print characters to stdout!
And this is why I consider PowerShell a more Unix-y shell than bash. Powershell commands don't just operate on character streams. They operate on CLR objects. For example, the PowerShell equivalent of grep, Select-String, doesn't return a string, or even a list of strings. It returns a list of match objects, which contain the regexp used, the matching string, and the file that matched. PowerShell often gets dinged for its verbosity, but honestly, I find it much easier to compose PowerShell commands because I know exactly what each command does, and the command-line flags are mostly intuitive. Meanwhile, Unix commands are shorter, yes, but it's a lot more difficult for me to remember, e.g. that if I want to run a command across a number of files, I have to pass the command to find.My initial comment was a response to the snarky "proprietary systems suck" comment, saying that UNIX itself was a proprietary design.
Even as the commercial Unix platforms split, flourished, and fell apart keeping their own source more or less private Ultrix, Solaris, Irix, AIX, SCO, SunOS, OSF/1, HP/UX, Digital Unix, and others were more or less "Unix" with a few caveats about compatibility. POSIX and the SUS were formed. Systems got certified.
Then the GNU tools were written and Linux became their kernel. Then, the Berkeley vs. AT&T lawsuit ended and we got the explosion of open source BSDs - NetBSD, OpenBSD, FreeBSD, Midnight BSD, DragonflyBSD...
So yes, technically the original AT&T Unix was proprietary. That's not all there is to Unix nor has it been for decades. There is no Single VMS Specification. There are no independent systems certified as VMS compliant. You run what DEC / Compaq / HP / VMS Systems supplies and you like it or you don't. There is no Single Windows Specification. You take what Microsoft gives you or not. Unix is an open standard.
Pity the poor VMS ppl. never having the pleasure of trying to figure out what "distro" with which kernel to run on their hardware.
I feel like PowerShell is a step forward in the sense that it's a new thing with a more consistent design than the cobbled, creaky Unix environments we've become accustomed to. But at the same time I feel like it throws away one of the most intuitive aspects of the Unix interactive shells (character streams).
% ls --libxo json
{"__version": "1", "file-information": {"directory":[{"entry": [{"name":"bar"}, {"name":"foo"}]}]}
}
% ls --libxo xml
<file-information __version="1"><directory><entry><name>bar</name></entry><entry><name>foo</name></entry></directory></file-information>Is it not reasonable to think of objects as having a description of some kind. Think of URLs, for example.
find -iname '*.c'
could emit file:// URLs rather than just filenames. Then anything can unambiguously understand the format of the output.That's one of the main downsides of object-pipeline systems; in UNIX you can stop and serialise a pipeline to a file, then unserialise later. But you can't do this with object pipelines in Powershell because the objects have to be "live".
Sure, everything can be made into strings at some point, but I've very rarely felt the need to do so in PowerShell (and regularly see on Stack Overflow that it only leads into problems down the (pipe)line if people try).
That's because you can't be both! The shell wants to be both an API and a user interface, and these are two goals that are fundamentally at odds. An API should be structured, strongly-typed, unforgiving; a UI should be discoverable, intuitive and unconstrained by the needs for automation or programmability. Programmers are stuck in the idea that a command-line shell running in a terminal emulator (emulating hardware from the 70's) is the pinnacle of interaction design, which is extremely sad.
But there's "find", and there's the convention of separating filepaths with a null character. Because null cannot occur in a filename, this is an unambiguous list.
Programs do "take arguments" (another unambiguous list of null-terminated strings, if you quote them). Like when you specify a command to find with "-exec", for it to run for each file. That's another fine way to do composition.
Another example of composition by arguments is "sudo". It switches to root, then runs the arguments as a new process. Another example is ssh, it runs the arguments as on the remote server, and stdin and stdout still work. Check it out:
tar -c testdir | ssh devbox sudo tar -C /opt -xv
That created an archive from a local directory and wrote it to stdout, into ssh which ran sudo on the remote system, which ran tar, which changed to /opt and extracted an archive from stdin. None of those tools have special handling for each other, but they all work together. "Interesting" filenames in the archive are no problem here either.Is there any basis for this claim? My understanding is that many of the GNU ls options are just shorthand for things that you'd traditionally do compositionally. It even reverts to a plain, more pipe-friendly style of output if its stdout is not a TTY.
It still sucks at the job if you assume a system where the convention of not using newlines in filenames is ignored. Unix-likes are typically filled with these conventions that will qualify or disqualify any particular approach to a problem, which imposes a certain load on the users, administrators and developers that I assume isn't as severe in Powershell. An inexperienced user will constantly have to second-guess their solutions and may not experience any issues until much later, e.g. when it turns out that their script will have to accept spaces in filenames, or when their `rm -rf "/$INSTALLDIR"` needs to handle unset variables.
I'm all cool with this since it offers brevity and a profound flexibility. If you want, you can implement typed, object-passing IPC like Powershell on top of Unix. If you don't, you may be content stream-editing and grepping in file listings to get what you want in the format you need. Besides that, the most compelling case against Powershell that I can thing of is that Windows programs typically aren't very composable. It's conventional for even the most mundane and simple tools to be operated using full-fledged GUIs, if only to offer a file selector and a run button.
> Is there any basis for this claim?
Of course. For starters, there's the fact that the date modified column format depends on how long in the past it was modified! This makes it impossible for `sort` to work with.
So suggesting that the `-t` option of `ls` could be removed in favour of using `sort` is entirely wrong.
I could go on and on. There are excellent criticisms of UNIX/Linux. This is not one.
Not if you do the sane thing and use the -T option. Of course that requires the user to read another huge man page to find out that such an option exists.
I think you have failed to understand the kind of alternative people are proposing. Nobody is saying that the current ls sorting flags should be removed and replaced by passing the output of the current ls command to sort. That would be broken and difficult to use for all the usual reasons that passing around unstructured text is broken and difficult. The idea is passing more structured data between processes so that the problem you describe could never even happen. What you describe is actually the type of issue people want to solve (but in the general case, not just for ls).
Formatting data for humans to read and formatting data for machines to read are different problems. Current unix cli tools have to attempt to do both, which often means doing neither well. If you are lucky there are at least some flags you can use to make the output more structured. If you are unlucky the command magically reformats the output based on the type of device it's printing to.
The truth is that shell is partly composable but with many concessions for friendly interactive use.
You could remove the compromises but it would not be much fun to use.
If that's not the case, if Powershell actually does integrate most of what are on Unix-like systems separate commands, then that's terrible.
>Meanwhile, Unix commands are shorter, yes, but it's a lot more difficult for me to remember, e.g. that if I want to run a command across a number of files, I have to pass the command to find.
You don't have to do anything of the sort. If you want to run things across a number of files, you need a list of those files, and you need to know what you want to run across them. That list of files can come from anywhere. Does it come from a text file? A script? Or find? Doesn't matter.
If you do get them from find, you run things on them in precisely the same way you run things on any other list of files: a for loop.
for file in $(find ...); do
rm $file
done
Yeah you can also do find ... -exec rm {} \;
or find ... -print0 | xargs -0 rm
or any other number of other ways of doing it. But you don't have to. I'd probably use xargs to delete those files, honestly.This is an important concept. A powershell script might need to change depending on the source, sometimes it'll be a file object and other times a string, in unix it will always be a string. This is important for composability.
You can't even encode lists of strings in a format that's reliably understood by all UNIX programs.
Sure, it's annoying in a shell when whitespace is used as a separator but that seems more of an issue with shells than with filenames.
Filenames should describe the file. In language, descriptions quite often require whitespace.
Systems should be better at handling files with unicode characters and spaces. Its the fault of the shell that whitespace has been overloaded, not filenames.
If the shell can't handle this well, the problem is that the shell is not very well designed, and we need something better.
Meanwhile, I never have problems with spaces in file names when I'm doing stuff in C++, C#, Python, etc.
find ... -delete>If you do get them from find, you run things on them in precisely the same way you run things on any other list of files: a for loop. for file in $(find ...); do rm $file done
If you know the number of files (and hence total command line length) will not be too high (i.e. not exceed the maximum length of a Unix command-line [1]), and that the filenames will not contain spaces or other problem-causing characters, you can shorten that to:
rm $(find ...) # or rm -i $(find ...) for interactive deletion.
This will only run rm once, instead of once for each file.If either or both of the above conditions are not met, your 2nd or 3rd snippets are preferable.
[1] IIRC, there is a constant defined for that maximum, somewhere in standard C library's header files (maybe something like BUFMAX). I've read that early versions of Unix commands used to crash in some way (core dump or seg fault) if the length was exceeded, and seen that a few times myself. Also may have read that it was quite short initially, like 512 or 1024 bytes. People might have done something to fix that, apart from xargs, later.
Update: I googled for this phrase:
maximum length of unix command line
and looked at the first few hits; found some info:
xargs --show-limits
can give info about the maximum length, which POSIX calls ARGMAX, not BUFMAX as I said above.
and a sample output of xargs --show-limits is:
Your environment variables take up 2446 bytes
POSIX upper limit on argument length (this system): 2092658
POSIX smallest allowable upper limit on argument length (all systems): 4096
Maximum length of command we could actually use: 2090212
Size of command buffer we are actually using: 131072
There is no limit per argument, but a total for the whole command line length. In my system (Fedora 15/zsh) its closer to 2Mb. (line 4).
(from http://stackoverflow.com/questions/6846263/maximum-length-of... )
One of the hits also says that
getconf ARGMAX
shows the value, on POSIX-conformant systems.
"Unixy" invokes very specific concepts; it is not a synonym for "computer sciency", "intellectually agreeable" or "non-sucky in some ways". (Relative to some things, it sometimes is these things, though!)
Dumb character processing is "unixy", and Bash is in fact "dead unixy".
There is all kinds of worse is better. A $20 Shimano-branded rear derailleur made in Indonesia on your bicycle is worse-is-better; nothing to do with Unix.
A funny cat picture posted all over the place is called a meme, but not all memes are cat pictures.
(Assuming that we're not doing anything interesting with the file descriptors; I can see an argument for excluding that possibility from this discussion, but certainly it's a simplification.)
Yes, `ls` isn't perfectly clean in a mathematical sense, but it's still tiny compared to modern software, and passing a couple of flags for common tasks now and then is far more convenient than writing `ls | map ... | sort ...`.
Unix is not a mathematically perfect Unix, but it is most definitely an acceptable one.
If we're going to use the overzealous puritism of the article's definition, why not go further? My keyboard driver isn't unix-ey, because I can use it to input a 'k'. Or a 'K'. Or a lot of other characters. The argument in the article would mean that for a keyboard driver to be unix-ey, it should only output a single character. Nevermind the fact that I have 104 options on the HID at my fingertips, if my driver allowed that, it wouldn't be "doing one thing"!
And by default `ls` does it's job well.
You're missing the article's forest for its trees. The author is claiming that "filter or sort the items in a list" and "get a list of files in a folder" are two different "things". That's not a frivolous, hair-splitting thing to say - the difference between generating a list and re-ordering it should be apparent to anyone who writes code.
After all, when you write a function that returns an array, do you typically have it take in a boolean argument for whether or not to reverse the array's contents?
And the granularity of unix is much more course than a single function. Have you never written a program or application that returns the array and sorts it?
Only rarely that isn't the right thing to do. For example: if the data comes from SQL, I tend to do the sorting in SQL, it I see that as a problem with SQL: it composes badly.
For example, a JVM program running on a client could send a _program_ to a JVM program running on a server to run it there, with the server using the security features of the JVM to enforce that that program is well-behaved by limiting what interfaces it can call, how long it can run, etc.
With that in place, a service could provide a call that produces the data and takes client-provided Java code fragments that filter and sort the result, and use hotspot to do optimizations across the stack.
(SQL, in both theory, can accept query fragments to insert into queries, too, and that can be done in practice, too, by concatenation of query fragments, but it always feels like a hack, and typically doesn't enforce security)
We'll have to agree to disagree here. I think of a 'tool' as more than just a named simple function - the author's definition of 'do one thing' is ridiculously narrow (hence the satire).
But it's just that a general purpose programming language like Python doesn't treat executing subprocesses in enough of a first-class manner, in the way the bash shell does, to make it feel like a seamless shell-scripting experience. You still have to go through the motions of "import os", etc. just to get started, and you can't directly execute other programs, instead having to use "system" or the more cumbersome subprocess module. But still, it's.... tantalizingly close to something like a "better" shell that includes support for lists and other data structures instead of just "everything is a string".
Ruby also does this well, since everything in backticks runs in a shell (just like in bash).
Trying to move more actively to Julia and others for this though. Perl6 looks like it extends on good things in Perl5, but, of course, in the process, breaks many things I am used to in Perl5.
The nice aspect is that there are many good tools, and quite a few mediocre ones to solve problems. Pick your tool sets and be productive with them, taking care to learn others over time.
;ls -al
Plus, it borrows the backtick notation from Perl:
run(`ls -al`)
And you can read the string returned for processing further within Julia:
x = readstring(`ls -al`)
And it supports pipelines... http://docs.julialang.org/en/latest/manual/running-external-...
For the other cases where you do need subprocess, the more sane escaping rules (compared to shell) almost makes the verbosity worth it.
That's my general impression on scripting, too. Shells are great at doing day-to-day things I want to do interactively on a system, but start to develop problems (awkward syntax, lacking readability etc) when it comes to more complex programming. On the other hand, languages like Python do that really well but have problems with the first part.
I sometimes wonder if there's really no way to combine the two in a practical way. One datapoint (on the negative side) would be Powershell, which (as a more recent shell) gives its user very useful tools but still fails (IMO) on the syntax front.
An idea I pondered on for a while was that maybe a shell-specific notation could help. Take xonsh, for example. The concept of it is something I could get used to, but I can't really get around the fact that there's some magic involved in differentiating between shell and python commands[1].
Perhaps a more modal approach (think vim) could help? Say I can switch explicitly between python- and shell-mode (avoiding all the magic) and can write shell-constructs in a script in a dedicated notation (like JS has its own notation for regexes, for example). Then, the main problem would most likely be the interaction between both modes (i.e. how one can pipe the stdout of a subprocess to a variable)...
[1]: http://xon.sh/tutorial.html#python-mode-vs-subprocess-mode
I think "a demonstrably inhumane environment" may be my favourite description to date.
There are many terrible assumptions in the design of the shell. Not least:
Use short commands (because they're quicker to type in theory, even though you'll waste time huge amounts of time in practice looking up commands and switches, because even the average god-like greybeard can't remember more than a few tens, and beginners have no chance)
Use terse switches (ditto)
Text is universal and computers are text processors (not true back then, even less true now)
Make no attempt at standardisation (switches are effectively random between different commands)
Quick hacks are better than a smart and structured environment (this was always bullshit and has caused endless grief, not just in Unix, but in everything touched by Unix)
Software forks are good (not in the shell, they're not)
And now computing is buried under the dead weight of accumulated cruft and poor theory and practice - probably forever.
But that isn't even the real problem. The real problem is that research into humane OS design seems to have stalled. CS academics could band together to design something smarter, better, and friendlier than the current shell. But there seems to be zero interest in starting a project that would generate huge productivity and reliability gains and also make low-level computing more accessible to ordinary users.
I find that baffling, and actually rather disturbing.
Even Plan 9 and Inferno are more close to that design model than UNIX, hence Rob Pike's complaints about UNIX, how ACME works or Limbo's design.
ls | sort | mc
to get an alphabetical list in multiple columns. That didn't last.This reads like someone discovered functional programming and now wants everything to be functional.
ls | sort | mc
would be something like ls -C --sort=name
and, as discussed, the ls program has to be much, much more complicated. ls | sort --key=modification_time | mc
For display, you'd have a convention that if there's a "text" field in the JSON, that's what's displayed in the console window. Otherwise you see the raw JSON. So "ls" would output an array of JSON records with all the usual fields, but if you just wrote ls
the user would see a list of file names.You could have "format" and "select" programs:
ls | select --exclude="owner=root" | sort --key=modification_time | format
--fields=name,modification_time,size
This would be quite powerful and would allow the creation of complex, functional pipelines.
The end result is basically equivalent to SQL, but it's functional!It would be fantastic to have a more cohesive environment, especially in (and including) the shell, but most people are happy with the status quo.
Even worse than *nix tools, there is really no environment outside the realm of text-based UI that follows the Unix philosophy. A GUI that does just one thing, let alone well? Ha. AFAIK, there is no existing way to create GUI software that can cohesively inter-operate with the expressive power that pipes and subshells give command-line tools.
Making the same set of tools be interactive (so we need them to be concise, simple) and used in scripting (so we need them to be constant over time) has led us here. They now form a programming API instead of a set of tools that each do one thing well.
Lets see, I often use ls -tr
* the good thing about this is that it is very fast to type
* the bad part is that you have to memorize the incantation.
You could probably do it by a pipe (ls | sort <now it gets complicated here>)
* good: the good part is that every program does its job
* bad: also two programs have a larger footprint than one
* good: you can have it as an alias to make it short again
* bad: aliases can't be used in bash scripts, unless you explicitly source the file that defines the alias
* good/bad: you just shifted complexity from ls to sort
....
In Powershell, from memory:
ls | sort -property time
https://msdn.microsoft.com/en-us/powershell/reference/5.1/mi...Know how am I supposed to know what type ls returns and what properties are available? For that matter, how am I supposed to know what type of object I'm dealing with? Time isn't displayed by the command and is different to LastWriteTime, how do I discover that it's a property?
So far in this ultra simple example powershell is looking over complicated and under featured.
The command you're looking for is Get-Member or gm. In this case 'ls | gm' will give you what you want. In much the same way that you were told to type 'man', rather than type all combinations of three letter words, you'll have to either pickup a book or read an article to learn a system. As a novice you can't expect to be proficient on day 1 now can you?
You're right that bash/unix requires some learning, but in my experience it requires less upfront learning and it plateaus fairly quickly, when I try and do something twice as complex it will only by 1.1 times as complicated. Powershell on the other hand doesn't seem to plateau, something twice as complex tends to be twice as complicated (or more).
LastAccessTime Property datetime LastAccessTime {get;set;}
LastAccessTimeUtc Property datetime LastAccessTimeUtc {get;set;}
LastWriteTime Property datetime LastWriteTime {get;set;}
LastWriteTimeUtc Property datetime LastWriteTimeUtc {get;set;}> In powershell from memory:
in that comment - I was on an iPad, and didn't have access to Powershell.
Reading HN :
irm https://news.ycombinator.com/rss
Stop all my VMs : get-vm | stop-vm
Stop all processes consuming more than 1GB: ps | ?{$_.WS -gt 1GB} | kill
Find all empty folders: ls -r | ?{-not($_|ls)}
All the PDFs that I downloaded this month: ls -r *.pdf | ?{$_.CreationTime -gt "4/1/2017"}There's a bunch of properties displayed by default, but if you want to see all of them:
ls | select *
Since ls (gci, get-childitem) can list Environemnt variables, Registry entries, etc (which is awesome) these all have different properties. ls Env: | select *
However I do see your point: the docs could be better about explaining that.- Writing a function
- Redirecting to the correct help topic
- Replicating all parameters and parameter sets so that tab completion works
Of course, if help and completion are irrelevant to you (wouldn't work correctly out of the box for bash either, I guess), then it just consists of writing a wrapper function instead of setting an alias.
It seems that UNIX has become the victim of cancerous growth at the hands of
organizations such as UCB. 4.2BSD is an order of magnitude larger than Version
5, but, Pike claims, not ten times better.
http://harmful.cat-v.org/cat-v/Yes, things like ls could be better - but is switching between "ls --plus -a --pile --of --options" and "ls | one | damn | filter | after | another" a particularly core concern?
I tend to agree with this bit... "Rather than re-evaluating the Unix command line with an eye towards improving its usability under the greatly relaxed technological constraints of modern hardware, we’ve written terminal emulators that faithfully reproduce the constraints of the mid-1970s",
... but feel that the author decided that it was a lot more fun and easy to complain about ls instead of properly exploring this idea.
Without flags you'd need "ls1", "ls2" etc. like you do with system calls ala. dup()->dup2()->dup3(), pipe() -> pipe2() etc.
Linux has gone further than some *nixs and started using flags to extend system calls as well see: https://lwn.net/Articles/585415/
ls | filter-hidden | sort | add-directory-slashes | mark-symlinks | mark-executables | colorize | display-in-grid
This seems really inconvenient, but who cares because it's ideologically sound!`ls | filter-hidden | sort | add-directory-slashes | mark-symlinks | mark-executables | colorize | display-in-grid`
I'd start with this:
`ls | nodot | sort | slash | dec | c | grid`
I imagine there's an even better way, I just haven't spent very much time thinking about it.
There's an argument that this could be even more convenient than existing flags—`slash` could work on output from other commands, like `find`—only I don't have to pull up a man page for `find` since I'm not as familiar with it as I am with `ls`. That's the idea behind this kind of composability.
If this is actually too much typing or you're not happy with the defaults, a few bash aliases would go a long way towards making this shorter. For instance, you could alias `ls` such that it colorizes, sorts, and hides dotfiles by default.
Also, these days the cost of doing an extra fs lookup + stat call for each file in each filter is pretty small, but if you want to run that pipeline in a directory of >100k files it might be noticeable.
Then "sort" becomes a utility that sorts a stream of objects by a chosen set of fields; "slash" becomes a utility that inspects the "file type" and "file name" field of an object and makes an appropriate "display name" field; and "ls" becomes a utility that outputs a stream of objects that have "file type" and "file name" fields, straight out of the directory entry. "ls" is the only part of the pipeline that looks at the disc.
Timeless reference: https://www.joelonsoftware.com/2001/04/21/dont-let-architect...
If you rely on communication between external utilities you have to be pretty clear where you draw the line - I don't want to end up with javascript like leftpad debacle on my command line apps.
Have keyboards gotten better/faster? How does the existence of a 80 character line length change the best way to specify a directory listing?
This seems to be based on the logical fallacy that since the old thing is bad that a better new thing must be possible.
Taken to extremes, this article seems to suggest that all command line flags for every executable ought to be replaced with separate executables and pipes. Just imagine how many files would be in /usr/bin, and how many of them you would have to remember just to do useful work every day. We'd have to increase inode limits just to hold all the manpages. An OS like this would not be one I'd like to use.
If you don't need all the options, ignore them. If ls is too complicated for you then do just as the author recommends: "write a simpler alternative to ls", and name it lf ("list files") or the like.
It communicates using byte streams. For the original authors of Unix where ascii ruled supreme, it didn't make a difference but now it does. For example, see this stack question: http://unix.stackexchange.com/questions/167814/get-first-x-c...