People praise perl, sed, awk, cut, etc. for being good at text processing. But the only reason they need these text-processing tools for pipelines is because they are trying to recover structure from the data that was already present before the previous stage of the pipeline threw that structure away by dumping it to flat text!
Text is obviously a convenient way for humans to view a program's output, so clearly it's useful that all Unix programs (ls, ps, etc) can dump their output as text. But there's no need to dump to text until the output is being sent to a human. If you're piping "ls | grep" there is no reason for "ls" to dump to text and "grep" to parse it back from text, especially since "grep" doesn't know anything about the format of ls's output. It would be way more convenient if you could say something like:
ls | grep 'file.size > 1M'
But the only way to do this today is to parse ls's output first. There would be no reason for this if ls could send structured data to grep.What I'm describing is similar to Monad, Microsoft's next-gen shell. AIUI it can send .NET objects between processes instead of flat text. But IMO it's too imposing to mandate a single object representation like .NET objects.
I'm experimenting with the idea of letting people specify the output of command-line utilities as a Protocol Buffer schema, for example:
message DirectoryEntry {
optional uint64 inode = 1;
optional string name = 2;
optional uint64 size = 3;
// etc.
}
I think this could be a compelling way of making the next generation of usability in command-line pipelines, by saving people from having to write ad-hoc parsers all the time.