Text Processing in the Shell
blog.balthazar-rouberol.com
blog.balthazar-rouberol.com
I am still learning tools designed around the constraints of teleprinters
Sure, it’s the same on Windows side (and macOS side with their classic OS compatibility layers still present, like all the HFS stuff). Not bashing bash here.
Surely our computers have very different models of operation than PDP-11, yet we are sometimes pretending it doesn’t
I've been on the job for 10 years and I still use these tools daily. For example, I sent a PR yesterday that was adding a configuration entry on 150+ config files by using find, grep and sed.
I'm not pretending these are the only tools that exist, but darn are they handy sometimes, and good to have in your toolbelt.
- formatting being in-lined via ANSI escape sequences,
- and there's a massive disparity between what escape sequences terminal emulators support,
- control codes being part of the same character set as printable characters,
- changing the behaviour of the TTY requires either terminal emulator support or OS support depending on the behaviour you require because the TTYs are defined partially via kernel drivers (which requires syscalls to alter) and partially by escape sequences,
- and in the case of kernel behaviour, those syscalls vary from one OS to another. Some OS's don't even support from TTY behaviours that other OSs do so you can't even guarantee that logic is cross platform and that you just wrap around specific differences in syscalls,
- resizing terminals UIs can be a nightmare -- often requiring capturing RPC signals and redrawing -- because there's no native layout system for drawing to the TTY,
This isn't meant as a criticism though because there's a lot the general design of terminals gets right (eg the kernel driver for TTY allows us to kill processes over remote shells like SSH and mosh). But I think terminals are one of those things that work "good enough" that most of the ugliness is hidden from everyday users. However if we were to redesign UNIX terminals from the ground up there is a lot of things most engineers would like to change and of lot of places where things could be improved. Like having an out-of-band channel for sending meta-data describing the pipeline but shouldn't be mixed in with the byte stream.
You even forgot to mention one of the things that was addressed: input. Terminal I/O input, done properly, requires a full ECMA-48 decoder state machine, with bodges to accommodate non-conformant warts from the Linux KVT, SCO Console, and RXVT. This is all too often not done properly, because people do not realize that there is ECMA-48 in both directions; and is a mess, just looking at function keys alone and not even accounting for keypad application/normal modes and a mouse/locator. Console I/O evolved into uniform input event records for HIDs that did not require state machines to decode.
Note that the lack of a layout system is only applicable to character-mode terminals. Block-mode terminals are a quite different kettle of fish.
> Note that the lack of a layout system is only applicable to character-mode terminals. Block-mode terminals are a quite different kettle of fish.
Indeed but the point isn't "are these solvable problems?" but rather "why are we still using archaic tech?"
Designing a solution to those problems is actually the easy part. It is shifting the ecosystem away from TTYs that's hard.
> You even forgot to mention
It wasn't intended as an exhaustive list :) There's plenty more issues I hadn't raised.
---
In an ideal world I'd love to see UNIX terminals reinvented. The reality is things are "good enough" for most people that they simply don't notice most of the issues and migrating to the next evolution of UNIX terminals would mean a break in backwards compatibility which will be more disruptive (initially in a negative way) than making do with the warts we currently have.
https://github.com/jdebp/terminal-tests/blob/master/PowerShe... is arguable, and depends from whether one sides with old actual DEC VTs or the ECMA-48:1986 standard, which the DEC VTs didn't keep up with. There are at least two existing terminal emulators that side with the new 1986 semantics, nowadays.
It really isn't anywhere near the most capable terminal emulator right now. This is acknowledged by its developers, and they even have long to-do lists of missing stuff, both as compared to a DEC VT and compared to the likes of XTerm.
It should also be noted that Windows makes a few mistakes when it comes to terminal design too:
- For starters cmd.exe is both a shell and terminal emulator and it's impossible to separate the two.
- To compound things, many common commands like dir, rm, etc are shell builtins (this will be a throwback to DOS). So you cannot even use an alternative shell on Windows without having to either invoke cmd.exe or rewrite existing utilities.
- And if that wasn't bad enough, cmd.exe builtins do not read from STDIN. They instead use DOS syscalls to read keyboard input.
- As you've probably guessed, cmd.exe isn't the only culprit which does this. Any "Windows" command line program designed for or which uses DOS APIs will not follow the standard streams idiom. Any command line software written for NT, however will. So you end up having to write all sorts of really nasty hacks just to get the command line working on Windows (far far nastier than any of the hacks that happen on UNIX/Linux).
- Then you have Powershell, which is an entirely separate command line in its own right and largely - though not completely - incompatible with cmd.exe.
- And WSL, which is also incompatible with cmd.exe and Powershell.
At least on UNIX/Linux, you have one terminal methodology. From there you can use whichever terminal emulator you want, whichever shell you want, whichever programming language to write CLI tools and/or whichever CLI tools you want to download. Where as on Windows you have 4 competing standards which don't cooperate well.
* http://jdebp.uk./FGA/a-command-interpreter-is-not-a-console....
And the fact that CMD and PowerShell provide different interpreted languages is no different to the Korn shell, tclsh, and Perl providing different languages.
Admittedly it's been a little while since I last played around with custom shells and terminals on Windows and I had also rushed my post so you're right that some details were wrong but you're just as far off with your corrections as the points you were criticising so you're really not in a position to be making platitudes about egregious mistakes.
I didn't say cmd.exe was a terminal emulator, I said it was multiple layers including the terminal emulator but not exclusively the terminal. Ok, techically it's conhost.exe that provide the terminal emulation, I'd lazilly lumped that together with cmd.exe because cmd.exe depends on conhost.exe when not run headless. The former requires the latter so you can't just drop cmd.exe into another terminal emulator and run it (other people have tried and there's extensive blog posts about the hacks they've had to do to get it to work, like running conhost.exe off screen).
If you want to be 100% technically accurate then cmd.exe is actually not like any of those things we've described. It's certainly not equivalent to Korn or other language REPLs as you stated. In fact the "language" part of cmd.exe is barely a macro language (again, due to it's DOS heritage). Plus shells orchestrate with byte streams where as NT's streams work very differently and cmd.exe doesn't even behave correctly when used as a CLI tool (which I'll get into below).
You're right that cmd.exe itself doesn't make DOS syscalls, as said aboveI was rushing my post which caused me to conflating two points. What I really meant to say was:
1. cmd.exe builtins read from the NT console API's stdin. Which means you cannot fork out to cmd.exe as a CLI command because any prompts ("Are you sure you wish to delete" type things) just whiz straight past without pausing for input. This was highly annoying when I was developing my alternative Windows shell and wanted to make use of rm, copy, etc rather than having to write those commands all over again.
2. Windows, and cmd.exe by extension, supports running other console applications which don't use NT's console streams because they favour some of the other hacks used in the DOS days. This means those applications also don't work with alternative shells let alone alternative terminal emulators.
Also I think it's disingenuous citing your own blog post as a source. I could link you to the Github repository where I've had to put in numerous workarounds for the shell I've written to work with Windows. But instead I'll link to something a little more recognised:
https://devblogs.microsoft.com/commandline/windows-command-l...
(I did have a hunt around for the blog posts from other developers building console solutions for Windows and the similar problems they've ran into but since it was around 5 years ago when I gave up first party Windows support, those blogs are now lost in the mists of the ether).
I've been doing this for 30+ years. I've written my own terminal emulators and UNIX shells. You're not the only nerd on here so take a step back and listen to the points people are making before assuming they need to be re-educated :)
Are you talking about MS-DOS-style memory-diddling to achieve things like colors and reverse video?
> Platforms with "consoles". These platforms provide a concept of a "console" to applications programs. Consoles support direct screen addressing, to the level of character cells at least, and are accessed through an API that is a first-class part of the overall system API.
> Platforms with "terminals". These platforms provide a concept of a "terminal" to applications programs. Terminals are not directly addressed, but are communicated with via byte stream communications protocols, involving control characters and control character sequences.
In short, you're deliberately being non-specific. I'd claim that a terminal system plus ncurses is a flexible and reasonably efficient console which is portable between many different systems, on the hosting end and on the client end. I could claim that the IBM block-mode "terminals" are consoles if the system is taken as a whole, but such things are markedly less flexible than what you can accomplish with ncurses, albeit more machine-efficient on the host side.
As far as first-class APIs go, you can argue with others about the kernel-vs-OS distinction, and the irrelevance of unbundling in the Open Source world. In short, shipping with ncurses is no more "odd" than shipping with Gtk or, say, a web browser.
Console I/O has a demonstrable evolution over the course of the 1980s, as I have already explained several times, a lot of which was to address the shortcomings of the 1960s terminal I/O model.
Compare languages like Turkish or Korean, which are still natural languages but got a well-thought-out tune-up in the not-so-distant past.
Why can’t my language have a pluralization rule (for example) that’s so simple and regular it takes 1 minute to teach, and 3 minutes to master? Or an alphabet that looks like how it is pronounced, so we don’t have to waste hours each week as children memorizing thousands of special cases? This is absurd.
And then, over the coming years, it would absorb words from other languages with different rules, and evolve according to what people find easy or convenient (or just at random), and in a little while it would be irregular once again.
As you said, Turkish and Korean were changed recently. Give them time and some of that regularity will get chipped away.
(Also, of course: make substantial changes to how the language works and suddenly no one can comfortably read any of the vast quantities of existing writing unless it's translated. And some of that existing writing is really good.)
We are still reading books using a 2k year old alphabet to represent ideas. Not sure why it would be surprising that text manipulation is still the norm pretty much in everything we do, including computing.
I know this sounds diffikult but migraxion to a new system is always diffikult. We're talking about a pretty serious rewrite here so everyone's kooperation will be nesessary. It's not an easy desixion but if you cek into it we can use other languages as a model. Spelling reforms are not a new konsept.
You also need to add more vowels, as there are way more than 5 vowel sounds, but that’s not hard. Umlauts are already common in many languages.
But people have tried this (spelling reforms/re-phoenetification of English). A lot.
Language drifts, that's just what it does.
Try 1725, especially for those of us who (still) set their editors to favor 80 columns.
Text is a very dense way to represent logic and ideas.
This is not like legacy software that you can rewrite.
You can add a GUI based on ideas from the late 70s/80s if you like but you’re still unlikely to come up with a more succinct way to represent logic than can be held in a text file.
So it follows that small tools that deal directly with manipulating text will be useful as long as text is useful.
In more practical terms at least for me, vast majority of stuff I mangle through shell is not really text but structured data, typically either some sort of tree or a table, or something in-between. Sometimes the structure is more ad-hoc, sometimes it is very rigidly defined, but it's still there
And no I don't think it's just a problem for the OS. Although, that's probably a popular idea.
Please make a suggestion or two.
(NB ultimately our character handling is based on 'writing' which goes back thousands of years, not 50, and it survives well).
* http://jdebp.uk./FGA/tui-console-and-terminal-paradigms.html
Recent versions of Perl also support UTF-8 so they can support text processing in different natural languages or internationalization needs. See https://en.wikibooks.org/wiki/Perl_Programming/Unicode_UTF-8
Make sure it's recent enough and released after 2002!
There have been other improvements and fixes in the versions up to 5.30, so the Unicode support now is pretty transparent.
Many of the classic command line text processing tools are not Unicode aware.
Another difference is that sort is optimized to handle large files [0]
[0] https://unix.stackexchange.com/questions/279096/scalability-...
"Unicode with various ways to represent é" is a shit show to parse using shell tools. e.g. Try scraping Spanish language Twitter feeds. When I have done this kind of work, I made a tool to canonicalize glyphs and had to put it between every step of a pipeline.
cat a.txt | perl -pe 's/banana-(\d)/papaya-$1/g'
Or in-place: perl -i -pe 's/banana-(\d)/papaya-$1/g' a.txtWhy do this kind of work in the shell? Isn't it better to do this in a programming language that can run on all operating systems? What are Windows users supposed to do?
It can be quicker & easier to write. Also individually these tools can outperform any code you write by hand. Shells are also common on a far wider range of operating systems than typical programming languages like Python et all. Additionally if you want to improve your productivity you may create your own shortcuts (e.g. "build_myproject" which understands what that entails & may involve some amount of text processing among other things). It's typically far more convenient (shorter, simpler & generally easier to understand) to invoke other programs from shell languages since that's what their programming interface is optimized around.
Sometimes it's good to even wrap the entrypoint for common scripting languages like Python in shell so that, for example, you can setup a virtual environment to run out of or use the proper version of Python.
> Isn't it better to do this in a programming language that can run on all operating systems?
Bash & coreutils have been ported to every operating system under the sun (including Sun operating systems). They're even more common & available than any other programming language that doesn't require a compiler (e.g. `adb shell` will get you into an environment where you can grep & do these operations even though there's generally no python or other scripting language available).
> What are Windows users supposed to do?
* cygwin
* WSL
* Windows ports of Bash (http://win-bash.sourceforge.net/, https://gitforwindows.org/, etc)
* msys
Let me conclude this post that you shouldn't really write anything complex or maintained by multiple people in shell if you can avoid it and if you can make guarantees about (for example) the available of a Python interpreter. That doesn't mean that shell scripts aren't a valuable and important part of the development ecosystem.
As to why use shell, that depends on your use case and working environment. Shell is something like an IDE [0] where you can solve multiple tasks from single environment. You don't have to use multiple programs (window manager, text editor, IDE, etc). Since it is all text, you can save and repeat a command, share it with others, edit a previously written command, etc. This is quite different from a GUI based workflow. Personally, I find using command line more productive, but as mentioned earlier, it'll depend on the task at hand.
Speed. Portability. Muscle memory. I've spent ten years troubleshooting UNIX applications, so most of these commands are fairly well-ingrained into my mode of thinking when I have data that I've got to parse.
To boot, these shell utilities were written by people way smarter than me. I have far more confidence that they will handle edge cases in the data stream infinitely better than whatever dinky little Python script I might try to hash out.
If you need to quickly extract something from a csv, you could break out python, or import it into a database, but using cut and grep (or csv-tools) will take 5 seconds.
The point is if you need to do a specific task many times, do it in a programming language. But if you have an ad hoc task, you're saving a lot of time by being proficient in the shell.
The way things are moving with containers, the idea you're even going to have these utilities on the server, and the idea that the server is writing this stuff to a file system- that's totally changing. So is this useful for local stuff? Maybe, but is Excel probably more useful there?
This includes man pages from tldr, and more! The command line utility had been a great help for me over the past few months.
some_file=example.sh; tail -n +$(( $(wc -l $some_file | grep -o "[0-9]\+") - 5 )) $some_fileThere is nothing new under the sun, it's all just rebranded.
It's designed from the ground up to support object manipulation while still retaining compatibility with the UNIX pipeline.
I do this by building a suite of builtin tools that are aware of structured data files (primarily because that information is passed down the pipeline as a data-type) but it still breaks into normal pipeline when forking an external executable.
What the world needs is the inverse program of "jc", where an unparseable json string is expanded into a flat list of lines all of the form "field.subfield=value"
With any gnu tools it is great that they will stay there for your life and are usually by default installed on every system
For "fun" projects (stuff I'm not paid for) and workflow optimization, I just care about portability, which means C (with heavy use of the C stream library, a simple collections library of about 200 LOC, and occasionally POSIX syscalls) and shell scripting. After spending a lot of time learning languages as a hobby, I just don't believe the dark corners and warts of C and shell scripting are any worse than other languages.