Unix Text Processing (1987)
oreilly.com
oreilly.com
Otherwise, I'd have to investigate groff and text conversion utilities to see if the metadata are readily accessible. Interesting question, though I suspect it might take a slight rewrite.
Exactly how useful are they? Would you say they're among the most important skills a programmer could have? Or do we just have a disproportionately large amount of sysadmins for HN readers? Isn't some general-purpose language like Python or Ruby almost always much better, much faster than these solutions? Isn't it worth it to rather invest your time learning Python well -- instead of getting familiar to the segregated and messy environment of modern Unix-land? Certainly, the learning curve is much steep for Unix utilities than, say, Python.
> Certainly, the learn curve is much steep for Unix
> utilities than, say, Python
I am not certain about that at all. Yes they may have trucload of options, but I have never had to memorize them. Many disagree that these tools really dont strictly adhere to the "do one thing but do it well" philosophy they do that to a satisfying degree of approximation. Coreutils, textutils, find and xargs can go a really really long way.I also think of things in an every unix way in my using of computers largely because of 15 years of using unix like systems, it's the same when I need to code more than a couple of lines and do not have vim, for me the brain became too much used to it, I believe this will be the case for many people in HN.
First, many of the things I want to automate are most naturally done at the command line. For these, I already know the commands I want to run, and just need to script the logic that glues them together. Second, I mostly want small scripts I can send to colleagues, and not have to worry about whether or not they have Python (or whatever) installed.
I do write scripts in Python and Ruby, but they tend to be longer, since they reflect tasks where the data have to be pulled apart and put back together in multiple ways (for example, a dependency-generator for some custom makefiles I maintain). This sort of task favors building up structures in memory, over pipelines, and I use the appropriate tool accordingly.
As for the question whether any of these tools are the most important skills a programmer could have, no, I don't think so. As a programmer, your most important skills are in the language you use all the time, the language your "deliverable" is written in. Most people probably don't use the unix tools as their primary programming platform. But those tools support and extend the environment in which we get our "real" programming done.
To borrow the woodworking analogy from (I believe) the Pragmatic Programmers, all those articles about the Unix tools are probably the equivalent of articles on keeping your chisels and saws sharp. No, the file you use to keep your chisel sharp isn't your most important tool; your chisel or saw is your most important tool. But the craft of keeping your most important tool sharp can be fun and rewarding in itself.
By the way, speed is never an issue with any of the scripts I'm likely to run. But if it was, I doubt Python would be faster, since those scripts involve a lot of calls on system resources.
The moment I see I'll need to do something many times, I'll consider writing a script, and then I'll often pick Ruby. But for one-off stuff, the command line is often faster once you get comfortable with a handful Unix tools.
In fact I often find that even when I do things many times, the mental overhead of remembering "yet another script" is often high enough to make it faster to just re-compose the command line I want.
For example, I very frequently do some variation over "grep [some term] | sort | uniq -c | sort -n" to get a sorted list by number of occurrences of [some term], but the key part is "some variation", and that makes adding and remembering an alias less useful.
Another big consideration is that these tools are present on all or most machines many of us use.
For larger applications, having to install other packages is often no big deal, but for example I don't want to find myself dealing with an emergency and suddenly having to pull down tons of packages to use Ruby because I'm not comfortable with the tools that are already on the machine.
Other people I know use perl. The important thing is to be able to automate anything non-trivial that you will do more than a single-digit number of times. One time I spent 30 minutes writing a script for things that a dozen people were doing a dozen times a day. It maybe saved a minute each time, but that's over 2 hours in the first day we had it. It paid off in time for me in 3 days, and probably less than that it terms of "damnit this is boring"
For me, at least, Python has a much higher friction to get from "here's the list of what I do" to "here's an automated version"
I do use Python for some of the more involved things (manipulating timestamps from bash is a pain), and I've converted some bash scripts into "real language" scripts when things started to get hairy, but most of the time, you get 90%+ of the time savings from the first 1% of effort.
I usually use Fabric, SaltStack, or https://github.com/kennethreitz/envoy
Now what makes you think that? UNIX shell tools have a very consistent interface which is really pretty simple (pipes in, pipes out,) they are well documented, and they're interactive and easy to play around with no zero setup.
Also, the UNIX shell has been around for 30 years now and isn't going anywhere. So what's the best investment? Python will change more in the next five years.
it's like anything else; each tool is just a building block. the number of creative things you can do with very little effort by piping the output of one to the input of another is incredibly powerful. it's not always computationally efficient, but many times that doesn't matter.
Python is another great tool, but sometimes a machete is quicker and easier. Different tools for different situations.
But... sitting here in this cubicle, in front of a PuTTy window working on a production server for a large financial institute to manipulate text files and I don't have access to either of those. Ruby is not installed.
Most of the time cut, comm, paste, diff, tr and ed|sed does my job. When they aren't, then there is an old awk binary. So, it depends on the job and the tools you have. IMO every developer should learn the bare minimum about UTP utilities. It won't hurt.
If you're trying to write shell scripts in Python, they're going to be 3-5x longer. Bash is a higher level language than Python. Every tool has its place, and bash and Python are complementary.
Also I realize there is some useless ceremony in Python standard practice. Do you want to know what my Python test runner looks like now?
$ find . -name \*_test.py | sh -x -e
That's it... no BS. I don't know what people are using in the Python world these days but some of it has drifted toward "framework land".
Also, it's not very hard to make this parallel, whereas it is somewhat annoying in Python. And you don't have to worry about global variables polluting each other -- tests stay independent.
I test a big part of my C/C++ code with shell scripts as well.
So basically I find it very helpful to think of yourself as writing shell utilities in Python. Python's not your world. It's part of your world.
You can download your last six months of skype activity from skype.com and they are in the form:
Date;Date;Item;Destination;Type;Rate;Duration;Amount;Currency "July 31, 2012 21:16";"2012-07-31T21:16:01+00:00";"+11234567890";"USA";"Call";0.000;00:00:10;0.000;USD "July 31, 2012 21:15";"2012-07-31T21:15:38+00:00";"+11234567890";"USA";"Call";0.000;01:17:02;0.000;USD
After 15 minutes or so, I came up with the following one liner:
cut -f7 -d";" call_history* | grep -v "Duration" | awk '{ FS=":"; s+=$1*60; s+=$2; if ($3 != 00) { s+=1 } } END {print s " minutes"}'
I could have done the same thing in perl or python in 5 minutes, but it was interesting to "program" only by hooking programs together to achieve the same thing.
After analyzing my skype logs, I found I used 1800 minutes. That would have cost me $180 with prepaid minutes, but only cost $30 with skype.
awk -F\; '
$7 != "Duration" {
split($7, t, ":")
s += t[1] * 60 + t[2] + (t[3] != 0)
}
END {print s + 0}
'
Note the handling of s == "" in END. sed '1s/.*/0/; s/;[^;]*$//; s///; s/.*;//
s/00$//; s/:..$/+1++/; s/:$/++/; s/:/ 60*/; $s/$/p/' |
dcSome of the best known advantages of Unix text processing tools is, as long as you can reduce your problem to 'Text'- There exist some very powerful, succinct and quick solutions to even some very difficult problems.
Well it takes some time to get a grip on how to work with Unix text processing utilities and tools like Perl. Once you are upto speed, you see how much work you can do so quickly with so much little effort. In fact the more you get into it, you realize how much useless code you have writing over the years. While all you needed was a command with a few options.
One thing that I do find interesting, is that the book (on text processing, no less), that is 680 pages x 30 lines x 80 columns (plus some minimal line art) - which, in theory, is around 1.6 megabytes of data, weighs in at 28 MBytes in this PDF.
Regardless of the Irony, great book which the authors clearly put blood, sweat and tears into.
http://oreilly.com/openbook/utp/UnixTextProcessing.pdf
has _scanned_ pages with text layer on top of it. That's the reason why it's so big.
I applaud retyping effort, as it's always better to preserve real content than images of it (even if OCRed), but I cannot say that I'm happy about troff being used for this purpose. It can be a matter of taste, but I don't like the way formatting is done in troff/nroff. That's why I never use it directly (e.g. using ronn to convert markdown text to man page, etc.).
But I understand it's done that way to preserve "the creation process" too, which is also appreciated. And the book is about troff/nroff, so dogfooding is present. ;)
Just curious, what was it for? :)
PS: I had a groff based CV (2006-ish). Now it is a LaTeX based CV.
A lot of my forays into programming started with the Bell Labs books (with troff | pic | eqn ...) prominently on the copyright page. So, it was one of the first things I looked up when I got access to a UNIX terminal. Then I discovered that all the UNIX books by Stevens were done up with troff/groff. From there, it was all "steadily downhill" for me ;-)
P.S: If you already know (La/Con)TeX then groff is literally a walk in the park. And it is always good to know more than one way to do things imho. Good luck with it.
I got a book on TeX coming soon because I want to get better at ConTeXt. ConTeXt is not shy about telling you that for better results you need to understand TeX and use it appropriately. I find it kind of absurd that I've been using TeX for so many years without actually understanding it.
I have been curious about roff since trying to use Plan 9, but in a pretty absent-minded way. They're quite unashamed of providing roff at the expense of TeX, and I believe all of its documentation is roff-formatted, including the technical reports, but I could be mistaken about that.
troff and friends were developed on Unix in its early days by the originators of Unix and it shows in what a good fit they are to the environment and in their elegance; they are Unix programs. TeX was born outside of Unix; it runs on Unix.
Also, do you find yourself writing your own macros much? I haven't the faintest idea what that would require with troff, but I rely on this with TeX quite a bit, mostly to elevate stylistic markup into semantic markup. My impression is that if you want semantic macros you use a macro package, and even then you probably freely intersperse non-semantic macros.
This project I've been working on for a while, uses LuaTeX so that I can connect to a local database, perform some queries and typeset them and their output. I imagine this kind of thing would not be difficult to do directly with a custom pipeline step using troff. Have you done that kind of thing before? If so, how unpleasant was it?
Thanks for talking with me about this.
I do write my own macros. They can be just short-hands for a combination of others in the same way my ~/bin/l is exec ls -l "$@", or sometimes for a simple document I start with just troff and have some macros on top of that. Yes, any distinction over semantics is purely convention.
You may wish to read Kernighan's _Nroff/Troff User's Manual_, http://troff.org/54.pdf, otherwise known as CSTR #54. It's original troff, not groff, but as a succinct reference with elegant prose we often refer back to it. At the end is a tutorial introducing simple macros.
Integrating troff and friends in pipelines and scripts is easy. They take line-based text as input and produce it as output, only switching to binary for some output formats at the last hop. You can also run system(3) from within troff documents, e.g. to include the output of a command, but often that's not the easiest fit.
I recommend again the groff@gnu.org list; they're friendly, patient with newcomers, and interested in showing how they tackle the task at hand.
Thank you for taking the time to answer these questions! I will plough through some of this documentation and make my way over to the list.
EDIT: Wanted to say, tbl nailed it for me that whatever conveniences WYSIWYG formatting might offer, it is never a good substitute for typesetting (a strong opinion to this day -- I still prefer typesetting for documents that need to 'travel')
And yes the letter(heads) came rather nice too.
Handing it over to someone else to maintain is a different thing though. Products like InDesign or PageMaker or Word understand that this is what eventually happens in reality. So, the not-so-steep learning curves that WYSIWYG tools offer have won out in the end.
If the manual is written by techies, and maintained in-house like all those old Bell Labs manuals that used troff, then markup based typesetting is not such a pain, imho.
My first CV was in (gasp!) MS-Word. ;-)
Then I did a groff one for larks, and was a bit surprised at how well it was received (Wow! how "professional" it looks), and can they please have the "original doc file" for their own modifications. Sent them the source/text file only to receive the rather "sailory" emails that came back. :-D
Then I recovered.
Sometimes it's better if ideals are replaced with pragmatism - I achieve a lot more than I used to.
XeLaTeX has finally moved LaTeX into the realm of directly embedding virtually any true type font into the final document. Literally 5-6 lines of LaTeX commands and you are done! I was amazed when I first did it. It was so convenient, when compared to messing with pfbs, then afms, and then.... you get the idea, no?
Is there a similar groff mechanism to pull TTF/OTF fonts into the final PS/PDF document with similar minimal effort? If, could someone be kind enough to point a resource to me. Thanks.
P.S: I was searching, but could not home in on the right keywords to drive me to an answer.
Gunnar's Heirloom troff has TrueType support. "troff can access PostScript Type 1, OpenType, and TrueType fonts directly, that is, it can read font metrics from AFM, OpenType, or TrueType files, and can instruct its dpost post-processor to include glyph data from PFB, PFA, OpenType, and TrueType files into the output it generates". http://heirloom.sourceforge.net/doctools.html
$ printf 'Hello ①②③\n' | preconv
.lf 1 -
Hello \[u2460]\[u2461]\[u2462]
$ printf 'Hello ①②③\n' | groff -k -Tutf8 | grep .
Hello ①②③
$ awk 'NR > 3 && !/foo/ {s += $(NF - 2)} END {print NR, s + 0}'