More shell, less egg (2011)
leancrew.com
leancrew.com
"Literate programming" was thus a way to monetize software if you were on the Stanford faculty. (This was before professors started doing startups.)
Yes, Knuth could have used a calculator, but the calculator still needs a division algorithm, and `sort` probably needs a complex data structure or algorithm. (At least if it's to be fast and work on large files.)
The only thing I can imagine, is that he wanted to advertise UNIX pipes (I would too!) and saw an opportunity to get attention by criticizing someone famous.
This is a form of wonderful magic, available only because of the discipline and vision of those who wrote the original tools.
As a side note, I think the shell script version would have been multicore friendly, at least up to six cores. I am sure Knuth's was single threaded.
Finally, to your point on horrible hacks, maybe Knuth's version dealt with a whole bunch of edge cases in each, like for instance dealing with dos/unix/mac endlines. But, probably not -- it was a toy program written as evangelism and out of courtesy. sort and uniq et al embed many programmer lifetimes of experience and bug fixing at this point. Getting all of that nearly for free is really, really great.
That's not how the sort command, in lines 3 and 5, works.
No, it has function calls instead.
> I think the shell script version would have been multicore friendly
Not unless 'sort' is implemented very cleverly.
> maybe Knuth's version
We'll never know unless someone manages to dredge up the source code.
But the specifics on Knuth's version are not really the point. Someone else might have been able to do better. And someone using a less brain damaged language that Pascal might have been able to do even better. In fact, here it is in four lines of Common Lisp (using Ergolib - https://github.com/rongarret/ergolib):
(defun histogram (path)
(bb l (split (file-contents path) t :test (fn (a b) (whitespacep b)))
l (sort 'string< (mapcar 'string-upcase (remove "" l)))
(for item in (remove-duplicates l) collect (list item (count item l)))))
Writing a pure CL version is left as an exercise. My guess is it would be 10-20 LOC.The shell script is easy enough to approximately fix:
tr -cs '[:alpha:]' '\n' |
tr '[:upper:]' '[:lower:]' |
sort |
uniq -c |
sort -rn |
sed ${1}q
(In my tests, it doesn't turn ẞ into ß, but it does turn É into é. "straße" is recognized as all one word, but not as the same word as "strasse" - but it's not clear whether it should be or not. I'm not sure how it'll handle normalization, if the same character is represented in two different ways.)1. Think of a task that can be done in six lines of shell
2. Ask person pushing a new programming technique to apply their idea to said task
3. Reply that it would have been better to just use six lines of shell
Call me cynical but it strikes me that Bentley's (edit: McIlroy's, see replies) reply could easily have been written without seeing Knuth's code. I don't know anything whatsoever about Literate Programming, but did Knuth mean it to be applied to such simple tasks?
It's impossible to know if Bentley intentionally picked the problem that would be easy to solve in shell script, but I don't think it's likely.
A lot of text processing problems are solved really well by the standard set of unix tools, I write similar scripts probably about once a month to extract counts out of log files.
I mean, one presumes that when Knuth wants to count the words in a text file he uses shell scripts, right? It would follow that he's only applying Literate Programming to a toy problem for purposes of illustration. As such, saying he's wrong to use LP for this task seems to miss the point (and doesn't say much about LP).
Knuth republished both, so they must have some things to say about Literate Programming.
That statement is not about Knuth’s presumably beautiful work. It’s that the least deterministic (and most important) part of programming is discerning the value of doing the thing at all.
This advice would be more actionable for me if there were more shell tools that operated on data structures (e.g. objects represented as JSON).
But more generally, I'm dealing all the time with a hodgepodge of json, csv, yaml, POM files, git logs, whatever. We need more tools that understand the semantics of data.