Quoting from your post:
And here is my initial solution in UNIX shell:
# bentley_knuth.sh
# Usage:
# ./bentley_knuth.sh n file
# where "n" is the number of most frequent words
# you want to find in "file".
awk '
{
for (i = 1; i <= NF; i++)
word_freq[$i]++
}
END {
for (i in word_freq)
print i, word_freq[i]
}
' < $2 | sort -nr +1 | sed $1q
So you invoke awk, and then run the output of awk through sort and sed.You're doing all the word counting in awk.
Yes, you're invoking awk from a shell script, but that's really not the same thing as "using shell." McIlroy’s solution is genuinely shell:
tr -cs A-Za-z '
' |
tr A-Z a-z |
sort |
uniq -c |
sort -rn |
sed ${1}q
"awk" is generally accepted as a full programming language, whereas "tr", "sort", "uniq", and "sed" are command line utilities. I don't think "awk" classes as a command line utility, so I don't class your solution as "shell".Perhaps you don't agree, perhaps you think "awk" is a command line utility. If so, then we'll agree to disagree.