tr A-Z a-z | tr -cs a-z '\n' | sort | uniq -c | sort -rn | sed ${1}q
equivalent to tr -cs A-Za-z '\n' | tr A-Z a-z | sort | uniq -c | sort -rn | sed ${1}q
but 4 characters shorter? Or am I missing something? tr A-Z a-z | tr -cs a-z '\n' | sort | uniq -c | sort -rn | sed ${1}q
equivalent to tr -cs A-Za-z '\n' | tr A-Z a-z | sort | uniq -c | sort -rn | sed ${1}q
but 4 characters shorter? Or am I missing something?McIlroy: "A first engineering question to ask is: how often is one likely to have to do this exact task’? Not at all often, I contend. It is plausible, though, that similar, but not identical, problems might arise. A wise engineering solution would produce—or better, exploit—reusable parts." [1]
[1] https://www.cs.tufts.edu/~nr/cs257/archive/don-knuth/pearls-...
Yours says:
Convert to lower case
Convert non-letters to CR
The original says: Convert non-letters to CR
Convert to lower case
Conceptually the same. The point is, most people don't seem to be able to do this at all, although there are many who can tinker with a given solution and improve it around the edges. It's the ability to come up with the initial version that seems to be rare. sort -f | uniq -ic
So I just allow the shell to do word splitting: for x in `cat foo.txt` ; do echo $x ; done | sort -f | uniq -ic | sort -rn | head -<num>And please be mindful of looping over user supplied data like that. Those things are typically the ones that works fine in test but blow up in production. The environment has a maximum size, and you shouldn't store large amounts of data there.
Split the stream of text into words.
Normalize with respect to case.
Accumulate a count of each distinct word.
Sort by the counts.
Print the number needed.
Furthermore, I would guess that the corner cases (such as hyphenated words) are as likely to trip someone up in either approach.Perhaps what makes it rare is not having seen an example before - but one example is enough to demonstrate the principle, thanks to its elegant simplicity.
The hard part used to be in remembering which utility has an option that performs the transformation you are looking for, but practice helped with that, especially once one knew enough to guess which utility probably had that feature, at which point one brought up its man page. These days, of course, it is usually easy to search for the solution you want.
$ cat test.file | tr A-Z a-z | tr -cs a-z '\n' > diffSecond
$ diff diffFirst diffSecond
$
No differences with my test file.