cut -d $delim -f $field1,$field2,... cut -d $delim -f $field1,$field2,...- outputting my selected fields in an order other than their order in the input lines: you can't `cut -f2,1`, for example.
- use a multi-character or regex field separator: -F'<space><star><comma><space><star>' [edit: rewritten in words to dodge HN's formatting] is a real winner when dealing with hand-written CSV.
- use a record separator other than \n: thanks, Windows, for all those times I've written ORS="\r\n" :P - you can even use a regex record separator, which not even perl supports; I don't need that often, but it's a godsend when I do.
- select based on criteria other than "number of fields from the start": I recently needed the last field from lines of varying field-count, and awk '{print $NF}' was just the thing.
But if you don't need anything like that, cut is a lot more succinct - especially if you need to split on tabs only, which it does by default, while awk defaults to splitting on spaces too.
It's a narrow tool. Other than the inherent advantage this brings in clarity/simplicity in a script, there's nothing else to it.
Consider a script that originally had now awk but just a
field=$(somecmd | cut -d: -f3)
but now has been modified to have awk elsewhere. Did the above just become less simple or clear? If so, is it worth changing it to the below for clarity? Does it matter if the awk predates the cut, instead? field=$(somecmd | awk -F: '{print $2}')
I say "no" to both, although I do recognize the argument for just using the same, consistent tool everywhere, instead.However, for ad-hoc command-line usage, I still prefer awk, because (1) the default delimiter is white space, which means I don't have to worry about if what I'm looking at might contain tabs intermingled with spaces (or how many spaces) and (2) it makes it much easier to incrementally increase the complexity of my ad-hoc operation, such as if I decide I want to take a sum or average of one of those columns, after all.
You can shell out if necessary:
cmd = "whois " $url
while ( cmd | getline ) {
# stuff
}
close(cmd)
Also: coprocesses.Can be done with tr, without writing a utility (if by "collapse" you mean what I think you do):
This file t:
$ cat t
the quick brown fox jumped
over
the lazy dog
containing many combinations of spaces, tabs and newlines (whitespace) can be changed to this file t2:$ cat t2
the
quick
brown
fox
jumped
over
the
lazy
dog
by this tr command:
tr -cs "[a-zA-Z]" "\012" < t > t2
That also makes the output more amenable to further processing, including common tasks like finding the frequencies of the words in the input, as mentioned in the "More shell, less egg" post mentioned in this post:
The Bentley-Knuth problem and solutions:
https://jugad2.blogspot.com/2012/07/the-bentley-knuth-proble...
(which I somehow managed to read as an awk invocation).
Re: python - I believe if using things like \w, \d or \s you should be Unicode safe.
Come to think of it, I seem to recall gnu tools should also have Unicode aware pattern/character classes, eg:
https://www.gnu.org/software/gawk/manual/html_node/Bracket-E...
> For example, before the POSIX standard, you had to write /[A-Za-z0-9]/ to match alphanumeric characters. If your character set had other alphabetic characters in it, this would not match them. With the POSIX character classes, you can write /[[:alnum:]]/ to match the alphabetic and numeric characters in your character set.
https://docs.python.org/3/library/re.html#re-syntax
[ed:
Apparently gnu tr is still out in the cold re:unicode;
https://www.gnu.org/software/coreutils/manual/html_node/tr-i...
> Currently tr fully supports only single-byte characters. Eventually it will support multibyte characters; when it does, the -C option will cause it to complement the set of characters, whereas -c will cause it to complement the set of values. This distinction will matter only when some values are not characters, and this is possible only in locales using multibyte encodings when the input contains encoding errors.
]