This pipeline may be significantly reduced by replacing cut's with awk, accommodating grep within awk and using awk's gsub in place of tr.
$ echo token:abc:def | grep -E ^token | cut -d: -f2
abc
$ echo token:abc:def | awk -F: '/^token/ { print $2 }'
abc
Conditions don't have to be regular expressions. For example: $ echo $CSV
foo:24
bar:15
baz:49
$ echo $CSV | awk -F: '$2 > 20 { print $1 }'
foo
baz //d
You can get a list of them with a single Awk line. awk -F'//d[[:space:]]*' 'NF > 1 {print FILENAME ":" FNR " " $2}' source/*.c
You can even create a GDB script, pretty easily.(IMO, easier still to configure your editor to support breakpoints, but I’m not the one who chose to do it this way.)
Would you have //d<0xA0>rest of comment?
Or some fancy Unicode space made using several UTF-8 bytes?
If tabs are supported,
[ \t]
is still shorter than [[:space:]]
and if we include all the "isspace" characters from ASCII (vertical tab, form feed, embedded carriage return) except for the line feed that would never occur due to separating lines, we just break even on pure character count: [_\t\v\f\r]
TVFR all fall under the left hand, backspace under the right, and nothing requires Shift.The resulting character class does exactly the same thing under any locale.
The isblank function tests for any character that is a standard blank character or is one of a locale-specific set of characters for which isspace is true and that is used to separate words within a line of text. The standard blank characters are the following: space (’ ’), and horizontal tab (’\t’). In the "C" locale, isblank returns true only for the standard blank characters.
[:blank:] is only the same thing as [\t ] (tab space) if you run your scripts and Awk and everything in the "C" locale.
Because it’s the one I remembered first, it worked, and I didn’t think that it needed any improvement. In fact, I still don’t think it needs any improvement.