The standard '.' (match any character) in a regexp matches an Unicode code point, not a grapheme cluster. To match a grapheme cluster, you have to use "\X", which is not universally supported.
For example, in Python 3, the builtin module module "re" doesn't support "\X", you have to install the "regex" module for it:
# text is 'e' followed by U+0301 (combining acute accent):
text = b'e\xcc\x81'.decode('utf8')
print(f'text: "{text}"')
import re
print(re.match('^(.)(.)$', text).groups()) # prints "('e', '´')"
import regex # must be installed
print(regex.match('^(\X)$', text).groups()) # prints "('é',)"You have the words "supported" and "implemented" mixed up. Kernighan claims Unicode support, so he is required by the standard the implement \X.
If a software does not implement \X, then it is not compliant, and it would be very wrong to say it supports Unicode. Does anyone have a deeplink showing the evidence for awk?
For example, I regularly say that Rust's regex crate has Unicode support. But it does not support \X. It's more precisely documented here: https://github.com/rust-lang/regex/blob/master/UNICODE.md
Which standard is that?
If you're talking about Unicode, the "standard" for regular expressions[1] is an "Unicode Technical Standard", which according to itself isn't required for Unicode conformance:
> A Unicode Technical Standard (UTS) is an independent specification. Conformance to the Unicode Standard does not imply conformance to any UTS.
So awk can claim Unicode support without supporting "\X" (like many regex engines).
If you're talking about POSIX, its regex chapter[2] doesn't mention "\X". In any case I don't think awk claims to conform to POSIX.
[1] https://unicode.org/reports/tr18/
[2] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
I think I once created one that was about a kilobyte. Is there an upper limit?
I created it using this page https://glitchtextgenerator.com/
The 73 byte X:
x̧̡̬̘͓̖̲̻̻̲̠̪̻͓͙̜̂̓̊̔̀̀͗̑̀̅̀̂̚͘̕̚͘͢͜͠