So, I see:
Python 3.7.3 (default, Mar 27 2019, 09:23:32)
[Clang 9.0.0 (clang-900.0.39.2)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> len("ẅ")
2
I wonder why this is? Is the Clang version relevant here?
EDIT: Your "ẅ" doesn't seem to be the same as the OP's "ẅ", although they look the same at first glance.
>>> "ẅ".encode('utf-8')
b'w\xcc\x88'
>>> "ẅ".encode('utf-8')
b'\xe1\xba\x85'
EDIT 2. More info:
>>> import unicodedata
>>> w1 = "ẅ"
>>> w2 = "ẅ"
>>> unicodedata.name(w1)
'LATIN SMALL LETTER W WITH DIAERESIS'
>>> unicodedata.name(w2)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: name() argument 1 must be a unicode character, not str
>>> unicodedata.name(w2[0])
'LATIN SMALL LETTER W'
>>> unicodedata.name(w2[1])
'COMBINING DIAERESIS'
So the second version (w2)
does seem to consist of two separate "characters", LATIN SMALL LETTER W and COMBINING DIAERESIS, which is apparently not the same as the single-character LATIN SMALL LETTER W WITH DIAERESIS. I guess these are actually Unicode code points and not so much "characters" to a human reader, but as another poster pointed out, what the number of characters should be in a string isn't always clear-cut.