while (*c) count += ((*(c++) & 0xC0) == 0x80) ? 0 : 1;
See https://stackoverflow.com/questions/9356169/utf-8-continuati... for more details.Counting the number of UTF-8 code units in a UTF-8 string is of course trivial. Counting the number of UTF-16 code units in a UTF-8 strings would take more work. But there's probably no reason you'd want to compute that anyway.
UTF-16's validation concerns are:
1. Broken surrogate pairs, which is mostly benign.
2. Byte-order confusion.
While UTF-8 has:
1. Invalid code points, for example, code points for surrogate halves.
2. Invalid code units, such as 0xFF.
3. Non-shortest forms, where a character may be encoded multiple ways.
4. Representation of NUL, and potential for confusion with APIs that expect null-terminated strings.
In practice the UTF-8 issues have caused much more serious vulnerabilities.
In any case, these are all concerns for a decoder, but not for an API, which is what we're discussing here. In fact, the original comment I replied to up there was advocating the opposite: UTF-16 internally, and UTF-8 for interchange!