I once observed a lot of bugs showing up in a Python project because of transparent conversion between byte-strings and unicode-strings. Programmers would test with byte-strings, everything is fine, customer is in east Europe, and you know what happens next. I think I decided that UTF-8 was actually worse than UTF-8-xor-69 so that the innermost code should be able to tell whether or note this was a bytestring or not. Explicit tagging found
loads of bugs that day.
An array of bytes is a storage medium. If you don't need the number of utf-8 characters (and given things like combining characters and invisible spaces, it's rare that's what I actually need) and just need to print and compare, byte-strings are efficient and acceptable. The text remains utf-8 encoded, but `length` is `O(1)` or a simple scan, and `substr` (or similar) are `O(1)`+the cost of allocating.
However if I need to recode things, an array of 32-bit (or 64-bit hah) ints is sometimes better. If the language has a real iterator, UTF-8 might still be okay, but then you are still using a bytestring anyway.
Lastly, if the "string" actually knows its character set, then using things besides UTF-8 might be good because it saves a lot of space. Maybe. I'm not really convinced because the tag for the character set (as you rightly think necessary) takes up space too. Dealing with these kinds of "strings" is really complicated, so I'd agree that anyone who would jump into complexity without some real justification is, as you say, a bozo.