1. completely backwards-compatible with ASCII
2. you can always tell if you're reading a part of a multibyte character or not, meaning that you can tell if you're reading a message that was cut in some random point and can tell how many bytes to drop before reaching the first start of a character
3. endian-agnostic due to being specified as a byte stream
4. contains no null bytes, so it can fit in any normal C string
Point 1 is absolutely vital for backwards-compatibility and 2 makes it better than a lot of other multibyte encodings. Consider for instance chopping one byte off the start of a ucs2-encoded string - you'll get complete garbage. And 3 means you don't get strange endian-related errors.
It's a robust, resilient encoding that's a drop-in replacement for ASCII and needs no special support from many utilities - tail and head for instance can work with UTF8 text as if it's plain ASCII, just looking for a \n byte and chopping the input in lines.
So yes, I do think it's elegant. It may not be the simplest possible encoding for unicode, but it's extremely practical.
This is false; the encoding of codepoint 0 is 0x00 per the standard; modified utf-8[1] makes 0 a special case with an over-long encoding: 0xC0 0x80.
And, if you really need to represent codepoint 0 in strings, you can use Java's Modified UTF-8, where codepoint 0 is represented by the byte sequence 0xC0, 0x80. (This isn't valid UTF-8 because in straight UTF-8, every codepoint must be represented by its shortest possible representation.)
No it doesn't, unless you are saying that one should treat it like that. But null termination is as dangerous[1,2] with UTF-8 as it is with ASCII and should be avoided as much as possible anyway. Also, ASCII doesn't mandate that \0 is end-of-string, that's just a "convention" from C.
(Did you notice that my original comment actually included the exact modified UTF-8 link you provided?)
[1]: http://cwe.mitre.org/data/definitions/170.html [2]: http://projects.webappsec.org/w/page/13246949/Null%20Byte%20...
MODULE_LICENSE("GPL\0for files in the \"GPL\"
directory; for others, only LICENSE file applies");
I'm not sure whether this counts as something other than terminating a string.[1] https://en.wikipedia.org/wiki/Loadable_kernel_module#Linuxan...
Well, yes. That's the point: 0x00 is only ever used to encode codepoint 0. It never shows up anywhere else. That's precisely what the text you replied to means.
even the C guys have moved on though, as Go (co-designed by Ken Thompson) illustrates.
This is indeed a tradeoff. The only alternatives are to either store the length of the string explicitly somewhere (the Pascal solution is to prepend the string with a fixed-size length parameter, which can be as awkward as it sounds when you get to really long strings) or to do something nasty with the string's actual contents, such as saying that the last byte of a string has its high bit set.
I think the C solution is the most reasonable when storage is really tight, and more flexible than the Pascal method in general, but I agree that it's potentially dangerous and it's annoying to have a byte which you can never represent in a string.
Why? C strings and Pascal strings can store exactly the same contents, except Pascal strings can store a literal \0, and are faster to manipulate (in many ways), at the cost of (sizeof(word) - 1) extra bytes.
C-strings are useful in extremely constrained environments when the extra few bytes of the length prefix vs the trailing \0 byte is too much to pay, but are essentially just a security risk in any other situation.
A lot of people complain about not being able to easily or programmatically manipulate UTF-8 strings on a character-by-character basis, like it's a typical programmers daily business to go messing with natural language, but I've never seen the great loss.
It also turns out that you very rarely need arbitrary string indexing. Most of the time, when you're indexing into a string, it's a fixed (and relatively small) number of bytes from the start or end. UTF-8 can do this in a tight inner loop that just checks for bytes that don't start with 0b10. If you need .startswith or .endswith, you can just compare bytes with a byte length offset. If you need to do substring search, you can do Boyer-Moore on bytes. If you need to test for equality, do it on bytes. If you need to chop a string in half for divide & conquer algorithms and only need to be approximately right, you can chop in half by bytes and then use UTF-8's self-synchronizing property to find the nearest character boundary.
[1]: http://erlang.org/pipermail/erlang-questions/2011-October/06...
Not all (base, combining+) have a precomposed version, so no that doesn't work.
* the first 14 characters were the length of the data stream in ASCII, padded with 0
* after this came the records - each record started with the size of the record, including the record header
* the record then had each field start with a letter that indicated what sort of record it was - integer, float, character, variable string - all were encoded in ASCII
* variable records were the letter "V", then a 14 byte length
Yes, the format was awful. But you asked when that would be useful. There's your answer.
This is why most modern programming & serialization languages (Go, Python 3, Protocol Buffers, Cap'n Proto) define separate byte[] and string types. Some things are just binary data and should be treated as such. Other things are encodings of world languages and should also be treated as such.