Indexing & substrings are common, too.
> These days code is so often bandwidth limited if anything
Right, and for 1 billion Chinese speaking people UTF16 is 2 bytes/character, UTF8 is 3 bytes/character.
Indexing & substrings are common, too.
> These days code is so often bandwidth limited if anything
Right, and for 1 billion Chinese speaking people UTF16 is 2 bytes/character, UTF8 is 3 bytes/character.
Indexing code points in both UTF-8 and UTF-16 requires reading the whole string up to index location. Substrings are the same as well.
> Right, and for 1 billion Chinese speaking people UTF16 is 2 bytes/character, UTF8 is 3 bytes/character.
That's true for a text file without markup. But most text is not like that in 2017. HTML is probably the most common text format nowadays.
So let's see how a popular Chinese language website does.
curl http://language.chinadaily.com.cn/ --silent | wc -c
52678
curl http://language.chinadaily.com.cn/ --silent | iconv -f utf8 -t utf-16le | wc -c
93368
So UTF-8 seems to be quite a bit more efficient in this case, 52678 bytes. When converted to UTF-16, same page was 93368 bytes.I don’t advocate using UTF16 for the web, but people still code native desktop apps, mobile apps, embedded software, videogames, store stuff in various databases, etc. For such use, markup is irrelevant.
* filenames
* identifiers
* config files
* text protocols
* host names, email addresses
* embedded scripts (including SQL and OpenGL shaders)
* command line interfaces
* translations for languages using Latin alphabets
I don't think 2/3 size reduction for some languages will offset the cost in all the other places.
Some of us use other languages and like to use them everywhere we can.
Other stuff like IDs, shaders before GL 4.2, and many text protocols aren’t Unicode at all.
For configs I usually use UTF-8 myself, because I don’t like writing parsers for custom formats and just use XML, and any standard-compliant parser supports all of them.
Java's String functions don't index by Unicode code points, though. Java strings are encoded in UCS-2, or at least the API needs to pretend that they are.
[0] https://www.mikeash.com/pyblog/friday-qa-2015-11-06-why-is-s...
> Right, and for 1 billion Chinese speaking people UTF16 is 2 bytes/character, UTF8 is 3 bytes/character.
The information density of a single hanzi character is roughly equivalent to 5 letters in English. A Chinese plaintext document in UTF-8 is still smaller in memory footprint than an equivalent English document in ASCII. Of course, most documents aren't plaintext, and where people use characters for metadata (e.g., email, HTML), there is a substantial corpus of ASCII metadata in those documents that UTF-8 is still smaller than UTF-16 even for East Asian languages.
Of course, it's moot since the people who don't like UTF-8 in China and Japan aren't using UTF-16 either. They're using GB18030 or ISO-2022-JP for their documents.
When you need to process Chinese text you don’t care how much an equivalent English document would take. You only care about the difference between different encodings of Chinese language. And UTF16 is more compact for East Asian languages.
> most documents aren't plaintext
That’s true for the web, and that’s why UTF8 is the clear winner there. In a desktop software, in a videogame, in a database — not so much.