The sad history of Unicode printf-style format specifiers in Visual C++ (2019)
devblogs.microsoft.com
devblogs.microsoft.com
And windows wasn't the only victim of using "two bytes for a char" for unicode support.
Java fell in this trap too. The team was split in deciding "whether to waste a single byte per character for string" -- that's in their words, and then decided to rather be ready for the future. That's in 1995.
And then everyone later watched UTF-8 take over the world.
The PHP team also wasted their version 6 and a few years of development effort trying to support multi byte chars by ditching all previous code. And then failed. There was never a version 6. They continued with the old code base with version 7.
Why didn't they go the UTF-8 round? May be Unicode consortium was late to introduce it. Maybe something else.
The original Unicode marketing / position statement from 1988[1] may provide a clue:
“In the Unicode system, a simple unambiguous fixed-length character encoding is integrated into a coherent overall architecture for text processing.”
“Unicodes [sic] are the most straightforward multilingual generalization of ASCII codes: - Fixed length of character code (16 bits); [...]”
“Are 16 bits [...] sufficient to encode all characters of all the world’s scripts? [...] Yes.”
“[A] fixed length-encoding is flat-out simple, with all the blessings attendant upon that virtue.”
Etc., etc.
The hypothetical possibility of more than 2^16 characters was introduced in Unicode 2.0 (1996), while actual such characters didn’t appear until Unicode 3.0 (1999). Windows NT shipped in 1993, OpenStep in 1994, Java and JavaScript in 1995. UTF-8 was presented at USENIX in January 1993; a contemporary exposition[2] says that “the 4[!], 5 and 6 byte sequences are only there for political reasons” (presumably referring to the fact that Unicode committed to 2^16 code points while the new, Unicode-compatible draft of ISO 10646 stuck with 2^31).
If you look at the text following the “Yes” quote, you’ll find that “all characters” is carefully defined to mean ”all characters in current use from commercially non-negligible scripts”. Compared to the current definition of “all characters we have reasonable evidence have ever been used for natural-language interchange”, it doesn’t sound as noble, but would also exclude a number of large-repertoire sets (Tangut and pre-modern Han ideograms, Yi syllables, hieroglyphs, cuneiform). Remove the requirement for 1:1 code point mapping with legacy sets, and you could conceivably throw out precomposed Hangul as well. (Precomposed European scripts too, if you want, but that wouldn’t net you eleven thousand codepoints.)
At that point the question seems to come down to Han characters: the union of all government-mandated education standards (unified) would come down well below ten thousand characters, but how well does that number correspond to the number of characters people actually need? One potential source of death is uncommon characters people really, really want (proper names), but overall, I don’t know, you’d probably need a CJKV expert to tell. To me, neither answer seems completely implausible.
On the other hand, it’s also unclear that a constant-width encoding would really be all that valuable. Most of the time, you are either traversing all code points in sequence or working with larger units such as combining-character sequences or graphemes, so aside from buffer truncation issues constant width does not really help all that much. But that’s an observation that took more than a decade of Unicode implementations to crystallize.
It is certainly annoying how large and sparse the lookup tables needed to implement a current version of Unicode are—enough that you need three levels in your radix tree and not two—but if you aren’t doing locales it’s still a question of at most several dozens of kilobytes, not really a deal breaker these days. Perhaps that’s not too much of a cost for not marginalizing users of obscure languages and keeping digitized historical text representable in the common format.
[1] https://apps.dtic.mil/sti/citations/ADA221614 (scanned PDF) or http://archive.adaic.com/pol-hist/history/9x-history/reports... (PostScript) or http://archive.adaic.com/pol-hist/history/9x-history/reports... (ASCII)
Every time I read about the history of character encodings I feel like I learn about a new encoding standard that attempted to standardize things. Reading about this led me to reading more about ASCII as well. I learned it was derived from the 1924 ITA2 standard which was itself derived from the "Baudot" printing telegraph encoding from 1874! It always amazes me how much history surrounds this topic! [1]
Also, that DTIC site is such a treasure trove of great information! :D
It seems the solution would be simple: once UTF-8 introduced (back in 1996!), add new "UTF-8" codepage, add new "SetDefaultCodepageToUnicode()" function, and tell everyone that wchar_t is legacy and should not be used anymore. They could probably do it in time for Windows XP!
Instead, UTF-8 codepage was only introduced in Windows 7 (but had incomplete support), and finally got the proper support in Windows 10. That's a lot of time Windows programmers were forced to use UCS-2.
It's one of those cases where you need to hit your head hard and repeatedly until you bite the bullet and accept that the cost of multibyte (for example, strlen() different from character count, need to parse to seek etc.) are trivial compared to the benefits.
Countries like Japan had already solved encoding differently while US was still using ASCII. (See JIS encoding).
As time passed tho, the byte savings matter less. Also combined with transfer compression makes the encoding size differences negligible.
In practice, basic 7-bit ASCII characters dominate a lot of texts, even those in East-Asian languages, as all the markup and such around the actual text is almost always in plain ASCII. For example, if I request a Japanese document from a HTTP server then the HTTP headers, HTML tags, CSS is all in single-byte characters.
Of course, there are still many scenarios where UTF-16 "wins" in terms of size – especially with plain text – but real-world size comparisons tend to be a bit more tricky.
For example, for the current Wikipedia homepage of ja, zh, and ko:
% file *
ja-16.htm: HTML document, Unicode text, UTF-16, little-endian text, with very long lines (2709)
ja-8.htm: HTML document, Unicode text, UTF-8 text, with very long lines (2709)
ko-16.htm: HTML document, Unicode text, UTF-16, little-endian text, with very long lines (2503)
ko-8.htm: HTML document, Unicode text, UTF-8 text, with very long lines (2656)
zh-16.htm: HTML document, Unicode text, UTF-16, little-endian text, with very long lines (5589)
zh-8.htm: HTML document, Unicode text, UTF-8 text, with very long lines (5589)
% ls -lh
210K ja-16.htm
118K ja-8.htm
193K ko-16.htm
106K ko-8.htm
189K zh-16.htm
104K zh-8.htm
The UTF-16 versions are all larger (and that's just the HTML, excluding the CSS, JS, etc).I wrote a small script to count the number of "wide" multibyte characters:
ja-8.htm 100,808 7-bit characters; 19,026 wide characters
ko-8.htm 93,888 7-bit characters; 14,149 wide characters
zh-8.htm 91,589 7-bit characters; 14,360 wide characters
It was higher than I expected, although not that surprising when looking at the
source when you have things like: <div class="vector-page-toolbar">
<div class="vector-page-toolbar-container">
<div id="left-navigation">
<nav aria-label="名前空間">Plus many identifiers come from libraries, and when creating their own identifiers many people use either full English or partial English no matter what language (it was a huge mistake to not use English for many identifiers in my first programming job, as you will invariably end up with a mishmash of two languages).
But it is easy enough to verify this with some actual websites: https://www.rakuten.co.jp is 330K in UTF-8 and 625K in UTF-16, https://ameblo.jp is 104K in UTF-8 and 187K in UTF-16, baidu.com is 360K in UTF-8 and 717K in UTF-16, sina.com.cn: 455K, 854K, daum.net: 666K, 1.2M.
And all of that is only the HTML document; if we'd add up the CSS – where there's almost no possibility to use non-ASCII outside of class and ID names – and JavaScript – where the filesize is usually dominated by React or jQuery or whatnot – things would skew even more in favour of UTF-8.
I'm sure there are examples where a page served over UTF-16 is smaller, such as pages with very little markup (like e.g. HN), but that is the common case, even for websites exclusively written and for users of CJK languages. Someone who does not speak a word of English will save many bytes of data every day with UTF-8. There's a reason all those websites are served over UTF-8 and not UTF-16.
But for the sake of the argument, let's replace all class="...", id="..", and data-event-name=".." with strings of the same length consisting of "回". That grows the filesize from 118K to 151K, and ... it's still larger in UTF-16 with 207K. We could start replacing more stuff and eventually UTF-16 may win, maybe. But you have to use a lot of CJK. Let's use a random excerpt:
<button
id="回回回回回回回回回回回回回回回回回回回回回回回回"
tabindex="-1"
data-event-name="回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回回"
Has 91 7-bit characters and 60 multibyte ones (this includes indentation, which may not be represented 100% accurately here). If we do the math this is: UTF-8 91×1 + 60×3 = 271 bytes
UTF-16 91×2 + 60×2 = 302 bytes
UTF-8 still wins.And to repeat, there are certainly cases where UTF-16 is smaller. Markdown documents and other plain text files is an obvious one, but HTML is rarely one of them.
But imagine actually checking things before making a claim...
You're missing the point entirely, the amount of characters you used is enough for 2 or 3 sentences. This was not an example constructed in good faith.
For example plain text emails still have considerable ASCII data in the headers, no matter which language you use to send them, and depending on the length of your email UTF-8 may be smaller than UTF-16. I just checked a very simple email in my inbox, and it has 165 lines of headers for 21 lines of actual email body text (plain email, no MIME). Even after removing the more modern headers (DKIM, X-Spam-, etc.) we're still left with 35 lines, or 1,886 bytes. The actual email text is 765 bytes, in basic English. If the email text was in CJK it would still be slightly over 1K smaller* in UTF-8 (4,181 bytes vs. 5,302).
Longer emails would be smaller in UTF-16. I don't know what the average works out to, but ~21 lines seems about average for an email, give or take.
I would also argue that the fast sync with at most one character consumed is a good feature but not essential for adoption. Any byte-oriented Unicode packing scheme that can gracefully consume ASCII is better than multibyte, codepages sent out-of-band etc. etc.
Languages such as PHP didn't get true UTF-8 support until 2015, Python in 2008 (and that switch took a decade), and Java and JavaScript continue to use UTF-16 as default strings.
The wastefulness of UTF-16 is mostly remedied by Unicode Compression (SCSU), which can be much shorter than UTF-8 in languages like Hindi and Chinese.
I have no problem with saving text files in UTF-8 and normally do. I just wish SCSU was handled by COM by default as well and not only sometimes.
You are probably right. Just slightly off topic: Hindi speakers who use computers mostly seem happy enough with English-language forms, instructions, and technical documents. This is true of most Indian users who speak languages that use one of the Indic scripts. Actual non-English text is useful mainly for publications like newspapers, magazines, and books, which native speakers do prefer to consume in their own language.
The above is, of course, not counting the significant numbers of Indians who are quite comfortable using English for daily communication. These Anglophones are just like US speakers in their preferences and have no use for internationalization.
Only if you never go out of the basic multilingual plane. And then find out so many bugs when people start using emojis.
The historical reasons of why UTF-16 exists are fully understandable, but by now it is clear that it's the worst of both worlds.
For East Asian languages UTF-16 uses 50% less memory, these characters take 3 bytes in UTF-8, but only 2 bytes in UTF-16.
utf16 decode
if (unit <= 0xD7FF || unit >= 0xE000)
/* one unit */;
else
/* two units */
utf-8 decode
if (unit < 0xE0) {
if (unit < 0xC0)
/* one unit */
else
/* two units */
}
else if (unit < 0xF0)
/* three units */
else
/* four units */
UTF-16 will take one comparison for all characters up to 0xD7FF; UTF-8 normally takes two. Same when encoding: utf16 encode
if (value <= 0xFFFF)
/* one unit */
else
/* two units */
utf8 encode
if (value <= 0x7FF) {
if (value <= 0x7F)
/* one unit */
else
/* two units */
}
else if (value <= 0xFFFF)
/* three units */
else
/* four units */
(We can tune UTF-8 it to take one comparison for ASCII if we expect mostly ASCII.)Things get even more complex when we are to read an untrusted string. In UTF-8 we have five byte types, invalid bytes, conditionally invalid bytes, and must recognize surrogates (as errors). In UTF-16 all 16-bit units are valid and there are only three unit types: character or surrogate, high or low. I once wrote a decoder for UTF-8/16 and here's my stats:
UTF-8 : 13 byte types, 13 states, 10 actions
UTF-16LE : 3 byte types, 4 states, 5 actions utf-8 decode
switch (std::countl_one(unit)) {
case 0:
/* one unit */
break;
case 2:
/* two units */
break;
case 3:
/* three units */
break;
case 4:
/* four units */
break;
default:
/* not code point boundary */
break;
}I think there was also some compatibility concerns, as Windows for a while only had a single global "ASCII" character set used by all apps[2], so you'd have issues with actual legacy apps expecting a real legacy code page breaking / having scrambled text if you tried to use UTF8 as the "legacy" encoding. Or the fun idea of an app that is clever enough to know that multibyte exists, but dumb enough to know they can only be one or two bytes long, and allocating buffers accordingly. I think there might have been some issues with how things like the clipboard handled text as well that assumed all "ASCII" apps used the same encoding. (Disclaimed: old memories of reading blogs, may be bollocks). At some point things must have been fixed so multiple character sets could be used though.
[1] But not on 9x, which was "ASCII" only, at least until late in the day when a Unicode compatibility library was created.
[2] The "Language for non-Unicode programs" as the Control Panel calls it.
Not quite, you still can't use some features like longPathAware with the "ANSI" APIs. The preferred APIs for file paths are still the ones using wchar, although the UTF8 codepage is apparently the preferred method for non-path APIs and console IO, at least on the GDK.
IIRC, the problem was that a MBCS codepage could have a maximum of 2 bytes per character, while UTF-8 could need 3 or even 4 bytes per character. Legacy applications could have fixed-size buffers with space for only 2 bytes per character, and would break if the default codepage needed more than that.
> Instead, UTF-8 codepage was only introduced in Windows 7 (but had incomplete support), and finally got the proper support in Windows 10. That's a lot of time Windows programmers were forced to use UCS-2.
It appears to me that lately Windows hasn't been as obsessed with backwards compatibility as it had been in the past. They are also gradually allowing for more use of paths longer than MAX_PATH (260 characters), which are not accessible to legacy applications, and as everyone knows, the compatibility with 16-bit applications has been dropped some time ago.
Make "SetThreadLocale" or something work with UTF-8, so apps could switch one-by-one (this may not work with hooks but still better than nothing...). Add UTF-8 support to MSVC runtime, so at least fopen can be safe (people have been asking for it forever). Add FILE_FLAG_UTF_8 to CreateFileA. Have Notepad be able to edit UTF-8.
But I suspect the real reason is there was a strong internal opposition to UTF-8 -- maybe it was seen as inferior to UTF-16, maybe it was NIH, maybe they hoped the world will change to follow Windows ways and to put BOM marks in each document...
Oh well, I am glad they have at least changed now.
It's not as if they were the only ones. I had a lot of utf-related pain in the Linux terminal back in the day.
Vi and Emacs took a long time to handle variable length encodings well. My old Unix editor of choice, NEdit, never managed the switch at all.
I got a mail from my boss in utf-7 once. God knows how that happened.
https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
Windows multibyte / mbcs is not unicode.
Windows NT started development in 1989 and was released in 1993. Considering the Unicode standard was first published in 1991/2, I think Windows NT can be counted as an early adopter.