> There's no inherent, fundamental reason in the universe other than momentum that says you've got to use an 8-bit encoding.
"Other than momentum"? That's a double-standard: Momentum is the only reason why UTF-16 is still relevant today. Ignoring momentum, UTF-8 still has a bunch of advantages over UTF-16, namely that endianness isn't an issue, it's self-synchronizing over byte-oriented communication channels, and it's more likely to be implemented correctly (bugs related to variable-length encoding are much less likely to get shipped to users, because they start to occur as soon as you step outside the ASCII range, rather than only once you get outside the BMP). What advantages does UTF-16 have, ignoring momentum?
The comparison between UTF-8 and UTF-16 is adequately addressed here: http://www.utf8everywhere.org/
I don't want to debate abstract philosophy with you, anyway. Momentum may be the reason, but there's no plausible way that UTF-16 is ever going to replace octet-oriented text. The idea that UTF-8 and UTF-16 are equivalent in practice is a complete fantasy, and I'm arguing that we should pick one, rather than always having to manage multiple encodings.
> I thought we software types are supposed to be big on abstractions and coming up with clever ways of managing complexity?
The best way to manage complexity is usually to adopt practices that tend to eliminate it over time, rather than adding more complexity in an attempt to hide previous complexity. It doesn't matter how "clever" that sounds, but it's generally accepted that it takes more skill and effort to make things simpler than it does to make them more complex.
> It seems rather rigid to say you'll only ever deal with one text encoding.
It seems rather rigid to say you'll only ever deal with two's-complement signed integer encoding.
It seems rather rigid to say you'll only ever deal with IEEE 754 floating-point arithmetic.
It seems rather rigid to say you'll only ever deal with 8-bit bytes.
It seems rather rigid to say you'll only ever deal with big-endian encoding on the network.
It seems rather rigid to say you'll only ever deal with little-endian encoding in CPUs.
It seems rather rigid to say you'll only ever deal with TCP/IP.
Why not eventually only ever deal with one text encoding? There's no inherent value in paying engineers to spend their time thinking about multiple text encodings, everywhere, forever.
Remember that we're talking about the primary interfaces for exchanging text between software components. Sure, there are occasions where someone needs to deal with other representations, but the smart thing to do is to pick a standard representation and move the conversion/negotiation stuff into libraries that only need to be used by the people who need them. This allows the rest of us to quit paying for the unnecessary complexity, and incentivizes people to move toward the standard representation if their need for backward compatibility doesn't outweigh the cost of actually maintaining it.
> I am pretty sure every Unix-like system I have set up gave me Latin1 by default, even quite recently.
Ubuntu, Debian, and Fedora all default to UTF-8, and have for several years now. You're going to have to name names, or I'm going to assume that you don't know what you're talking about.
> I think you are underestimating the extent to which UTF-8 is a crude hack designed to avoid rewriting ancient C programs that did very much the wrong approach to localization.
Really? How would you have done it so that UTF-16 wouldn't have broken your program? Encode all text strings as length-prefixed binary data, even inside text files? It's ironic that you say it's immature to call UTF-16 "wrong", but you've basically just claimed that structured text in general is "wrong".
Let's not forget that nearly every important pre-Unicode text representation was at least partly compatible with ASCII: ISO-8859-* & EUC-CN were explicitly ASCII supersets, Shift-JIS & Big5 aren't but still preserve 0x00-0x3F, and even EBCDIC, NUL is still NUL. Absent an actual spec, it was no less reasonable to expect that an international text encoding would be ASCII-compatible than it was to expect that such a spec would break compatibility with erverything. Trying to anticipate the latter would rightly have been called out as overengineering, anyway.
In that environment, writing a CSV or HTML parser that handles a minimal number of special characters and is otherwise 8-bit clean is exactly the right approach to localization.
Also, "ancient C programs"? Seriously? Are you really saying that C and its calling convention were/are irrelevant?
> There was a time before UTF-8 existed and became popular when it was a fairly common viewpoint that proper Unicode support involved making a clean break with the old char type.
Sure, and it was a fairly common viewpoint that OSI was the right approach, and that MD5 was collision-resistant. Then, we learned that all of these viewpoints turned out to be wrong, and they were supplanted by better ideas. UTF-8 became popular because it worked better than trying to "redefine ALL the chars".
> You are right to say that UTF-8 "won" in most places but the fact that the NT kernel or the JVM use 16-bit chars reflects that prior history.
The JVM isn't comparable, because its internal representation is invisible to Java developers. I can write Java code without ever thinking about UTF-16, and a new version of the JVM could come out that changed the internal representation, and it wouldn't affect me. Python used a similar internal representation, and recently did change its internal representation. Most Python developers won't even notice.
If NT used UTF-16 internally, but provided a UTF-8 system call interface, I wouldn't care. The problem is that, in 2013, people writing brand new code on Windows still have to concern themselves with UTF-16 vs ANSI vs UTF-8. This is a pattern of behavior at Microsoft, and that's what I'm criticizing.
> I think a more mature attitude would be to accept this, that it came from a time and a place and is a different way of working, rather than call it "wrong".
Look, I understand that mistakes will be made. My criticism isn't that the mistakes are made in the first place, but that Microsoft doesn't appear to have any plan to ever rectify them. The result is a consistent increase in friction over time, the cost of which is mostly paid for by entities other than Microsoft.
One could argue that Microsoft's failure to manage complexity in this way is one of the reasons why Linux is eating their lunch in the server market. Anecdotally, it's just way easier to build stuff on top of Linux, because there's a culture of eliminating old cruft---or at least moving it around so that only the people who want it end up paying for it.
As for your ad hominem arguments about "attitude" and "maturity", I could do without them, thanks. They contribute nothing to the conversation, and only serve to undermine your credibility. Knock it off.