As a Java developer, I'm very happy with the compromise that UTF-16 strikes. For day-to-day unicode use, UTF-16 covers all the bases. In the rare cases I need to step outside of UTF-16 to the higher planes, it would work transparently for me up to the point where I start slicing strings naively.
To be honest, I've never actually had to use anything outside of the Basic Multilingual Plane. That's not from lack of breadth in my day-to-day job either - I wrote the first implementation of our web crawler/search engine at DotSpots and had to deal with fetching and indexing pages in languages from most of the BMP (mainly English, Russian and Chinese).
Anyways, I strongly disagree with your assertion that Java was "generally considered" to have "made a mistake" by using UTF-16. It's made my life far easier and the memory costs aren't an issue on today's machines. We don't live in an ASCII world most of the time and storing strings internally as UTF-8 completely ignores this fact.
[edit]
Just so it's clear: UTF-8 only wins when you are representing pure ASCII text. For everything else, it either breaks even or loses. For most Chinese text, UTF-8 is 50% larger:
For characters <= U+0080 (ie: ABC), you win by one byte.
For characters > U+0080 but <= U+07FF (ie: Ȁɐ), you break even at two bytes.
Everything higher than U+07FF in the BMP, < U+10000 (ie: 丂且⬄☃), you lose (by one byte).
For characters >= U+10000 you break even again at four bytes.
If you really want to optimize for ASCII text in Java, there's always byte[] and you're free to wrap a CharSequence around it.
In general practice, however, this is a non-issue. Characters outside of the Basic Multilingual Plane are not in common use, especially on the web. It's not a perfect programming practice, but it's very pragmatic.
Making everyone pay the development tax of variable-sized characters for any sort of multi-lingual code just means that more code will be written incorrectly.
This is a red herring. Because of combining characters, it is rarely valid to slice between Unicode code points, regardless of encoding. Even in the BMP, a semantic symbol can be composed of multiple code points.
Unfortunately since PHP developers don't like to reinvent the wheel, PHP is also based on a large number of 3rd party libraries that aren't unicode aware.
UTF-8 is nearly ideal for transmission and storage and is fairly robust for manipulation, albeit it can be the least straightforward to implement(not that app developers actually have to implement it).
UTF-32 is probably most useful as an internal optimization for tasks that can really benefit from a straight scan/cut/paste over even-sized memory cells. You wouldn't store it in your database, but you might want to make use of it in a document editor, for example, to speed up search+replace type operations.
UTF-16 is still substantially more heavyweight than UTF-8, but it can't be optimized into straight memory cells like UTF-32 without breaking the spec. So - unless your needs are extremely specific and you discover a sweet spot in UTF-16 after extensive profiling - it's just not a likely candidate.
For search-and/or-replace (or anything that can be done with regexes), I'm pretty sure that UTF-16 has no advantage over UTF-8, as you can make the state machine operate on the bytes directly.
Thank you for that clear and concise explanation of the dangers of using UTF-16.
Yes, I know that wasn't your intention, but it was the end result. One of the most dangerous library failures you can have is a function that works 99.99% of the time. Or in this case, 100% of the time on the input the English-speaking developer provides but distinctly less than 100% in the field.
In this specific case, you can't actually optimize anything because all your optimizations are bugs. You can't just divide by two for character count; that's not an optimization, it's a bug. You can't just multiply by two for a substring operation, because you might chop a character in half, that's a bug. And so on. You'd need a separate type that indicates you've scanned the string to verify it never has split chars and now you might as well be on UCS-2, and that has its own dangers w.r.t. working 99.99% of the time.
Much better to use UTF-8, where the dangers are much more apparent, all you have to do is leave the base ASCII case and you're testing UTF-8. Even I, an English-speaking developer, manage to test that case (once I know it exists, anyhow). There's still ways you can screw up but you're off to a much better start.
A file is very very very simple to convert once. Tell your developers, "If you don't encode in UTF-16, you will have a performance penalty. Set your file encodings as UTF-16 too. You weren't doing complex internationalization work before, it's really not that big a deal."
I worked with someone who had been on the ICU project, and he argued that UTF-16 is the best compromise for most cases. If you're working primarily in the western character set, UTF-8 is attractive, but that comes at the expense of others.
And frankly, if you don't roll it yourself, what are you going to use other than ICU?
GNU libunistring: http://www.gnu.org/software/libunistring/