Resetting PHP6 (or: Unicode claims another victim)
lwn.net
lwn.net
For search-and/or-replace (or anything that can be done with regexes), I'm pretty sure that UTF-16 has no advantage over UTF-8, as you can make the state machine operate on the bytes directly.
Thank you for that clear and concise explanation of the dangers of using UTF-16.
Yes, I know that wasn't your intention, but it was the end result. One of the most dangerous library failures you can have is a function that works 99.99% of the time. Or in this case, 100% of the time on the input the English-speaking developer provides but distinctly less than 100% in the field.
In this specific case, you can't actually optimize anything because all your optimizations are bugs. You can't just divide by two for character count; that's not an optimization, it's a bug. You can't just multiply by two for a substring operation, because you might chop a character in half, that's a bug. And so on. You'd need a separate type that indicates you've scanned the string to verify it never has split chars and now you might as well be on UCS-2, and that has its own dangers w.r.t. working 99.99% of the time.
Much better to use UTF-8, where the dangers are much more apparent, all you have to do is leave the base ASCII case and you're testing UTF-8. Even I, an English-speaking developer, manage to test that case (once I know it exists, anyhow). There's still ways you can screw up but you're off to a much better start.
As a Java developer, I'm very happy with the compromise that UTF-16 strikes. For day-to-day unicode use, UTF-16 covers all the bases. In the rare cases I need to step outside of UTF-16 to the higher planes, it would work transparently for me up to the point where I start slicing strings naively.
To be honest, I've never actually had to use anything outside of the Basic Multilingual Plane. That's not from lack of breadth in my day-to-day job either - I wrote the first implementation of our web crawler/search engine at DotSpots and had to deal with fetching and indexing pages in languages from most of the BMP (mainly English, Russian and Chinese).
Anyways, I strongly disagree with your assertion that Java was "generally considered" to have "made a mistake" by using UTF-16. It's made my life far easier and the memory costs aren't an issue on today's machines. We don't live in an ASCII world most of the time and storing strings internally as UTF-8 completely ignores this fact.
[edit]
Just so it's clear: UTF-8 only wins when you are representing pure ASCII text. For everything else, it either breaks even or loses. For most Chinese text, UTF-8 is 50% larger:
For characters <= U+0080 (ie: ABC), you win by one byte.
For characters > U+0080 but <= U+07FF (ie: Ȁɐ), you break even at two bytes.
Everything higher than U+07FF in the BMP, < U+10000 (ie: 丂且⬄☃), you lose (by one byte).
For characters >= U+10000 you break even again at four bytes.
If you really want to optimize for ASCII text in Java, there's always byte[] and you're free to wrap a CharSequence around it.
In general practice, however, this is a non-issue. Characters outside of the Basic Multilingual Plane are not in common use, especially on the web. It's not a perfect programming practice, but it's very pragmatic.
Making everyone pay the development tax of variable-sized characters for any sort of multi-lingual code just means that more code will be written incorrectly.
This is a red herring. Because of combining characters, it is rarely valid to slice between Unicode code points, regardless of encoding. Even in the BMP, a semantic symbol can be composed of multiple code points.
Unfortunately since PHP developers don't like to reinvent the wheel, PHP is also based on a large number of 3rd party libraries that aren't unicode aware.
A file is very very very simple to convert once. Tell your developers, "If you don't encode in UTF-16, you will have a performance penalty. Set your file encodings as UTF-16 too. You weren't doing complex internationalization work before, it's really not that big a deal."
I worked with someone who had been on the ICU project, and he argued that UTF-16 is the best compromise for most cases. If you're working primarily in the western character set, UTF-8 is attractive, but that comes at the expense of others.
And frankly, if you don't roll it yourself, what are you going to use other than ICU?
GNU libunistring: http://www.gnu.org/software/libunistring/
UTF-8 is nearly ideal for transmission and storage and is fairly robust for manipulation, albeit it can be the least straightforward to implement(not that app developers actually have to implement it).
UTF-32 is probably most useful as an internal optimization for tasks that can really benefit from a straight scan/cut/paste over even-sized memory cells. You wouldn't store it in your database, but you might want to make use of it in a document editor, for example, to speed up search+replace type operations.
UTF-16 is still substantially more heavyweight than UTF-8, but it can't be optimized into straight memory cells like UTF-32 without breaking the spec. So - unless your needs are extremely specific and you discover a sweet spot in UTF-16 after extensive profiling - it's just not a likely candidate.
Why is it dysfunctional?
- every discussion leads to bikeshedding (and almost none of the bikeshedders actually commit code to the Zend engine)
- there are 'rules', but they don't apply to most people (ie the 5.4 thing in the article)
- no firm hand to guide them (Rasmus has deliberately not provided this)
- the mailing list has a complete lack of civility
- highest concentration of poisonous people to non-poisonous that I have ever seen
- votes for everything
- patches are not discussed, either pre or post commit, so the code is bad, and people won't work on it.
I was so glad to be the hell out of there.
PHP is a band aid, but as a band aid it served it's niche remarkably well, imagine if clojure or some other better designed language would attract such an enormous following and would be so easy to deploy.
Even today mod_php runs rings around mod_wsgx in that respect (and it's already a lot better then mod_python).
PHP has tons of shortcomings, but it is relatively good at what it does, and that's what drives it forward, not the people behind the project. Say python and everybody things 'Guido van Rossum', say Clojure and 'Rich Hickey' jumps to the foreground.
As long as I've been using PHP I would have a hard time coming up with the full name of it's lead developer. That 'lack of personality' and the chaotic development process may actually contain some hidden benefit.
Absent a strong leader there will be many people pushing and pulling in different directions, it may have gone too far but there is a lesson in there somewhere.
And if I were Zed Shaw, this is the part where I'd threaten to kill you if you don't meet my demands.
How on earth did they decide this was a good idea? The points about getting rid of register_globals and safe_mode are great, but why add a feature to a programming language that is highly likely just to result in lots of awful code?
By the way, the docs on that page are way out of date so if anyone is interested feel free to contact me.
EDIT: I'm an idiot, for tail calls of course it's a bit more complicated than just using a loop.
I've seen this problem "solved" with a do while false and a break, but isn't that even hackier and less expressive?
At this point I think goto-phobia is well understood enough that adding it for occasional use wouldn't ruin the language.
[Direct to comic] http://imgs.xkcd.com/comics/goto.png