Unicode is, in 2012, still “cutting edge.”
golem.ph.utexas.edu
golem.ph.utexas.edu
http://www.unicode.org/versions/Unicode6.1.0/
The latest published version of the standard (5.0) runs to almost 1500 pages:
http://www.unicode.org/book/aboutbook.html
(Best quote from that page: “Hard copy versions of the Unicode Standard have been among the most crucial and most heavily used reference books in my personal library for years.” -- Donald Knuth)
People think that Unicode support is just a matter of implementing multi-byte characters, but it's so much more: you've got collation rules, ligatures, rendering, line-breaking, punctuation, reading direction, and so on. Any technical standard that aims to cover all known human languages is going to be a little bit complex.
Half of those are irrelevant to MySQL or any other database. Those are front end problems. Even reading punctuation, text direction, and the like will only be important in more advanced collation orders (as opposed to just binary ordering, which he was using).
http://stackoverflow.com/questions/5567249/what-are-the-most...
Is it "excusable" that MySQLs implementation of UTF-8 isn't standard? That's a judgment call (they are up-front about it in the docs). But given that most unicode characters "in the wild" lie in the BMP, I can see how they'd make that trade-off. There might well be a technical limitation lurking somewhere in the database internals that made 4-byte characters a problem.
The options are simple: implement UTF-8 correctly and call your implementation UTF-8, or implement just the BMP and name accordingly. They did neither: effectively, they lied to end-users. That's deeply, deeply problematic.
Yeah, that's not exaggeration at all. Because it's not as if they very clearly document exactly what they support:
Bonus points for the documentation brazenly ignoring that they're implementing something that's not spec-compliant and naming it like it is.
And while, as a Postgres user, my tone here may be a little snide, I also say this with grudging respect: I think there is a point at which implementing n% of a feature X and calling it X (rather than MaybeX or MostlyX) does give you some momentum and practical compatibility that you wouldn't have otherwise. Is it dishonest to hide the limitations regarding the edge cases in some documentation no one will read? Maybe. But will providing the feature solve more problems than it causes? Quite possibly.
I don't agree with MySQL's decision with respect to UTF-8, but I do understand it.
While I don’t know the history of MySQL, it seems to me that when they implemented it, their implementation was indeed in compliance with the standard (Unicode 3).
The standard has since grown from 16 to 32 bit code points.
Why MySQL had to introduce a new name for the UTF-8 encoded tables that can contain 32 bit code points is strange, but I assume there is a technical explanation (probably having to do with binary compatibility with existing tables / MySQL drivers or similar).
Edit: you can try Symbola at http://users.teilar.gr/~g1951d/ (and yes, I keep editing this post :p)
I installed it and now I see it as well.
That right there is the problem.
Other DBs, when you do something they can't handle, bail with a full-fledged error that stops what you're doing. MySQL doesn't do that for quite a few data-losing cases, and that's incredibly dangerous.
Most of those pages are the just repertoire: a printout of a long uneventful table that lists all the available code-points. For each code-point is gives you a reference pre-rendered glyph, the languages where you can find it, the kind of character it is (numeric, alphabetic, symbol) and so on.
From an implementer point of view there are about 200 pages of interesting and extremely detailed stuff. The rest of the pages can be downloaded as text tables from the Unicode site and its companion sites.
Wow. That is pathetic.
1, is faster, better tested and string handling (it's a database!) is much faster but it only handles the 65000 most common characters
2, this one can handle upside down characters from a 1930s paper on formal logic in Turkish. But is slower for all other cases and we haven't really tested it as much,.
Do you have a redundant,self powered , asteroid impact proof internet connection? No? Pathetic !
"UTF-8 (UCS Transformation Format—8-bit[1]) is a variable-width encoding that can represent every character in the Unicode character set," says Wikipedia. The UTF-8 implementation in MySQL does not meet this definition because it cannot represent every character in the Unicode character set.
Unicode 2.0 introduced multiple planes, i.e. more than 65536 characters. That was in 1996. If that was the case, then MySQL has had more than one-and-a-half decades to introduce multiple planes and seems to have done so less than a year ago. I disagree with being 'redefined out from under them', when it was defined a year after MySQL started, at a time when it probably didn't even have Unicode support yet anyway.
1996.
China even made it a legal requirement for computer systems in 2000, through mandating GB 18030.
There's the Private Use Area if nothing else. There is NO excuse to not support anything other than the BMP. Adding support is trivial unless you have been using UTF-16 in the erroneous belief that it's two bytes long always (in which case you've really been using UCS-2).
More than 2 years ago. March 2010.
So, why would they want to store things internally as UCS-2? Or rather, why should they?
Or do you mean in-memory being different from the file format AND different from the I/O format? That doesn't sound terribly efficient.
Not exactly. A varchar will store it as is, but a char column will allocate a fixed 3 (or 4) bytes for each character.
All data stored in memory (for sorting and such) is always as char, even if it started as varchar.
So by allowing 4 bytes per character they use more memory.
The hell of it is, UTF-8 expands gracefully to the astral planes; it's UTF-16 that you need to worry about, either because the people designing the software never heard of surrogate pairs, in which case they didn't give you UTF-16 but UCS-2, or implemented surrogate pairs incorrectly.
Luckily Perl's Unicode support is fantastic, and saved my ass
I worked on a product for a couple of months geared at minority languages in developing countries - doing linguistics work etc. It was a pain to support Unicode, because there's lots of code points, lots of weird cases (what's capital?), ICU is a good lib but it's not up to date to the latest Unicode version and it's C/C++ (and thus a pain in C#). Oh, and there's the Private Use Area where characters go while Unicode decides to include them or not...
Java, for instance, implemented 2-byte encoding and uses surrogates for the higher planes, which means you get the worst of both worlds... You double the size of ASCII text (that is, half the speed of bandwidth-limited operations on text) and you've still got a variable length encoding... but you've got lots of methods and user-written code that assume that the text is fixed length encoded. what a mess
That said, if you think Unicode is a pain, try storing and retrieving "𝒜" in any other encoding. I'll stick to Unicode, thanks. :)
The only real issue is handling bad input, as you never get an error with decoding e.g. ISO-8859-1; for a lot of applications you need to handle potentially malicious input, so you can do it there, but even for trusted input, there is a lot of really-broken external programs that output "UTF-8" or "UTF-16" (scare quotes intentional).
I really think a lot of the problems with unicode is that a lot of languages/libraries try to handle it transparently, and that just doesn't work; encoding/decoding is part of dealing with external formats, and trying to do it transparently means that it will fail unexpectedly.
I think that depends heavily on what programming language and frameworks you are using.
Slides for that talk and two other Perl-related Unicode talks are at http://training.perl.com/OSCON2011/index.html
Specifically the talk titled "Unicode Support Shootout: The Good, The Bad, & the (mostly) Ugly"
It's a year old now, but it's still relevant. It gives a very detailed look at unicode support across JavaScript, PHP, Go, Ruby, Python, Java, and Perl.
But storing a stream of UTF-8 and retrieving it on command is not remotely difficult. You almost have to go out of your way to screw that up.
The same experiment with cat in place of textmate works fine, so it's textmate that is buggy.
According to "od -x1", textmate is writing:
0000000 ed a0 b5 ed b6 99 ed a0 b5 ed b6 8a ed a0 b5 ed
0000020 b6 98 ed a0 b5 ed b6 99
So textedit is right to complain.It turns out that this is such a common mistake that there's even a name for this encoding, CESU-8: http://en.wikipedia.org/wiki/CESU-8
dos2unix
all the time. How sad.However, it's true that Unicode is (relatively speaking) very new for such a fundamental technology. Support in applications still varies widely. I wouldn't characterize it as cutting edge though, since we have many mainstream programming languages built using Unicode internally.
TFA notes that this is not fixed, the `utf-8` mysql encoding still isn't utf-8. And as TFA also notes related technologies (aka drivers) may not be compatible with it (the example he uses, mysql2 for Ruby, still hasn't had an official release supporting utf8mb4[0])
> it's true that Unicode is (relatively speaking) very new for such a fundamental technology
That's becoming quite hard an argument to swallow when encountering astral planes issues in 2012 when Unicode 2.0 was introduced in 1996.
> That's becoming quite hard an argument to swallow when encountering astral planes issues in 2012 when Unicode 2.0 was introduced in 1996.
I don't get your argument. MySQL was also released around that time and we don't call it "cutting edge" because we found a bug. There are bugs in old stuff all the time but (most) people don't throw a fit.
Not supporting astral planes and saying you're supporting utf-8 is not a bug, it's a lie.
Creating a new character set `utf8mb4` was the right thing to do, as annoying as it is. Just clearly label the `utf8` collation as 'deprecated' in the docs or something.
Or they could just have implemented it correctly to start with, considering unicode "support" was introduced in mysql 4.1.
In 2005.
> Who knows what software out there depends on it breaking on 4-byte characters, or whatnot.
Then again, mysql routinely drops and corrupts data anyway, I'm sure its "users" could have dealt with it corrupting data slightly less than before.