How to properly use UTF-8 in Perl
stackoverflow.com
stackoverflow.com
The key to handling Unicode correctly is to think about the data that comes into your program and the data that leaves your program. If it's text, it has some sort of encoding, and you need to decode those encoded characters before you can use them as text in your program. Similarly, you can't just leak Perl's internal representation of characters to the world: you have to explicitly encode the characters to octets.
The problem that people run into is not knowing that you have to do this (in every programming language), and then deciding which encoding to use. In some cases, it's easy: HTTP includes the encoding in the headers, so all you have to do is get the request, decode according to the headers, and you're set. Similar, when you send an HTTP response to someone, just set the headers, encode the characters, and everything ends up 100% correct.
The problem comes when poor design causes you to assume the character encoding. What encoding are file names on a random removable disk? What encoding is this text file in? What encoding are those database rows? If you don't know, you simply can't process that data correctly as text. But most people assume "something will magically decide for me and it will all work out". Nope, it won't. Don't rely on magic: be explicit.
The reason why things work most of the time is because you treat the data as binary: opaque, meaningless octets, that don't have semantics like "make the first letter capital". This will work if your world is entirely UTF-8-encoded: UTF-8 filenames, UTF-8 files, UTF-8 source code, UTF-8 database results, etc.
The reason why people run into trouble with Perl and not with other languages, is because Perl assumes that when you treat binary data as text, you really have Latin-1 text. It then says, "hey, this is text", upgrades it to Unicode, and then transforms it as such. When your text was Latin-1 (the backcompat case), it all works. When it's UTF-8, though, then you get double encoding: the dreaded "æ¥æ¬èª" instead of "日本語". But if you just tell perl, via Encode::decode_utf8, it will know that you actually have unicode text and not latin-1 text, and everything will work!
Anyway, the solution is to always specify the encoding via Encode::decode_utf8($str) or Encode::decode('my-encoding', $str) when reading data, and to Encode::encode_utf8 or Encode::encode when outputting data (even to the terminal!). The hardest part is making sure your libraries do this (DBD::SQLite does, Catalyst does, LWP does), and making sure that your libraries have enough information to do it right. HTTP is easy, there's a header. But that blob in your database is not easy, and you may have to handle it yourself... because the information about the encoding exists only in your brain, not anywhere the computer can find it.
(Oh, and this only gets you to the "I'm not trashing any information" stage. If you want your Japanese text to sort あ い う え お instead of in codepoint order, that's going to require a module. Even simpler cases will require you to learn about collation and normalization.)
Edit: distilled the key points I raise here into an SO answer: http://stackoverflow.com/questions/6162484/why-does-modern-p...
I won't pretend that Perl's unicode support sucks; Working with Unicode in Perl is far more complicated than it ought to be and rife with bugs and corner-cases. However, working with Unicode properly in just about any language requires a whole lot of careful thought and consideration.
Maybe Java (and by extension JVM languages) is the best you can get right now, I dunno...
On most of those issues being the only one that gets it right is either pointless or retroactively wrong.
PHP has no separate Unicode type, the built-in string type works pretty well for UTF-8. I'd liken its use to the concept of duck-typing (as seen, for example, in Python): if it walks like a duck^W^W UTF-8, swims like a duck^W^W UTF-8, quacks like a duck^W^W UTF-8, it can be handled (consumed and produced) like a duck^W^W UTF-8.
Note in Python Unicode is a separate data type, so the above comparison is related to duck-typing only.
• You need to be aware that most built-in string functions won't work and use mb_ equivalents where available (and not all are available).
• You need to set internal encoding to UTF-8 for some extensions in php.ini.
• You need to normalize all input yourself (thankfully PHP5.3 has Normalizer class. In 5.2 it was a nightmare)
• PCRE library doesn't support all uncode ranges (e.g. {InFoo} is unsupported). All gotchas of regexes apply to PHP/PCRE as well.
The list goes on…
"Binary-safe" strings and UTF-8 cleverness will let you roundtrip characters safely in most cases, but full, correct Unicode support is really hard.
Actually the non-mb string functions will work perfectly fine for almost everything. For example splitting on space or comma, or joining strings will work just fine with UTF-8. The only things that don't work are those that split by character position. Even search works fine.
> You need to set internal encoding to UTF-8 for some extensions in php.ini.
ini_set('default_charset', 'UTF-8');
mb_internal_encoding("UTF-8");
> You need to normalize all input yourselfOnly if you need to compare against it (for example to see if the username was taken). Most of the time you just store the input as is.
For PCRE make sure to add the /u option.
setlocale(LC_ALL, 'pl_PL.utf-8'); # if not set explicitly defaults to system's $LC_ALL, thus fits most of the time
mb_internal_encoding('UTF-8');
mb_ereg_match('[[:alpha:]]', 'Ł') -> true
mb_ereg_match('[[:alpha:]]', '♙') -> false # that's a white chess pawn character, U+2659.
mb_ereg_match('[[:graph:]]', '♙') -> truesetlocale() can be switched at run-time, even multiple times if you want. You could imagine running most of the code with user's preferred locale to output UI or report any errors, switching to another locale to format and send message to a foreign customer, and then back to normal settings to continue with outputing UI, all in one pass of PHP.
Strange that the libs are named PCRE (where "P" stands for "Perl"), then? :-)
You can use:
<FORM accept-charset="UTF-8">
Or just leave it out since the browser will submit all forms using the character set of the page.So yes, UTF-8 in PHP is pretty easy - and I speak from experience, not theory.
So far, you've gotten lucky because your browser wants and provides UTF-8-encoded text, your xterm assumes everything is UTF-8, your filesystem metadata is UTF-8-encoded text, constants in your source code are UTF-8-encoded opaque blobs, and your database results are returned as UTF-8-encoded text. The second you throw in some other source, things are going to break, and you will need a plan for how to handle things. Otherwise, you inadvertently end up with garbage responses or a database where half the data is UTF-8 and the other half is Latin-1. Not good.
Why would I include some other bad source if it's going to mess up my data?
Also a nice property of UTF-8 is that it can be recovered midstream. So if the client sends me bad encodings, he'll get back bad encodings, but it won't harm anything else.
For example a bad encoding followed by a double quote can't do anything to the double quote.
So don't do that.
Don't accept some latin-1 data. Accept only UTF-8, and if you have to accept latin-1 you have two choices: Convert it to UTF-8, or use only latin-1.
> In some cases, it's actually a security problem because of the memory corruption that could result from treating latin-1 as utf8.
What are you talking about? There is no memory corruption. UTF-16 can have serious issues, but UTF-8 is quite resilient.
I think I get it. You are a web programmer and hence always get encoding specified and are never fscked by something else?
Or you've never had the problem of recognizing whitespace (as was the jrockway example) or sorting?
(English isn't my first language, but I'm not good with UTF-8 anyway. I tried to understand why you guys are talking past each other.)
And if you specify one encoding and get a different one then exactly how do you plan to handle that? Obviously you need to know what encoding you get.
Recognizing whitespace works just fine in PHP - you can use ordinary string function and they will work with UTF-8 for that.
Sorting is more complicated, but I typically sort in the database which handles UTF-8 sorting correctly for me.
You should learn more about UTF-8. There is no reason to use any other character set.
coughencodingcough ;-)
There is a good reason to use another in general:
http://man.cat-v.org/plan_9/6/utf
http://man.cat-v.org/plan_9/2/rune
Plan 9 uses UTF-8 for IPC and most simple string handling, but programs that do more involved string processing (text editors, regular expressions stuff etc), use Runes instead. Runes are similar to UTF-32, one Rune directly represents one character. Most of standard library string functions come in two varieties: using UTF-8 and Runes.
UTF-32 is a very bad idea. You have issues with nulls and byte order. And on top of that it's basically useless since it solves no problems at all, while wasting space.
Rune is not UTF-32, but a similar constant-width format. It is not multibyte, but rather single-int-per-character. Gotta stress this point: never interpreted as bytes. Resides only in process memory, not in files, nor for IPC. Held in arrays of of 32-bit ints, not bytes, thus no problem with nulls. There is a trailing zero int, quite like in good old strings.
Individual Rune character access (again, an int, not byte) is O(1); various string processing functions are simple to implement for Runes; character ranges are easy to compare, etc., strings have `predictable' length, etc. etc.
As an internal-only format, it may be improved at will. Indeed, once it was -- from 16 bit short to 32 bit int, to support newer Unicode character set ;-) The change was internal to library and header files; programs were just re-compiled.
A character encoding is an algorithm for reading/writing characters from/to octet streams. A character set is ... a set of characters.
(Unicode has additional semantics; collation, etc.)
As to why you'd waste space with a 32-bit character representation: algorithmic efficiency. UTF-8 is efficient in storage, but it's not efficient for string operations: splitting a string is O(n), finding the nth character is O(n), determining the length of the string is O(n), etc. When each character takes the same amount of memory, though, these operations become constant time.
Finally, nulls don't cause problems because these are character arrays, not null-terminated octet arrays ("c strings"). Most big C apps don't use null-terminated octet arrays anymore; they use something like bstrings, or std::string, or their own homegrown implementation (SVPV in perl, etc.).
Uploaded files etc are also so simple for web programmers?
If you don't then you are stuck in any case, and PHP doesn't hurt anything since it's binary safe. And presumable since you don't know what it is you avoid using any text functions, and treat it as binary.
The software must handle every input in a predictable way. For wrong encoding, it may stop and signal error, or may continue processing the data -- depends on owner's wishes. But must not do anything stupid, no undefined behavior. Where doing something stupid could be giving user information he may not access (like SQL injections; IIRC wrongly-encoded data was successfully used for SQL injections by fooling code that should escape it), damaging data stored in system or saving invalid data if that would break consistency or any work later on.
Two things to be done:
1) instruct browser and other systems that work with your software to send and read UTF-8
2) handle situations when something else is sent instead
With UTF-8, the 2) is mostly done by PHP's mb_...() functions, IF configured correctly.
You can't tell though. If you get data with the wrong encoding it's almost impossible to tell.
At a minimum you'll need to know what language, and from that you may be able to check. But in that case what if the language is wrong?
Nope; it's treated differently by different functions. Like ars wrote, with UTF-8 a lot of properties of old ASCII strings holds. A lot of PHP functions did not need any changes to handle UTF-8 just fine, for example explode() and join(). Other functions do handle actual characters -- when they must.
In contrast, using UTF-16 or UTF-32 would require treating strings as opaque binary blobs in some cases, because those encodings lack certain properties of UTF-8. For example, the old ASCII-age strpos() is reliable with UTF-8, but unreliable with UTF-16.
> Also, "trust the client" as your correctness mechanism? Not very safe.
Never the case here. Standard library is trusted, and the average developer team is not much more secure than maintainers of such libraries. I know my code is not better than the standard library, anyway, so I'll rather rely on it to handle correctness.
You use two kinds of functions: some care about properties of characters -- like mb_strtoupper(); others just work on strings -- like explode(). Only the former kind matters in aspect of correctness.
When such function tries to operate on the mis-encoded character, it follows the Unicode standard: replaces the character with U+FFFD, a.k.a. `REPLACEMENT CHARACTER'. Result is valid UTF-8, with one character replaced with the marker. Any subsequent processing will go just fine, unless you want to stop and warn the user. The condition is easy to detect: there is one unique byte sequence for the REPLACEMENT CHARACTER, and it's clear where in the string the problem was.
UTF-8 was designed to be self-synchronizing in case of such errors.
> This isn't "handling unicode", it's "doing nothing"
That ``doing nothing'' was my original point; thanks for re-iterating? You -- the developer -- do nothing, in particular you don't insert boilerplate code everywhere (and what if you forgot one rarely used execution path?). The functions that must handle individual characters perform the mundane work. Unicode is handled where it has to be -- and only there.
> (...) your xterm assumes everything is UTF-8, (...)
Nothing's assumed there; the process is instructed to do so via configuration. On POSIX platforms, that's environment variables: $LC_ALL and related. The main point of such configuration is that you can adjust whole system and make programs speak common language to one another. And they do -- following your instruction.
Unless explicitly instructed to do otherwise -- when interfacing with remote systems and a different encoding has been agreed upon.
> The second you throw in some other source, things are going to break, and you will need a plan for how to handle things.
When you extend software to interface another software, you either configure both to use the same protocol (here: encoding), or implement protocol (here: encoding) negotiation. You -- the developer -- configure or implement it once, as part of creating the interface between them.
Plain and simple: by not handling Unicode, you are not handling Unicode. Echoing back UTF-8 when you get UTF-8 input is handling binary, not handling text. Yes, I know this works for your particular application, but it not correct in general. It won't work at all if you ever have any data sources or sinks that are not UTF-8, and it won't work if you actually have to treat the data as text ("match all characters that are whitespace").
Yes, some PHP functions take binary blobs that contain UTF-8 octets and return correct-looking answers. That's all they do, though; that's handling UTF-8, not handling Unicode.
In Perl, characters are abstracted away from their character encoding. Perl strings are not UTF-8 or Latin-1 or anything; they are character strings. That makes producing correct results much easier; if you try to write a Perl character to a file, it will say, "hey, you can't expose characters to the outside world, you need to pick a character encoding, encode, and output that". Your PHP method just assumes everything is UTF-8, which is simply not a safe assumption to make.
(Consider the case of binary data: if you start treating binary data like it's UTF-8 text, you break the binary data.)
Also, regarding LANG, yes, that will change your terminal's encoding. So if you take one of your UTF-8 blobs and print it when the LANG is en_US.UTF-8, it will show up. But change the locale to en_US.ISO-8859-1, and it doesn't work anymore. That's because you can't just print binary blobs to the terminal and expect it to interpret that as text.
Nope. UTF-8 is text. We don't live in an ascii world anymore, text requires all 8 bits and a variable number of bytes per character.
> It won't work at all if you ever have any data sources or sinks that are not UTF-8
So don't do that. (Or convert as necessary.)
> and it won't work if you actually have to treat the data as text ("match all characters that are whitespace").
Wrong. Actually it works just fine.
> Yes, some PHP functions take binary blobs that contain UTF-8 octets and return correct-looking answers. That's all they do, though; that's handling UTF-8, not handling Unicode.
So? Since that's completely sufficient to do the job why do you want more?
> In Perl, characters are abstracted away from their character encoding.
Why? Why would I want to do that? Is this complexity for the sake of complexity?
> Your PHP method just assumes everything is UTF-8, which is simply not a safe assumption to make.
Yes, actually it is a safe assumption for the simple reason that the only types of strings I handle are UTF-8. In todays world there is no reason to use anything else. All other character sets are obsolete and should never be used.
And if you do need to handle one then convert as necessary, or define your app to handle only that character set. Programs that need to handle multiple character sets in the same applicate are extremely rare. There is no need for extreme complexity just to handle the rare case.
> Consider the case of binary data: if you start treating binary data like it's UTF-8 text, you break the binary data.
So don't do that.
> But change the locale to en_US.ISO-8859-1, and it doesn't work anymore.
So don't do that.
There just isn't One True Way about validating user input; please don't get religious there. For sake of an example, if an user stated his birth year as 1905, do you accept it? The answer is, ``depends on who pays the bill'', in other words, on business rules. Sometimes it makes sense to accept and store data that's unprobable or flat out wrong -- perhaps no entity cares, ever; perhaps the user knows better than us, developers, this time. Sometimes data must fit strict criteria -- say in processing legal documents -- or else you raise big red error message and rollback transaction. And raising an error on that REPLACEMENT CHARACTER is trivial. Do so when you need.
However, don't let rare cases dominate your application with boilerplate code. That's time/effort/money down the drain.
That REPLACEMENT CHARACTER ``is valid UTF-8'' was only to point out there won't be any undefined behavior. Thus no weakening of security, which was your point.
> Your PHP method just assumes everything is UTF-8, which is simply not a safe assumption to make.
That was exclaimed twice in above posts [1][2].
As a developer I pick and choose all the data sources and sinks the app interacts with. It's either an UTF-8-ony entity (source code, the files created by my software, etc.), or my code negotiates some common encoding at the start of communication (the browser, the database, etc.) and perform sconversion if needed. Once, in the implementation of interface to the source/sink.
Conversion to UTF-8 is simple, takes O(n) time and O(1) memory.
If there is no metadata describing content's encoding, no amount of abstraction will help, ever. You may report an error or try to go on using defaults; again depends on needs of users/owners/whatever.
I don't know how you arrived at binary data getting mixed up with textual. The only shared matter is the `string' datatype itself; all else is separate. The execution paths are different: very, very different classes read, mangle and write back user-uploaded binnary files (to re-scale images etc.) and different ones handle text submitted by user, fetched from DB etc. Even the name of uploaded file is held in different variable than the content. Normally most of the files have metadata associated stating the type of the content, either in some header or implied by format. Again, the developer chooses which files are handled by what code, if at all.
> In Perl, characters are abstracted away from their character encoding.
Good in theory, but usage -- as shown in OP -- is elaborate. Error-prone, one could say.
Abstracting stuff away is means, not an end in itself. The elaborate OP suggests the means may hamper achieving the end in some cases. What are you gonna do?
> That makes producing correct results much easier;
Let's see if others agree with you: http://stackoverflow.com/questions/6162484/why-does-modern-p...
----
[1] > and convert only when explicitly requested by user or remote system to do so.
[2] > When you extend software to interface another software, you either configure both to use the same protocol (here: encoding), or implement protocol (here: encoding) negotiation.
EDIT: on a website, what is worse: mojibake (garbled characters) or an error message? May depend on audience: AFAIK some internet users ``understand'' mojibake in some way (may for example attempt to adjust browser's encoding settings and actually get correct text, yay!), but most hate cryptic error messages. God forbid an error message tell them they sent ``illegal character'' [3] ;-)
Any distration (modal dialog, workflow interruption) during user registration, order submission etc. may end up scaring away a good deal of potential customers. May be easier for site's owner to actually manually correct the profile/order/whatever, or ask the user to correct it later on.
[3] http://forums.thedailywtf.com/forums/p/24359/253469.aspx#253...