A Localization Horror Story: It Could Happen To You
search.cpan.org
search.cpan.org
One architectural takeaway suggestion I've learned over the years, which is not obvious when reading that article:
"You should not assume that you can generate any part of any string visible to the user without the full context."
Whenever you design a localizable application, it isn't enough to provide a string that can be translated. You have to allow for delegation to the most specific piece of code dealing with the string, because only that piece of code will have the appropriate context to properly produce the string.
This means your code can't just assume it can generate strings somewhere deep inside the guts of a library. The programmer writing the final application that uses your library needs to be able to generate/override those strings on a per-case basis in the actual code that displays them to the user. The strings might be different between two UI windows.
Trust me, I know. I'm Polish. Few languages are as insane as my native tongue. If you don't believe me, take a peek at this concise 252-pages long introduction to Polish numerals: http://www.amazon.com/Liczebnik-grammar-numerals-exercises-l...
I'm learning Polish and the crazy numbers business has really opened my eyes - doesn't the number 2 have something like 16 or 17 possible forms?
Hrm, not sure I entirely agree with you there, and I'm both stupid and living as a foreigner in Poland :) The writing system is fantastic in that regard, you can look at a word and have a very good idea of how it should sound.
For me, there's only one really difficult sound and that's ń: the best description I've heard of it is the first N in onion. But it's still a bugger to pronounce and hear the difference between, say, koń (horse) and koni (horses).
I don't think I've ever met any American who can pronounce ы - though it's fun watching them try.
My Russian tutor says that my accent (which is mimicry based on Russian from songs, films, etc.) sounds almost native EXCEPT for how I pronounce ы. I can't quite get it down correctly. :(
Compare that to french, chinese or english...
http://en.wikipedia.org/wiki/Dual_(grammatical_number)
quoted // Of the living languages, only Slovene and Sorbian have preserved the dual number as a productive form. In all of the remaining languages, its influence is still found in the declension of nouns of which there are commonly only two: eyes, ears, shoulders, in certain fixed expressions, and the agreement of nouns when used with numbers.
The guide was somewhat whimsically named bluepill.doc and subtitled Welcome To The Real World. You have no idea how deep this rabbit hole gets. I did this for years and I am regularly surprised by novel, hard problems. It is like security. (It even intersects with security sometimes: since approximately no application developers actually understand encoding issues, there are virtually boundless classes of vulnerabilities arising from their (mis)understandings not matching technical reality.)
(I only found out later that the blue pill was the escape-back-into-comfortable-fantasy option. Whoopsie.)
* In the code, messages are specified abstractly, e.g.
print getMessage('found_x_files_in_x_dirs', $fileCount, $dirCount);
* Languages each have their own class with a 'convertPlural' function that maps the quantity to the forms. So in english, that function might be simple, for Arabic, it's complex: http://svn.wikimedia.org/svnroot/mediawiki/trunk/phase3/lang...* Lexicons use a simple wiki markup to define the different forms of their language. To illustrate that the arguments don't have to be used in order, I did it in the reverse of how the code passes arguments.
'found_x_files_in_x_dirs' =>
"I searched $2 {{PLURAL:$2|directory|directories}}
and found $1 {{PLURAL:$1|file|files}}"
So for a language like Arabic you write a similar pipe-delimited list of forms. You just have to know how to lay down the six different forms in the order that LanguageAr.php defined.Note how this side-steps most (but not all) complicating issues like case or gender, so you don't have to mark it that way in the lexicon. If the word is used in the feminine gender, accusative case plural in the sentence, that's what the translator writes.
All this is mediated with the amazing http://translatewiki.net/ website, run mostly by volunteers.
You can define the number of plurals and rules to select a plural case. Then you have as many translations as plural forms in your translation (po) file.
Arabic for instance has 6 plural cases with the following rules [2]:
nplurals=6; plural= n==0 ? 0 : n==1 ? 1 : n==2 ? 2 : n%100>=3 && n%100<=10 ? 3 : n%100>=11 ? 4 : 5;
See also http://wiki.amule.org/index.php/Translations#Plural_forms for an example of both a rule and the resulting code in the po file.[1] http://www.gnu.org/software/gettext/manual/gettext.html#Plur... [2] http://translate.sourceforge.net/wiki/l10n/pluralforms
Imagine your phrase/function is called:
"Found %n1 matching files in %n2 directories"
You could pattern match for one particular language like this:
%n1 == 0, %n2 == 0
%n1 == 1, %n2 == 1
%n1 > 1, %n2 == 1
... and so on ...
With the matching being any (simple?) boolean function of the operator, applied in order. At the point where this becomes too cumbersome, you could fall back to proper code. I bet that this would be much easier to use for translators, with an optional fallback to a programmer if it gets too complex to spell out all combinations.In general, there is no easy way out — and you have to allow for exceptions. Please see my other comment, about an architectural takeaway.
Japanese, for example:
http://en.wikipedia.org/wiki/Japanese_counter_word
Take the counter 本, for a commonly-used example. That's used for: "Long, thin objects: rivers, roads, train tracks, ties, pencils, bottles, guitars; also, metaphorically, telephone calls, train or bus routes, movies (see also: tsūwa), points or bounds in sports events. Although 本 also means "book", the counter for books is satsu (冊)."
But the real problem is that we have a different problem in each language. Every single language has its own way to do things and that's where we run into trouble, because you need generic ways of doing things and there really isn't one for all languages. You essentially need one function to create text for each language, even though some of them can share helper functions, like that one for numerals.
{l}Found {$d} directories{/l}
and the replacement for each language could have its own switch on $d to decide how to translate it.One (unavoidable?) downside is that the translators have to know some basic if/then/smarty syntax, and if they mess it up your template won't compile. Also you have to trust them somewhat, since they essentially get to execute PHP on your webserver.
(unless I'm misunderstanding something in your templating scheme)
{switch $d}
{case 1} foo {$d}...
{case 2} {$d} bar..
{default} ba{$d}z
{/switch}
We provided the translators with some basic examples of if/then/switches and so on, they were free to do all sorts of crazy things (especially in polish).This is a classic case of "make things as simple as possible... but no simpler". You've achieved the first bit. To go any simpler would lose vital granularity.
After all, there's a vast amount of open source software that has been localized with gettext without problems.
In my experience, most translation strings are straightforward, and the corner cases as mentioned in the article don't happen very often. If you can tolerate a few hacks in your code, or a few extra translation messages, you don't have to go all the way down the rabbit hole. The perfect is the enemy of the good in this case.
Also, it requires hard decisions, like should this language that has falled 20% out of date be disabled, or is it better for the program to startle the user with English 20% of the time? (Which 20%?)
And finally, after all this work and pain, technical computer users will complain that they prefer the English version because their native language has clumsy terms for computing terms, or the translation to their native language is not idiomatic enough, or whatever. And you as the coordinator can't begin to judge translation quality unless you're fluent in N languages.
I've been lucky to have very skilled people handling the translation coordination in some free software projects I've been involved in, and localization has still had most of these elements of pain in most of them.
(Or if you prefer a proper rant: http://kitenet.net/~joey/blog/entry/on_localization_and_prog... )
What I wanted to point out was that the article puts focus on a problem which often times is not that big of an issue.
I've mostly been doing web development and localization of web applications, and I guess this makes the build process a bit simpler than what you describe.
I have had experience localizing web applications and it was not pleasant. We had to design the site from the start to account for multiple languages. It complicated the design. The alternate languages were, at best, several days behind the original English. A couple weeks was more typical. Localizing an existing site where we had not planned for it from the start was not practical.
Personally, I would recommend outsourcing localization if that is an option. It just isn't worth it to do it yourself. It is time consuming. More importantly, the time spent localizing is time not spent on your real business. Then again, I am completely biased because I make my living at a company where we provide this exact service.
(It may be a shameless plug, but it's on topic and might be useful to someone... The name of the company I work for is in my profile.)
printf("Directories scanned: %g", $directory_count);
That's precisely the problem the article addresses :) You could just about get away with it in English where the only number that won't work is 1, but what about Polish where "Directories" will take a different form if the number ends in a 2, 3 or a 4 but isn't 12, 13 or 14?
Thanks for the clarification. My point was that a workaround in English may not be a workaround in every other language in the world.
Good ideas: 0
Usability is everything: True
Which is one reason to speak or code in more than one language.
Game translators need to know all sorts of metadata about the speaker and listener characters' "social context", such as gender, age, and "honor". Is the old king speaking to a young peasant girl or an entire village? Is the young prince speaking to an old peasant woman or to his old grandmother? Is the grandmother speaking to her son, who happens to be the king?
1. Just don't bother internationalizing. I think that's often not a bad solution for small software businesses anyway. It's not much use making a software package that works in French/German/Russian/Chinese/Swahili unless you have the language skills or partnerships to sell and support the software in those languages as well.
2. Design the software such that messages are whole sentences or phrases that stand alone and can be translated one-to-one. Nothing fancy like the %g type stuff for number inserts. Keep it simple stupid.
What would you say if Japanese were trying to sell Japanese software with Japanese-only indications/manual ? Do you think many American would buy it ?
Dammit, there has got to be a way for me to make money there.
Directories scanned: %g
Files found: %g, Directories with files: %g
seach result: directory: NN, file: NN.
as in "search result: directory: 1, file: 0." or "search result: directory: 4, file: 23."
ie: nouns only, singular form, no verbs, no plural, etc....
Sure, it does not look "good" but its probably much more easy to translate.
Worse is better
Is this the best we, as an industry, can do?
form.applyPattern( "There {0,choice,0#are no files|1#is one file|1<are {0,number,integer} files}.");
http://download.oracle.com/javase/6/docs/api/index.html?java...