Fluent 1.0: a localization system for natural-sounding translations
hacks.mozilla.org
hacks.mozilla.org
For a slightly contrived example to demonstrate this, let's say you have a string like this:
"Please click here 7 times to confirm"
Where you want to make the "click here 7 times" look like a link by wrapping it in a <a> tag, or just styled differently using a styled <span>.
Using something like react-intl, which is what I've used in the past, you'd have to do something like this:
<FormattedMessage
id="confirm"
defaultMessage={`Please { confirmLink } to confirm`}
values={{
confirmLink:
<a>
<FormattedMessage
id="confirm-link"
defaultMessage={`click here {clickCount, number} {clickCount, plural,
one {time}
other {times}
}`}
values={{ clickCount: 7 }}
/>
</a>
}}
/>
If then some language happens to require a completely different sentence structure that changes the ordering such that the "to confirm" part needs to be interleaved somewhere in the middle of the "click here 7 times" message to sound fluent, this would not be able to accommodate that.I'm wondering how people generally deal with this, in React and elsewhere.
In the React case, a component oriented approach to i18n could maybe look something like this:
const DefaultMessage = ({ clickCount }) =>
<span>
Please <a>click here {clickCount}
{pluralize(clickCount, {one: "time", other: "times"})}
</a> to confirm
</span>
// some theoretical language that requires putting "to confirm" between
// "click here" and "7 times", using English for clarity
const MessageInSomeOtherLanguage = ({ clickCount }) =>
<span>
Please <a>click here to confirm {clickCount}
{pluralize(clickCount, {one: "time", other: "times"})}
</a>
</span>
// some theoretical component that renders different components
// based on the language and passes through props
<FormattedComponent
id="confirm"
defaultMessage={DefaultMessage}
props={{ clickCount: 7 }}
/>
This feels a lot more elegant and flexible to me. Though it would make it more difficult for non-technical folks to contribute to translations, which might not be much of a concern if your company has the resources to support localization teams in-house. Am I overlooking any other obvious downsides to this approach? Does anyone know of any libraries that offers a similar API, or have experience using a similar approach?I've found this really hard. At my last place we just ate the cost (and ugliness) of included HTML in this strings and dangerously inserting them into the page.
We're also working on creating richer and more streamlined authoring experience in Pontoon, Mozilla's translation management system. You can read about the current state of Fluent support in Pontoon in my colleague's post at https://blog.mozilla.org/l10n/2019/04/11/implementing-fluent....
Though if the i18n string files are present at build time, then this sanitization step could be done there.
Then you could write an alternative set of functions that returns say Vue components or raw html templates.
It still doesn't make individual translations _always_ reusable across paradigms, but I'm not so sure if the impedance mismatch associated with working with raw strings in a modern UI frameworks is worth the translation portability of that de-facto approach. And at the end of the day the only transaction functions you'd have to duplicate are the ones return more than raw strings, so it's not the end of the world.
We've taken a layered approach to designing Fluent: what we're announcing today is the 1.0 of the syntax and file format specification. The implementations are still maturing towards 1.0 quality, but let me quickly describe what our current thinking is.
For JavaScript, we're working on low-level library which implements a parser of Fluent files, and offers an agnostic API for formatting translations. On top of it we hope to see an ecosystem of glue-code libraries, or bindings, each satisfying the needs of a different use-case or framework.
I've been working on one such binding library called fluent-react. It's still in its 0.x days, but it's already used in a number of Mozilla projects (e.g. in Firefox DevTools). In fluent-react translations can contain limited markup. During rendering, the markup is sanitized and then matched against props defined by the developer in the source code, in a way that overlays the translation onto the source's structure. Hence, this feature is called Overlays. See https://github.com/projectfluent/fluent.js/wiki/React-Overla....
Here's how you could re-implement your example using fluent-react. Note that the <a>'s href is only defined in the prop to the Localized component.
<Localized
id="confirm"
$clickCount={7}
a={<a href="..."></a>}
>
{"Please <a>click here {$clickCount ->
[one] 1 time
*[other] {$clickCount} times
}</a> to confirm."}
</Localized>
I'd love to get more feedback on ideas in fluent-react. Please feel free to reach out if you have more questions!I feel there's a fundamental impedance mismatch here because we're defining messages as strings but the rest of our UI as React components. I described here a potentially different component-oriented approach as an attempt to get rid of this impedance mismatch: https://news.ycombinator.com/item?id=19681129
I'd love to hear some thoughts on that approach from folks with more real-world experience working with i18n than I do (which is not a whole lot sadly, given the nature of the kinds of projects I've worked on in the past).
Another advantage is that the components compile to vanilla JavaScript, so we don't rely on a runtime library to run the application.
How does this work for languages that have more complex pluralization rules?
E.g. in Russian it's "1 раз", "2 раза", "11 раз", "12 раз", "22 раза" and "55 раз" - the case depends on the number ending, with exceptions for 11, 12, 13 and 14.
https://en.wikipedia.org/wiki/Arabic_grammar#Cardinal_numera...
Fluent relies on Unicode Plural Rules [0] which allow us to handle all (as far as Unicode knows) pluralization rules for cardinal and ordinal (and range) categories :)
Authoring tools can help here, too. Pontoon, Mozilla's translation management system, pre-populates plural variants based on the number of plural categories defined in Unicode's CLDR.
Another example is when you need to inject an image as a text decoration that includes text, or only makes sense in a particular part of the sentence.
One workaround I've considered is to add the decoration directly to the font you're using, so you can literally translate the decoration as text, but that doesn't usually feel like a reasonable solution, and still might not solve the problem for all languages.
For the sentence ordering we include all variables in the translations, but split translations on styles. We then give the translator the text in order of html appearance in source code for context. Translators can then rearrange everything but the variable across the string and also leave stuff blank when necessary. It's not perfect, but works in most cases.
And in the end developers and designers must consider the "translateability" of the UI. It's always possible to create untranslateable UI.
Our opinion is similar to Unicode's - Gettext is fundamentally flawed design for internationalization purposes.
Here's you can find more detailed explanation of our position - https://github.com/projectfluent/fluent/wiki/Fluent-vs-gette...
Please, don't take it as a criticism of using it. We just don't think it scales and we don't think it's possible to produce high quality sophisticated multilingual UI's with it, but if it works for you, don't touch it :)
Did the Unicode consortium express critics about gettext? Could you provide some reference about this?
I can also point out to ICU MessageFormat - which has been designed much after Gettext and, I'd dare to say on purpose, bares no resemblance to it.
> Secondly, it makes it impossible to introduce multiple messages with the same source string which should be translated differently.
This is false. The gettext message format uses msgctxt to deal with this. It's a fundamental part of the format. The unique identifier is the combination of msgctxt and the singular string. I wonder how you could miss that? We actually use an automatically generated msgctxt for some part of our app to avoid accidentally translating the same source text incorrectly in different context.
Also I couldn't quite follow the point about interpolation of fluent vs gettext (probably because I don't know fluent). Message interpolation in gettext works and can be absolutely readable. E.g. "You have {count} items". The big drawback is that you can't move this variable across strings. Can you do that with fluent?
Thank you for the feedback! I updated the article to include the mention about `msgctxt`.
Personally, in my experience, many project environments end up with partial support for this feature (for example many react/angular extractors don't support it) which leads to limited use and requires the localizer to request adding a context by the developer.
I did not include that since it's just my personal experience and I assume more mature projects tend to recognize the feature and use it, hopefully, extensively :)
> Message interpolation in gettext works and can be absolutely readable. E.g. "You have {count} items".
As far as I understand this is not part of the system (gettext), but its bindings and in result is underspecified and differs between implementations. For example [0] uses `%{ count }` while [1] uses `{{ count }}`. If I'm mistaken here, please, point me to the spec :)
Since it is a higher level replacement, this approach likely suffers from multiple limitations. First of all, I highly doubt that there is any BiDi isolation between interpolated arguments and the string leading to a common bug when RTL text (say, arabic) contains an LTR variable (say, a latin based name of a person). Fluent resolves it by wrapping all interpolated placeables in BiDi isolation marks.
Secondly, I must assume that any internationalization, such as number formatting, date formatting, etc. is also not done from within of the resolver in gettext. That, in turn, means that it may be tricky to verify that a number is formatted using eastern arabic numerals when used in arabic translation, while formatted to western arabic when used in english translation. Fluent formats all placeables using Unicode backed intl formatters (for example in JS we use ECMA402), allowing for consistency and high quality translations where placeables get formatted together with the message.
For example, in your example, will the `You have { count } items` be translated to `لديك 5 عناصر` or `لديك ٥ عناصر`? And what will happen if instead of `count`, you'd have `name: "John"`? Will it be RTL or LTR?
[0] https://hexdocs.pm/gettext/Gettext.html#content [1] https://angular-gettext.rocketeer.be/dev-guide/api/angular-g...
clickHere: { one: 'Please' two: { singular: ' click here {0} time ', plural: ' click here {0} times ' }, three: 'to continue', link: 'google.com/en/' }
<span> {{ this.$i18n('clickHere.one') }} <a href="{{ this.$i18n('clickHere.link') }}" >{{ this.count > 1 ? this.$i18n('clickHere.two.plural', this.count) }} : this.$i18n('clickHere.two.singular', this.count) }}<\a> {{ this.$i18n('clickHere.three') }} <\span>
Don't forget to localize your links! English might not need it but many languages eventually will point to a different url.
Screen reader users will often navigate your page by cycling through the links that are on the page and then they'll get only the link-text read out, not the surrounding text.
It’s a pity that the programming world is still super bad at i18n :-(
And the problems go quite deep, all the way down to standard libraries and programming languages. The Swift debate about correct and performant Unicode string processing was very interesting – people are very hard to convince that string is not a collection of characters randomly accessible with integer indexes.
If your clients are happy to only ever have a product in one language which will never have to deal with anything outside the 7bit ascii range then ignoring the complexity required to do it is (probably) fine.
As soon as you hit some requirement which violates the above if you haven't considered how this might affect you you're likely in for some horrible problems when suddenly you need to handle these things.
The complexity of this Fluent library shows what a massive problem it is. I'm not surprised that we continue to be this bad at it.
Especially the advantages and drawbacks of using the source string as a message identifier, compared to a developer provided ID.
I'm wondering if fluent has something similar to xgettext, to extract the IDs from the source code?
Edit: Looks like there is some discussion about extraction here: https://github.com/projectfluent/fluent.js/wiki/React-Bindin...
And using the source string as ID is a pretty clever trick. Of course, there are some downsides, but there are certainly also downsides with separate IDs.
Having said that, Fluent looks interesting.
I have, in fact, been using Gettext for quite a few years, but of course, as you pointed out, I am also biased.
If you have suggestions on how to improve the article to better represent the reality, please, file an issue and provide a PR! Our goal is to express our design differences, but we don't want to mislead anyone!
It's rather disingenuous to say this without mentioning that in gettext, the actual identifier is a combination of the source string and the string context, which is blank by default. But the context can and should be provided in cases where disambiguation is required.
https://www.gnu.org/software/gettext/manual/html_node/Contex...
This, in turn, breaks the design principle #1 of Fluent - https://github.com/projectfluent/fluent/wiki/Design-Principl...
I'll update the Wiki to reflect that!
With Mozilla's experience and adherence to ideals of interoperability and openness, I can see Fluent as a solid "golden standard" solution for a great chunk of i18n needs :)
https://metacpan.org/pod/distribution/Locale-Maketext/lib/Lo...
We Czechs have it easy!
There was a presentation years ago, how Perl handled Unicode right and every other programming language didn't (with Python 3 pretty close, IIRC).
Does anyone remember the URL?
My point is that while there might be a concept of "few" that does map uniquely to that range, I am not sure naming the keyword "few" is the right name for this.
Quite honestly it would be easier to understand if it were explicitly referring to the range. After all these strings are provided specific to a language anyway. As such why not encode the rules in them explicitly instead of relying on keywords?
Or are those merely user defined abstractions to accomplish reuse? I guess it would help, but I'm still not sure why this needs a whole new framework.
So for example "1002" is "few" in this sense (even though it's bigger than 100) but "14" is not, even though it's less than 100 and ends in 4.
It's a great abstraction that makes a lot of Fluent easier :)
As far as the cases used go, 1024 would take the genitive singular in Russian. 1025 would take the genitive plural. 1021 would take the nominative singular. Nominative plural is not used at all when counting things in Russian.
tabs-close-warning-multiple = {$count ->
[2..4] Chystáte se zavřít {$count} panely. Opravdu chcete pokračovat?
*[] Chystáte se zavřít {$count} panelů. Opravdu chcete pokračovat?
}
Specify a range (2..4). The second option shouldn't need to match as it is the default value anyway (signified by the *) n % 10 = 1 and
n % 100 != 11 or
v = 2 and
f % 10 = 1 and
f % 100 != 11 or
v != 2 and
f % 10 = 1
…where n is the absolute value of the number, f is the visible fractional digits with trailing zeros, and v is _the number_ of visible fraction digits with trailing zeros. Some rules can get even more complex than that; see [0] and [1].It's safer and more robust to rely on the plural categories defined by the Unicode: (zero | one | two | few | many | other), and by the APIs provided by the platform (ICU, Intl.PluralRules, etc.).
[0] http://www.unicode.org/cldr/charts/latest/supplemental/langu...
[1] http://unicode.org/reports/tr35/tr35-numbers.html#Operands
That's an interesting suggestion! Right now, the identifier between the brackets is required, but we could relax this in the future. In 1.0, we erred on the side of more conservative and explicit design, to improve the readability and discoverability of the syntax for translators.
Great work by Mozilla; it's clear there's a lot of experience in the organisation feeding into the design of this system, and it's great that they're sharing it with the world.
I'm German and I disabled spell checking almost everywhere, because most implementations are extremely poor in German. Word lists are a poor solution to capture different word forms, and I find it surprising that even in 2019, only very few programs get that right (for instance Microsoft Word; it understands some but not all grammar rules). This is another thing where I think a modern (OSS) spell checker could make a difference.
($count -> Jana added {n} {apples|apple}) ($gender -> to {his|her} profile)
the $gender will affect what form the word "added" should take. You're suddenly dealing with possibilities($count)*possibilities($gender) variants of the sentence- Gender agreement is not trivial. In French, in "Mary bought it" the verb needs to match the gender and number of the object not the subject: "Marie l'a acheté" vs "Marie l'a achetée" vs "Marie les a achetés" depending on the gender/number of the "it" object. But in most other cases the verb needs to match the subject in gender and number, in Polish "Maria kupila" vs "Stas kupil" vs "Oni kupili".
- In many languages nouns need to agree in case, gender, and number with the phrase they're in, even in English we see this with pronouns: "this is he" vs "this is his".
- And not to mention number agreement between pronoun and either subject or object, depending on context: "this is the button" vs "these are the buttons" - but also "this hovers over the button" vs "these hover over the button" etc. Pronouns in general are a world of hurt, as are copulae (is/are/etc)
- So once we want more complex sentences, simple word tagging like $gender becomes insufficient, because now there's multi-party agreement to worry about, we have to worry about $gender_subject and $object_gender_number_case, etc.
This becomes completely untenable for all but the most technical translators. Maybe those are easy to find for a world famous project like Firefox. Unfortunately, not so for a run of the mill commercial project.
($count -> Jana ($gender -> {addedHis|addedHer}) {n} {apples|apple}) ($gender -> to {his|her} profile)
(obviously with addedHis and AddedHer substituted for the correct word.)https://projectfluent.org/play/?id=2d7ab4b7ed1c4d9656475614f...
It's a complex piece of UI and consequently, the resulting Fluent message is also quite complex. But possible to build :)
https://github.com/projectfluent/fluent.js/blob/master/fluen...
For example in Polish:
"Page 3 of 4" is "Strona 3 z 4"
"Page 3 of 100" is "Strona 3 ze 100"Context: Polish preposition "z(e)" is spelled "z" if the following word starts with s and z and alikes (problematically enough, the exact rule is not systematic) and "ze" otherwise. Korean has a similar case with postpositions "은(는)" and "이(가)" where the former is for words ending with a consonant and the latter is for words ending with a vowel, and the ko-KR localization of Firefox seems to completely ignore and/or sidestep this; the last letter is assumed (e.g. "{-brand-name}는") or a static word is inserted (e.g. "{$user} 사용자는" instead of "{$user}은(는)").
So I think it's only words starting with "s", not "sz" nor "si". And only 0 starts with "z" so it's not a problem (you never have "X out of 0"). So the only special case is for 100-199, 100 000-199 999, etc.
Ignoring the issue altogether :) Or, if you're pedantic - implementing the special cases in the source code. But that's unmaintenable if you have lots of languages.
The Fluent Syntax is a simple declarative DSL. By design, it doesn't allow translators to build complex conditionals or use arithmetic. There is, however, an escape hatch. The problem you described can be solved in Fluent with a little bit of one-time help from the developer of the source code, through a feature of Fluent called custom functions.
Translations in Fluent can use functions to format values or decide between variants. There exist built-in functions like NUMBER and DATETIME. They are rarely used because the Fluent runtime calls them on numeric and temporal values implicitly, but they can be helpful when localizers wish to use custom formatting options.
weekday-today = Today is {DATETIME($today, weekday: "long")}.
See https://projectfluent.org/play/?id=a3540d4f02c104a634adbfc0e... for a live example of DATETIME.There can also be custom functions, defined during the initialization of the runtime. In Firefox, we use one such function called PLATFORM: https://searchfox.org/mozilla-central/rev/d33d470140ce3f9426.... It can be used as follows:
open-preferences = {PLATFORM() ->
[windows] Open Options
*[other] Open Preferences
}
The logic of custom functions is entirely up to developers and the localization needs of the UI. In https://github.com/projectfluent/fluent/issues/228#issuecomm..., for instance, I suggested using a custom function to handle negative and positive floor numbers.A custom function can also cater to the use-case you described. A simple and possibly naive implementation in JavaScript could look like the following one:
function NUMBER_HEAD(num) {
while (num > 999) {
num /= 1e3;
}
let first = num.toString()[0];
return num < 10 ? first
: num < 100 ? first + "x"
: first + "xx";
}
I wrote this with Polish in mind, but it could be useful to other languages in which numerals are named after the first thousand-triple, in a left to right order. Depending on the exact product requirements, the function could be called NUMBER_HEAD_POLISH, or perhaps NUMBER_HEAD_TRIPLE_FIRST_DIGIT :)Once defined, the function can be used as follows:
# The Polish copy can take advantage of the custom function.
page-of = {NUMBER_HEAD($pageTotal) ->
[1xx] Strona {$pageCurrent} ze {$pageTotal}
*[other] Strona {$pageCurrent} z {$pageTotal}
}
This method still requires some work from developers, but it only needs to happen once and in a single palce in code: where the Fluent runtime is initialized. Because they are code, custom functions can be reviewed and tested just as any other code in the code base, to help ensure that they do what they claim to :)Importantly, the use of the custom function is completely opt-in.
# The English copy doesn't need any special handling.
page-of = Page {$pageCurrent} of {$pageTotal}
All localization callsites remain unchanged, and all existing translations remain functional.Right now I'm using qt translation system, it handles nicely various plural forms, just like Fluent.
But it requires special code in each message that has "X of Y" to handle the "ze 100" correctly in Polish. It might be done for Polish, because we have Polish developers, but many languages have similar quirks and it's not done for them. And it would result in combinatorial explosion of translation message versions if source code had to add special case for each quirk in each language.
This seems to be a much better solution.
> And it would result in combinatorial explosion of translation message versions if source code had to add special case for each quirk in each language.
This is the exact problem we designed Fluent to solve. If you get a chance to try it out, feel free to reach out to me if you questions. I'll be more than happy to help and to hear feedback.
The magic is a tiny expression language which understands plural cardinalities, ordinals, etc. so a translator can encode all required logic in a JSON file - the application code can be "dumb".
Here's our take on the differences between MessageFormat and Fluent - https://github.com/projectfluent/fluent/wiki/Fluent-and-ICU-...
[0]: https://github.com/projectfluent/play/blob/a4f49a4a7eeb93535...
```properties
# A comment
hello = Hello, world!
```
This is in fact by design. Properties files are quite nice for simple things. Fluent builds on top of them, and provides modern features like multiline text blocks (as long as it's indented, it's considered text continuation) and the micro-syntax for expressions: {$var}, {$var -> ...} etc.So far, I haven't had much time to invest in building proper highlighting modes for popular editors. There are ACE and Vim modes mentioned in other comments here, and also a slightly outdated https://atom.io/packages/language-ftl20n written by a contributor. I'd love to see more such contributions, and I'll be more than happy to help by reviewing code!
Most notable being that using the source language string as identifier a) discourages changes (and improvements) to the source language strings and b) makes it hard to handle strings that appear the same in the source language but need different translations.