Localisation is too hard for Gmail
shkspr.mobi
shkspr.mobi
If you called that the "bin" label in England, then functionality could break when the user switches languages... Now they have a bin and a trash, with different things in!
Or you could write the code to relabel all the users emails when they changed language... But what if they had already manually created a label called "bin" then the auto-re-labeler will cause loss of data...
There is no perfect solution here... Any solution has unexpected behaviour for some things a user might want to do. Considering that, I think what is there is one of the least bad solutions...
In the case of Gmail, it would mean that you could search both for in:Bin and in:Trash. The only issue would be if someone creates an additional Bin label in a US locale, then in UK, when searching for that label it would also show items in the trash but that's a better trade of I think.
The vocabulary doesn't matter when the grammar is an incomprehensible pidgin.
Apple script is the quintessential Jobsian "form, not function"
So it might be a disgusting hack, but it's been whatever it is since launch, although incompletely.
This isn't a new concept in computer science.
Probably an oversight or a low priority bug for the Gmail team.
Could also be something complicated in the underlying infrastructure that makes it difficult but I see no reason why it couldn't be abstracted and managed in the frontend as part of localization. Its always easier to talk about things and handwave without understanding problems that may cause with other underlying architecture.
https://erikbern.com/2020/03/10/never-attribute-to-stupidity...
When people have pride in their work and want to put out a quality product, instead of "optimizing" for pennies, then things get fixed.
The user now changes their interface language to England, and suddenly they have two "Bin" labels that map to different sets of things. They behave inconsistently and frustrate anyone who gets into that state. Some messages might even be in "Bin" and "Bin" at the same time! Searching for "in:Bin" returns just one set, or the other, or both.
Doesn't really sound like a good solution...
How many people use Gmail in a language other than US English? How many people change their UI language and rename folders such that there is a clash?
Optimise for where it will have the most impact, and develop a work-around for the minority of users where this causes a problem.
You'd be surprised. There are only about 400 million native English speakers [0], and Gmail has 1.5 billion users [1].
If we assumed that all native English speakers use Gmail, they'd represent not even a third of the users. Of course, there are many non-native English speakers, but I'd assume a significant fraction of them would prefer to use Gmail in their language.
I'd argue i18n has a huge impact, but Google really doesn't optimize for it.
[0]: https://en.wikipedia.org/wiki/English_language [1]: https://www.cnbc.com/2019/10/26/gmail-dominates-consumer-ema...
I, personally, have 3 "GMail" accounts: one @gmail.com, one @ my own domain, and one for work. So if I were the average English-speaking user of GMail, having every English speaker be a user would pretty much cover the 1.5 billion.
(I do not, in any way, think I am the average—I merely seek to point out that it's significantly more complicated to calculate than your post suggests.)
Don't assume that only native English speakers (want to) use things in English.
How?
> Alternatively save all the built-in labels (Trash, Bin, etc) as reserved
So now all localizations are reserved words? And ultimately, this may be a reasonable approach from the beginning, but how many people's accounts are you willing to break to institute this change now?
Now, when the user does a search, they need to search for "In:'Bin (built in)'". Except when they delete the label called Bin, now the search terms they must use changes... All their saved searches and bookmarks with search terms in the URL would break too.
Reserving the names in all locales might work, although I could imagine quite a few users being frustrated they can't name something without realising that the word they are trying to use means something in some language they don't even speak.
I think to solve this 100% you need namespaces.
No one is suggesting that there is, but there are definitely better solutions than what GMail is doing.
So if the query is "in:bin" just convert that to "in:trash" and use that for searching, without the user knowing. Then you also have no problem with switching languages.
Not if you designed the system with localization in mind from the beginning.
I agree, localization is hard. I have to do it for the web sites I build. But I build them with this kind of abstraction in mind, and am able to convert the mechanics of a web site to another language in just a week or so working with a translator. Then they do the content, and I add that in, and it's done and good.
I've added up to five languages to some of the sites so far, and the language audits have turned up no problems, and no complaints from the users. (We do actual human testing with actual humans. No bogus "telemetry.")
Admittedly, my sites don't serve nearly as many people as GMail, but they have to be as close to perfect as possible because I'm in healthcare, and conveying important information.
At the same time I'm able to do it with just me, while Google has a trillion dollars to throw at the problem.
The professional translators I work with say that to them, it looks like Google translates its web sites with Google Translate, rather than with a native-speaking human being.
for example, don't (canonically) identify accounts with usernames. that way when "joesmith" becomes "josephbarnes", the user can just change their username and everything keeps working (because it's really all user ID 19829).
any "name" you get from a user should always be considered nothing more than a non-canonical label for a resource with some other persistent and canonical identifier.
Should they all break just because I change the language settings?
"inbox", "trash", "drafts" are labels and localization shouldn't pretend that they are something else.
It is nice if the UI presents us a named shortcut displayed in our chosen language setting, which applies a certain filter on labels when clicked.
Or worse yet, outsource to a low cost firm whose translators add their own semi-illiterate or colloquial flourishes to Google Translator output.
this basically boils down to "because they didn't consider internationalization or localization when coding it."
It is NOT a valid excuse for this not working from what was already a multinational company when GMail was introduced.
> Now they have a bin and a trash, with different things in!
translation: "but there's tech debt"
Not an excuse. Especially from a company with enough money to fix it without even noticing the effect on their bottom line. It didn't have to be coded to have the concept have a 1:1 correspondence to the name. "spam" and "trash" could have been conceptual objects with multiple localized interfaces for humans.
But yes, it's getting complicated. I can see why they just stopped with changing the GUI labels and left the workings in canonical US mode.
That's something that the highly paid engineers at Google should figure out, not something that random Internet people on HN should have to figure out when replying to your comment.
Google has the resources to find an answer to your question. It's embarrassing to all of us that they refuse to do so.
Typical "smart" solution, that completely fails when meeting the real world.
I wonder how many horrible software issues we've created trying to force smart/elegant solutions and then having either to: completely rebuild the feature, create horrible workarounds, or just let users deal with it however they can (which seems it's the Gmail way).
Custom user labels can be arbitrary, non-namespaced free text using any string except “gmail”.
The UI should absolutely be localized and clicking the "filter messages in the bin" button should add "in:trash" into the search bar.
It means all other groups become second class citizens, and you might not even understand who will be your real primary group at the time of coding the service.
In practice, I'm not sure how that's possible. I18N is hard, really, really hard, to the point that only very large teams can really get things even close to right. Really, you almost need as many people working on the issue as there are written languages, or at least some people are going to be second-class users, right?
But let's says we stick with EFIGS, since that covers most of the western world and supporting Arabic and several different Asian character sets is well outside the expertise of the majority of people reading this page. That's still really, really complex.
Maybe most of all for Americans, who can easily live our entire lives without encountering anything other than American culture.
I mean, I've spent a lot of time trying to tease Traditional Chinese subtitles apart from Simplified Chinese subtitles, and then watching how the resulting Chinese subtitles handle large numbers, to know that I'm not the person for the job!
As a counter example, why aren't URLs in arabic too, for example? "https://news.ycombinator.com/login" doesn't make sense to an arabic speaker. Should we also force service providers to alias "/login" to "/تسجيل الدخول" as well?(forgive me, I just googled 'login in arabic' and picked the first result)
You could make the same argument about every programming language in use today. I don't see a mainstream programming language that doesn't use english based reserved words; for better or worse, the legacy of ASCII and english lives on.
What's the API? The search string?
>localization of en_UK tries to share a filter with a user with en_US and suddenly it doesn't work is not an acceptable outcome IMO.
Why not? Does anyone actually want to share filters? Is it even possible to share filters?
>The UI should absolutely be localized and clicking the "filter messages in the bin" button should add "in:trash" into the search bar.
What if you have a 'trash' folder? If anything, they should just use an internal identifier with no semantics attached.
Why wouldn’t you as a user expect that the search strings be portable across accounts? It would be surprising, especially if the changes across locales was small such as the one in this article. I have copied and pasted filters from blogs on the web for example to get some more complex queries. It’s not like you get a lot of feedback from gmail if you get it “wrong”.
Your comments on naming are on target and that’s why we end up with ugly guids across Windows, for example. I mean who wouldn’t know that your trash folder is really {4f342ebb-d392-4a7d-8db8-3f718c1bcd71}. But that would make for an entirely unusable experience for 100% of users. I would assume that, just like a programming language, “trash” is a reserved word and you can’t use it for the name of your own tags. I have not tried it though.
This is ultimately a subjective decision to make. I’ve done this hundreds of times when designing APIs. Balance usability across multiple user bases and maintain compatibility with decisions that I’m sure were made decades ago when gmail was a beta project in the early 2000s. It’s a problem with no “right” answer, so I think having a guiding principle - such as prioritizing consistency - can make those decisions easier and result in a more sane api.
I admin a few g suite users. A common question I get is "I can't find that email" Which is easily resolved by sending them a search query.
It'd be weird if search worked differently based on localisation.
As a non American english speaking British person, the americanisation of our language continues to irk me. It's a constant situation of scrolling down localisation dropdowns triyng to find out if my locale or language, is 'English (UK)' or 'British' or 'Great Britain' or 'United Kingdom'[1], or simply not even there with just 'English' listed (which of course is US English). Then to add more salt to the wound, I'll find a dropdown that is in Alphabetical order, except for 'United States' which for some reason is pinned to the top.
Because we Brits share a language with the US, and because most software is written by americans, or by people who learnt american english as a second language, our language is being steadily eroded. If it were any other language being abused, there would be campaigns to stop it happening.
---
[1] No, Great Britain, and the United Kingdom are not the same, and are therefore not interchangable.
My first encounter with this came from "How to be an Alien" by George Mikes:
> "You must understand that when people say ‘England’, they sometimes mean ‘Great Britain’, sometimes ‘the United Kingdom’, sometimes the ‘British Isles’ – but never just England."
:P
But let's talk about "bin". That is the second word of "rubbish bin" I assume. So why adopt the general, non rubbish word as your standard trash container name? Because of that odd choice, now you have to explicitly state what kind of non-rubbush bin a bin is else it is confusedly assumed to be the trash.
When my parents grew up, people in cities still burned their trash in their backyards, or in special cans on their balconies. The ashes were put in an "ash can" and put out for the "ash man" to collect.
When burning became illegal, people started throwing garbage in their ash cans, and the word transformed into "trash can."
At least that's the story as related to me by my parents and grandparents. The ones who are still alive still refer to it as an "ash can" for this reason.
I disagree and vehemently. The strength of the English language is its ability to change and absorb words from other languages. Your proposal to have campaigns restricting change leads to the decay of a language and to oddities like "courriel" instead of email in Quebec French which - mind you - tries to keep itself free of English borrowings. See also the needlessly long compound words in languages like Dutch and German. Even Finnish has "speaking box" instead of telephone.
This is a road you won't like. I think you should accept that English is no longer "your" language. It belongs to everyone now.
German for telephone: Fernsprecher, or "far speaking"
So how is that better or worse than "speaking box".
Finnish is a weird outlier that is similar to Hungarian but different to most other Germanic/Latin/Greek derived European languages.
But with regard to Finnish, while Finnish is a Uralic language and not an Indo-European one, the Finnish language today is replete with calques on Swedish and German, because those were the prestige languages in the region when the modern Finnish literary standard was being created and new words coined.
Also, many of the "needlessly long" compound words in Dutch and German are just the same as in English and French, except those languages use equally long Greek or Latin compound words instead.
Great Britain is the island that contains England, Scotland and Wales.
A citizen of the United Kingdom is British, even if they live in Northern Ireland.
What you're complaining about is basically about spelling, not usage or "language".
Many of the so-called "misspellings" in en_US vs en_GB aren't if you look at the OED, for example, the -ise and -ize. Having "u"s in spelling (eg colour) is a left over of the French influence on English in the UK.
A lot of the changes in spelling are because Noah Webster, when creating the first US dictionary, simplified things, eg axe vs ax. Some of the word usage, eg "gotten" came over to the US but faded out in the UK.
Oh, and it's spelled "americanization" :)
Keep speaking and writing proper English at home, of course!
Because we Brits share a language with the US, and because most software is written by americans, or by people who learnt american english as a second language, our language is being steadily eroded.
I'm not sure there's any solution to that, unfortunately... Possibly building new software from scratch, and putting in place a strict style policy that code and comments must be in British English?
While I appreciate the importance of formality and precision there's way less time wasted (most times) when people don't get hung up in pedantic ortographic or even (relavitively) common signifiers.
I do insist in britishisms when I write to American colleagues but obviously have to fallback to en-US when coding. The cost is: I have struggled a bit with grey/gray in stylesheets in the past.
It goes both ways. Especially on the internet, a lot of Americans pick up wonky spellings and phrases from your end of the ocean.
It's just a fashion that comes and goes. For example, back in the late 17th/early 18th century, England went through a wave of Francophilia which altered the spellings of a number of common words. By then, the United States was becoming more independent and because of geography wasn't part of that craze, so some of the legacy spellings remained here.
In the early 20th century, American newspaper publishers pushed for simpler spellings of some words ("thru"), in order to save money on newspaper and ink. They became common in the United States, but less so in other English-speaking countries.
Don't let it irk you. It's not worth it. Humans are messy.
Go read medieval english and see how you like it. I’m sure as it became popular Latin speakers despised it.
- Probably only a tiny fraction of users actually know about the "in:folder" search syntax.
- A simpler UX pattern for users who don't use a keyboard for everything is to click on the folder first and then add to the search string to limit search to that folder.
- Folder names must be normalized in some way before being added to the search string, otherwise it would be impossible to search for folders with for example spaces in the name. This normalization could happen in many ways, but if it reduces the character set in any way (such as replacing spaces with hyphens) it's trivial to create "duplicate" folders for search purposes ("foo bar" and "foo-bar", for example). This problem gets worse if they introduce aliases into the mix.
- What about the "in" part of "in:bin"? Surely localizing only part of the API is more confusing than localizing all of it? I've never seen an API framework which encourages this though, probably because it's a shedload of extra complexity for very little gain.
still a few phrases came out awkward
This is a pretty common localization problem. Tooling doesn't always make it easy to share context with translators who are often contractors for outside agencies. Unfamiliar developers try to do string math instead of laying out full sentances to translate. People are surprised to find the same original sentence needs different translations on different screens.
Of course, Google has lots of experience with localization, so should really have this down by now.
The issue is a text-based search where the keywords and commands are not localized.
It's what Excel has tried as well, renaming functions to localized versions; I cannot imagine how big of a headache that must have been / be for its developers.
But can't help but think its made harder because there's this sense (among those managing resources) that its a nice to have, not a crucial aspect of product development. Which is funny because software, specially web software aims at a global market (at least implicitly).
In my experience localization is even less prioritized than accessibility (non locale related accessibility). Maybe because there's more awareness around a11y. Likely because unless you're actively working in a specific market (with a specific locale) you don't care about making it readable for that market, and yet you gladly take credit card details from that market.
On the other hannd, and while I'm an ignorant in these matters, I'm really surprised that AI/ML is not helping us more on localization and yet, we can make it make up new'ish text (in en-US), replace people's faces instantly in photos and recognize faces in google photos.
Like Excel renaming all its formula names which is a huge problem with cross-language excel sheets. There is a reason we don't translate programming languages and this is the same thing.
I used computers before localisation was a thing (we didn't even have localised keyboards) so I always have everything set to US English (even though I endeavour to speak British English).
However I agree, if you localise it should all work properly.
In that context, I think it is wrong to expect people to know multiple languages in order to interact with it.
In the Anglo-Saxon world, we write 3,000.00.
In many continental European countries and countries like Brazil, the same number is written 3.000,00.
I once generated a CSV in my program which could not be read correctly in Excel in Brazil. I later figured out that I had to use another delimiter than a comma, because the comma was the decimal separator.
I always have to do "File -> Open" instead of just doubleclicking because then I get everything in one column. Indeed if I do file -> Open it defaults to semicolon. Grrr.
I never understood why this happened, I thought the developers of Excel were just stupid. Actually now that I understand why, I still think so :P Indeed CSV means Comma Separated Values.
Excel is very dumb in how it deals with this. When it thinks dates should be DD-MM and you paste a bunch of dates in MM-DD format it will parse the 'invalid' combos correctly (like 01-31) but it will leave the others alone. It should just realise "hey this is not the format I expect" and do them all correctly). Because if you don't notice you now have a whole bunch of half correct / half incorrect data and no way to tell which is which. it doesn't even flag the cells with a warning or comment. Or even pop up a warning to the user that something is off.
Also, if I paste a whole load of numbers starting with zeros then YES I want them as text, not for them to be changed to 1.374E+22. Great with stuff like serial numbers or IMEIs.
Microsoft with all their self-proclaimed AI chops should really apply some of that to Office. All these things have worked this way since the 90s.
PS: Part of the issue is also that MS doesn't have an intuitive application for databases, and because nobody groks Access everyone uses Excel as a database which it isn't. Causing a stinking pile un unmaintainable VLOOKUP crap.
While you should be using YYYY-MM-DD. Please at least use a different separator if you don't.
Number formatting in C and many languages following the system locale by default? Better remember to disable that for config files or anything shared between computers.
Localised error messages (especially compiler errors)? Have fun searching for a solution.
Websites (hey Google) assuming that you want things to be in some language based on your location.
Command-line utilities printing using localised date formats which are then parsed elsewhere...
I suspect it is easier if everything is a foreign language to you. But it's really easy to trip up between eb_GB and en_US because of their similarities.
You might enjoy this article on if PHP were British. https://aloneonahill.com/blog/if-php-were-british/
Single items of information are still data, not datum.
Among many things in HTTP/HTML/etc that bug me is "referer".
At least that’s equally wrong for everyone.
what's unfortunate is that http is self-consciously inconsistent about the spelling: https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Re...
Do localize english text but not the digital concepts.
But then, my native language uses "mouse" (the english word as-is) for the computer peripheral, not the word we use for the mouse animal. There were some silly nationalists who wanted to translate everything, but they got laughed at so much that all such proposals died.
We also get a lot of "french language warriors" who call our call centers and complain about ANY mistake they can find. One that pops up frequently is if we use the incorrect punctuation mark (" as opposed to « and » for French), which happens a lot if a naïve marketing manager is copying and pasting text without regard for the format.
As I said, actual text is nice to be translated, and translated right... but not the menus and stuff like that.
I'm guessing the situation is radically different in, say, Latin America. They can probably afford to use english IT terms. We can't, because it won't stop at IT.
I am not dogmatically against using english words if they are commonly used but I am against avoiding native words becouse it sounds abit like corparate speak to me.
I speak English but regularly correspond with people who speak Hebrew. Hebrew is written from right to left, so "john doe" is written "eod nhoj." This is normal. But when Gmail writes Hebrew names in English, it often transposes the letters only and not the first/last fields, so you end up with "Doe John." The problem is that this happens inconsistently, so you never know if it's going to occur.
I can't tell you how many times I've mixed up people's first and last names, which is embarrassing! Because I don't speak Hebrew, I don't have a natural sense for which names are first names vs which names are last names, so sometimes it's impossible for me to decipher which is which.
Very frustrating.
Example from a quick DDG search for the name "Arnon" which was the first thing that came to mind. Both Yigal Arnon and Arnon Segev are names of attorneys, so Arnon is definitely ambiguous. The latter's law firm employs an associate named Chen Moshe, where Moshe is a common biblical first name (Moses) and Chen could be a man or a woman.
For the German example in the article, would you expect “im:Papierkorb” to work?
If I click "Inbox", Gmail takes me to my "primary inbox" where my non-spam emails land. As far as I'm concerned, that's my inbox. It takes another click to "promotions" or "social" to see all the unwanted advertising semi-spam mail that I probably signed up for but don't want to read. Google is really good at this bit.
But to search for unread emails in this inbox, you can't do "in:inbox" because that brings back all the mails from "social" and "promotions", and the interface for search results doesn't have the same 3 tabs, so the relevant emails are buried a hundred pages down. "in:primary" doesn't work either. The magic incantation is "category:primary is:unread", but at various points in the past you've needed instead to use "is" or "in" or "label" instead.
Google supports you searching for messages by content. Type the words or the sender's name and you'll find the email. Advanced search is an afterthought or downright discouraged, in the same way it's been made less powerful on Web searches than it was in 2001.
This label rather than mailbox approach is clearly visible if you ever use the IMAP interface to GMail; it breaks assumptions / guidelines / rules in the IMAP protocol.
On top of that GMail sometimes treats some special labels as exclusive to other special labels (i.e. some of the built in always available labels). i.e. something can't have the inbox and trash labels at the same time.
in is fine, just interpret tags as sets of emails.
This very much limits ediscovery flexibility and is unbelievable for a company that specializes in search
I bet someone decided to the L10N (it was not there at launch) for the label display name (as you observe), and certainly knew they were not addressing the search bar (that would be a much harder change), but figured it's ok since so few people use it.
It sort of makes sense to do the relatively small amount of work needed for the daily use of 99% of your users rather then get stuck on the really hard part for the last 1%.
Localization is hard. This is one of the big reasons why Europe is not able to re-create Silicon Valley. If a startup in Silicon Valley creates a US only version of a service, they have access to a massive, wealthy market without having to do any localization. While the EU as a whole is larger than the US market, it is split into a lot of different localities with different languages.
I once tried to remove the US language from Windows, in a vain attempt to stop it continually switching from British to American in random places.
Turns out it breaks all sorts of things.
I don’t know of such a language, but I’m sure it exists... At least you shouldn’t assume that it doesn’t exist.
You would have to resort to quoting, maybe? in:”Trash can”
Or do you expect users in language A to also know language B in order to work with the interface?
Users should not be expected to know a second language if you offer products in their language. I agree that Arabic and Farsi are somewhat outliers here, but this behaviour impacts languages like German 'Papierkorb'
I would expect searching in:papierkorb to work as expected.
I am familiar with 'im' and 'am', from a translating them back to English (while learning), but lack the ability to know the correct times to use them (as in this case).
Of course, it's not quite grammatical without the determiner, making it "in dem Papierkorb" which contracts to "im Papierkorb", but then you could say that about English as well, still we don't write "in:the bin", so "in:Papierkorb" should be OK.
If you want a challenge, there are languages that prefer postpositions ("Trash:in"?) or case suffixes ("Trashin"?) or even circumfixes ("iTrashn"??) – and some case suffixes may happen to be formally equal to the nominative for certain words (so "Trash" could be ambiguous between "in:Trash" and "Trash"). Happy localising.
Despite the advances in AI, computers still need humans to meet them in the middle in their own language quite a bit.
Here in particular, since “trash” is not just a technically imprecise word but also a highly idiomatic one (to North America in particular).
I wonder if the user switched languages at some point, or if someone was overly lazy with en-GB and missed this.
[1] https://stackoverflow.com/questions/2185391/localized-gmail-...
Faster, easier to navigate, more intuitive, less clutter, and can be set up to not open images by default (if you don't want the sender to know whether or not you've opened their email).
I use it for a few years now, and I have zero complaints.
Just thinking about the command implementation in french, the inbox is called `boîte de réception`, good luck on having every user writing it with all the correct tildes.
Gmail's choice is not because it's too hard to localise commands, it's because it creates a lot more stability with the product.
Another difference with of british people. “Doesn’t” is the correct word here, to me.
Why would you open the article with a pot shot at Americans and American English if the goal of the article is to get an American company to cater to your linguistic needs?