CSS for Internationalisation
chenhuijing.com
chenhuijing.com
[0] MDN reference: https://developer.mozilla.org/en-US/docs/Web/CSS/ruby-positi...
[1] W3C in-depth article: https://w3c.github.io/i18n-drafts/articles/ruby/styling.en.h...
[2] https://www.w3.org/International/articles/ruby/markup.en
> Have you ever wondered how Chrome knows to ask you if you’d like a web page’s content to be translated? No? Okay, maybe it’s just me then. But it’s because of the lang attribute on the <html> element.
The lang attribute is a signal to whatever's inside Chrome that detects language. Here's a simple counter example: https://twitter.com/chewxy/status/1253076770066010112
The image is an SVG generated by graphviz. With no lang attribute (or even a HTML tag) Still Chrome thinks it's Luxembourgish.
Nonetheless, I appreciate the tip on the pseudoclass for lang.
> Chrome first checks the HTML lang attribute and if it's not present it checks the Content-Language HTTP header. Then it gets a prediction from cld3.
Which, even if they don't still do this, I think makes total sense and is what I would do because obviously programmers make mistakes and if your page claims to be English but your language has a high chance of being Armenian and a low chance of being English I would consider it was Armenian.
> Google uses the visible content of your page to determine its language. We don’t use any code-level language information such as lang attributes, or the URL.
But what about the Chrome browser?
[0] https://support.google.com/webmasters/answer/182192?hl=en
> Chrome uses CLD3 to perform language detection on all webpages. This language detection is generally very accurate. Chrome will override the language detection if there is a language attribute on the HTML tag or a language specified in the HTTP content-language header.
> However, both these signals are often incorrectly specified by the site/page, and in particular many non-English sites report the language as English (presumably based on default values from authoring tools, etc.) Currently a whitelist is used to determine if the detected language should always override the language attribute / content-language.
Compact Language Detector v3 (CLD3) is a neural network model for language identification [1]. Sometimes it fails [2].
[0] https://bugs.chromium.org/p/chromium/issues/detail?id=771861
[1] https://github.com/google/cld3
[2] https://bugs.chromium.org/p/chromium/issues/detail?id=105923...
Nice easter egg, first time seeing emoji used in the url fragment. It also loads a new one.
Edit: As an idea to share, I wonder given the topic, whether different emoji's can be localized too. An emoji can mean something different depending on country.
border-top-color: tomato;
border-right-color: limegreen;
border-bottom-color: dodgerblue;
border-left-color: gold;
Looking back, my comment was more due to my own inability to visualize the simplicity of the logical properties as presented vs. what I've seen in other demonstrations and realizing that I lost sight of a resource that I think would be a helpful reference for myself as this new method becomes commonplace. .i18n-hello:lang(:en) {
content: "Hello";
}
.i18n-hello:lang(:nl) {
content: "Hallo";
}
<span class="i18n-hello">hello label</span>
That doesn't work, so you'd have to use ::after, but that requires some more verbose styling.How difficult is it to automatically detect the language? Probably naive but how far can you get by counting how many common words on the page are from a particular language (e.g. "the of then because I" for English and "je pour des les" for French)?
I suspect it would be reasonably simple to add a "Translate this" button somewhere in Chrome (perhaps it's already there, I don't have Chrome installed on this machine.)
Perhaps automatically doing it for every site would be a little bit intrusive on the privacy front.
Yep, I'm asking how hard it is to build a quick and simple auto-detect feature. Is it harder than it looks?
build requires Chromium, not so hard
It will use the property "aspect-ratio" with values like "16/9" or "1/1".
MDN: https://developer.mozilla.org/en-US/docs/Web/CSS/aspect-rati...
Article about it by Rachel Andrew in Smashing Magazine: https://www.smashingmagazine.com/2019/03/aspect-ratio-unit-c...
Recent Chrome and Firefox use that to determine aspect ratio for jank-free loading.
If you want an element to have the aspect ratio 16:9 give it a parent div with width: 100%; height: 0; padding-bottom: 56.25%; position: relative;
then just absolutely position the child element with width 100%; height:100%;
This works everywhere, even in IE8. I am happy I will not have to resort to this anymore in the future, but it will do for now.
In the olden times, browser was not smart (ie: still set the wrong charset). You may have come across for example, an asian site (japanese, korean, chinese), and see a lot of text with "?"'s and "□"'s like so ??? □□□ in 90's and common still through beginning 2000's. Glad these days you don't have to spend time trying to match it up, but it was available readily in browser's menu for you to match.
I think here, the lang tag is useful so you can explicitly tell the browser. And if you are designing the page, you have more control for localization to target by the international code.
You do sometimes see mojibake in web pages, (question marks and □□ in place of the real text). These are caused by incorrect encodings. The web server tells the browser which encoding to use using the HTTP header Content-Type or using the <meta charset="UTF-8"> HTML element. You should always set the encoding rather relying on the browser guessing.
Here I was just sharing my encounters of the browser rendering ?? and □□ character marks, because it did not know. The browser had a "Auto Detect" charset mode, so one always had to toggle (and remember to revert back again when viewing another page). For the very reason you have indicated, "should always set", but in those days (and maybe even today), not always set.
Again, just a memory sharing.