> Chrome uses CLD3 to perform language detection on all webpages. This language detection is generally very accurate. Chrome will override the language detection if there is a language attribute on the HTML tag or a language specified in the HTTP content-language header.
> However, both these signals are often incorrectly specified by the site/page, and in particular many non-English sites report the language as English (presumably based on default values from authoring tools, etc.) Currently a whitelist is used to determine if the detected language should always override the language attribute / content-language.
Compact Language Detector v3 (CLD3) is a neural network model for language identification [1]. Sometimes it fails [2].
[0] https://bugs.chromium.org/p/chromium/issues/detail?id=771861
[1] https://github.com/google/cld3
[2] https://bugs.chromium.org/p/chromium/issues/detail?id=105923...
Which, even if they don't still do this, I think makes total sense and is what I would do because obviously programmers make mistakes and if your page claims to be English but your language has a high chance of being Armenian and a low chance of being English I would consider it was Armenian.
> Google uses the visible content of your page to determine its language. We don’t use any code-level language information such as lang attributes, or the URL.
But what about the Chrome browser?
[0] https://support.google.com/webmasters/answer/182192?hl=en