Your code displays Japanese wrong
heistak.github.io
heistak.github.io
So, if I have to display user-entered text (usernames, posts, comments, messages, form data, etc), and I want to do The Right Thing™:
- I cannot rely on user locale, because it might be set to something generic like English, or the user may be bi-lingual.
- I cannot rely on location, because the user may be traveling to a different CJK region, or somewhere else altogether.
- I cannot set a single lang: attribute for the whole page because it'll be wrong for the other two languages.
- The string alone is not sufficient to identify the language because you can write valid sentences in different CJK languages with the same codepoints.
- I cannot have a per-user language setting, because users may be bi-lingual.
What does that leave me? A dropdown list "C/J/K/Other" besides every single text field?
I'm chucking this on my pile of examples of software development being hopelessly broken by design, along with "unix time is non-monotonic and discontinuous at random" (hint: what's the unix time exactly 1e8 seconds, ~3 years, from now? Answer: it's up to the astronomers[1]!).
[1]: https://en.wikipedia.org/wiki/Unix_time#Leap_seconds
Edit: actually, even the dropdown list is insufficient because it only allows one language per string! How is a Japanese user asking for help learning Chinese supposed to write?
Or have I misunderstood?
[1] These readings didn't take account for systematic variations like the initial sound law (두음법칙, for example 이 vs. 리 at the beginning of words).
On top of that, Kanjis aren't the only type of letters Japanese folks use. They also use Hiraganas and Katakanas, which are phonetic symbols and totally unrecognizable to non-Japanese speakers.
As many websites do not convey the language properly they try to fix that using some heuristics, but especially for small texts (like messages) that can easily fail hard leading to completely wrong pronunciation.
Mail for example tried to fix that by allowing you to annotate the language for UTF-8 embedded in mail headers and the content. Ironically while that mechanism works for display names it doesn't work for email addresses itself. And I'm not sure any mail program uses it.
> What does that leave me? A dropdown list "C/J/K/Other" besides every single text field?
If you need to control the language display at the granularity of a single text field (rather than a user or a page or a website), then yes, you need a tool that operates on a single text field. This shouldn't be too surprising.
Surely you can get away with one of your other solutions, though. In particular, you can guess the language and be right most of the time.
Representation matters. There should have been plenty of experts who understood that Han unification was problematic. It seems they were not in a position to do anything about it.
A few years later they increased the character space by 5 bits and it wasn’t an issue anymore, but the original legacy of han unification remains.
> if the equivalent symptom was happening with English text, ιҭ wѳuld bє lѳѳκιng sѳmєҭЋιng lικє ҭЋιs.
From a page linked by the article, "I Can Text You A Pile of Poo, But I Can’t Write My Name":
> To help English readers understand the absurdity of this premise, consider that the Latin alphabet (used by English) and the Cyrillic alphabet (used by Russian) are both derived from Greek. No native English speaker would ever think to try “Greco Unification” and consolidate the English, Russian, German, Swedish, Greek, and other European languages’ alphabets into a single alphabet.
If there had been a proposal to sacrifice English in order to cram Unicode into a certain code space size, is there any question that the people on the panel whose first language was English would have quashed it?
But people who might have spoken out for Japanese, or for Indian languages, etc. were seemingly not in a position to do anything.
Han Unification is just as absurd and just as unacceptable to a native Japanese speaker as Greco Unification would be to a native English speaker. But Han Unification went through because native Japanese speakers were not in a position to block it. Representation matters.
¢σηѕι∂єяιηg тнє нιѕтσяι¢αℓ αηтιραтну вєтωєєη נαραη αη∂ ¢нιηα ιт'ѕ ρяєρσѕтєяσυѕ тнαт тнє נαραηєѕє ωσυℓ∂ υηιℓαтєяαℓℓу ѕα¢яιƒι¢є тнєιя ¢υℓтυяє ιη ѕυ¢н α ωαу.
(EDIT: Considering the historical antipathy between Japan and China it's preposterous that the Japanese would unilaterally sacrifice their culture in such a way.)
Edit: s/variant/variation/.
Either way this seems to be the most painless "fix" for this given the current situation. Setting a language for the whole text fails as soon as the languages are mixed.
It really is not, a screen reader using the wrong language is way, way worse than a human with the wrong pronunciation. The first time voiceover decided to switch to russian in the middle of english text I thought the os had crashed, the mangling is quite extreme.
As far as I know textfields only have one font applied, so entering both languages in a single field won't be optimal. And if you're not doing anything fancy with your fields, they will all take the same font as well.
So even at the input level, the user switching languages will already be mildly screwed, and the best solution would probably be to change pages for each language.
1) The language setting on your website if you have one and have translated it to C/J/K. You may use different TLDs for the different languages and discern that way, too.
2) The list of preferred languages from the browser. This is usually unreliable, but if someone has gone to the trouble of inputting "english=1;japanese=.9;chinese=.8", then it's a fair bet they want Japanese Kanji usually and will be understanding if you use them in place of Chinese Han characters.
3) The country to which the user's IP belongs. The least ideal option, but if you're in Korea and reading a random string of Hanzi, you probably expect them to look like Hanzi.
You will show the wrong characters to some users, but the behaviour is understandable. "Oh, the site is showing me Korean characters because I'm in Korea." is a lot easier to grasp than "The site is showing me Chinese characters because I clicked a dropdown one time that I forgot about and now I have no idea why my name is written wrong!"
You can argue about point 2) that some users might set their language preferences and forget about it, but so far I have never observed a user who doesn't know about them messing with the setting.
They write with japanese kanji. I imagine minuscule differences in kanji forms are negligible compared to general unfamiliarity with the foreign language.
Well, depends on your definition of unix time; if you use time() as the definition then it is actually monotonic because the integral part only repeats on leap seconds?
1. Websites that are in Japanese are likely tagged with lang=ja already. So they will display fine. Unfortunately, this practice seems to be less followed by Chinese sites. I checked a few top sites, qq.com do have lang=zh-cn, while baidu.com and sina.com.cn don't.
2. Majority of UI elements in OS will prioritize the display language you set when choosing variants. This means, if the users are reading content in Japanese while also using Japanese UI, the glyphs would be correct. Of course, this will cause problem if a Japanese is reading Chinese or vice versa, but such scenario is in minority.
Another scenario, which I think is more common, is when someone is using a Latin-language UI. For example, lots of my (Chinese/Japanese) friends are using English UI while reading Chinese/Japanese a lot. The OS in this case will default to one variant (I believe Apple by default would choose Japanese) and therefore display another language's glyphs wrong (side note: for web pages, desktop browsers often have their own font/glyph fallback logic above the OS one).
3. Most of people are just not sensitive to such thing. I pointed it out to lots of people (when due to their setting, some glyphs are displayed wrong, like 门), and they can't care less.
Also, there is no simple "fix" if you have multi-language content. Without manually assign <lang> tag to every single string, you can't display both Japanese and Chinese correct at the same time. It isn't worth the hassle for just a few phrases in text. A good example is Wikipedia, they have templates for all kinds of languages so you can display them correctly even if it's just one Japanese word on, say, English Wikipedia. And Wiki editors do use them all the time!
Han unification have been a huge mistake, to save a few thousands of characters, and now we keep piling on more and more stupid emoji.
> Han unification have been a huge mistake, to save a few thousands
> of characters, and now we keep piling on more and more stupid emoji.
This is my exact problem with Unicode. I've very grateful for the efforts that they have made in the past, but the change from "spare valuable codepoints at the expense of causing ambiguity in text" to "assign a new codepoint to every cartoon permutation of intangible nouns" is infuriating.Unicode may have been never adopted at all if it had even larger set for CJK and made all CJK texts 1.5-2x larger than in Han-unified version, due to longer encoding. Also: UTF-8 did not exist yet and most systems treated text as arrays of fixed length characters.
The original CJK Unified Ideographs block from 1992 consists of 21k codepoints. It was impossible to do that four times since 4×21k is more than 65k, and we're going to need space for some other languages as well. Why not make Unicode larger? Well, size was a real concern back then (still is, to some degree, but less so).
Since then Unicode has expanded and now we have slightly under 1 million codepoints. Han character blocks have extended to about 93k codepoints today, and 4 times ~93k codepoints is actually feasible. But now you run in to compatibility issues: you can't remove all the old Han unification stuff (it will break text, big no-no), so you need to re-define it all anew. Is that better? How about mixing "old" Han unified codepoints with new Japanese or Chinese stuff? Will it really improve things or just cause endless confusion (see: combining characters)?
For scale, all of Unicode currently defines about 145k codepoints; so even with Han unification we're talking about two thirds being taken up by just these three languages.
In comparison there are currently about 3,000 emojis, although the number of codepoints is much less since many codepoints are re-used (e.g. "firefighter" is "person + firetruck", flags use the country code, etc.). In a quick check it looks like there are about 1,000 to 1,500 codepoints reserved for emojis. In comparison, this is nothing.
What I'm trying to say is that the (comparatively) very low number of emojis has absolutely no bearing on this and that going off on a tangent about it is very misplaced.
I have no problems with this either, at least not principally. But historically this was literally impossible. Someone thought of a clever hack that seemed like a good idea at the time, but turns out it doesn't work all that great after all (at least, according to some – opinions seem to differ and I can't really judge myself) and now you're stuck with it and fixing isn't so easy – I don't know if people have made concrete proposals for fixing this, but if it was easy it probably would have been done already. Sometimes sticking with a suboptimal "legacy" solution is better than replacing it with a new better solution due to the friction and issues involved.
It could be stored in-band with the text with little changes to existing systems. The only change would be on the presentation layer, and if the tag were to be a non printable character, it would be backward compatible. An input device could implicitly tag input texts depending on the default lang.
You need some form of sanitization, but you need it for right-to-left and left-to-right already.
Its use is also deprecated and discouraged. According to [1] it's often not needed, and [2] states that it puts a lot of burden on implementations and best done at a higher level such as HTTP, HTML, etc.
I have no opinion on [1] as I don't speak these languages, but I do know I really hate working with these "invisible characters" in Unicode both as a user and developer. Copy an extra invisible LTR thingy or display variant codepoint and stuff can look and behave different, and it may not at all be obvious what the hell is going on (especially for those without a technical background).
[1]: https://www.unicode.org/versions/Unicode14.0.0/ch05.pdf#G115...
[2]: https://www.unicode.org/versions/Unicode14.0.0/ch23.pdf#G301...
"ja-JP" part is also written in tag characters, so it's actually E0001 E006A E0061 E002D E004A E0050 E007F and doesn't render even in unsupported environments.
It does, but they aren’t widely supported and their use is not recommended, see http://unicode.org/faq/languagetagging.html and https://datatracker.ietf.org/doc/html/rfc6082. The recommendation is to use markup languages instead to carry that information.
The reasoning is that anything related to styling is out-of-scope for Unicode (except where needed for round-trip compatibility with other character sets), or else people will also want tags for bold, italic, monospace, or (expressed semantically) for emphasis, code, etc. That’s what markup languages like HTML are for.
My experience with web crawling is that the use of the lang-tag seems inconsistent at best. To make matters worse, sometimes content is straight up mislabeled, although Japanese sites often helpfully declare that they are using the Shift_JIS charset rather than UTF-8, which is at least somewhat helpful in figuring out that it is Japanese.
OSs and Browsers having their own logic for it actually makes things _worse_ in some cases. Windows is especially bad (different types of UI elements care about different settings or don't care at all, so good luck having apps render correctly if you don't change your entire OS locale), and Chrome is pretty bad too, again especially on Windows. Overall MacOS/iOS and Safari does the best job by far.
The failed attempt at Han Unification[1] is the worst decision the Unicode people have ever made.
At first I nodded my head in agreement, but then I decided I still think the failure to include separate code points for "lower case Turkish dotted I" and "upper case Turkish dotless I" is worse.
You can't have 'ı' ≡ lc( uc 'ı' ) unless you already know you are processing Turkish ... completely unnecessary complication.
Consider, İ/i where it is _possible_ to do lossless case conversion:
> lc( 'İ' ) becomes i followed by COMBINING DOT ABOVE which means uc(lc 'İ') becomes LATIN CAPITAL LETTER I WITH DOT ABOVE as a by product of the fact that perl6 deals in graphemes[1]
say 'İ' eq 'İ'.lc.uc.lc.uc;
True
If an extra codepoints existed for Turkish dotted I, such contortions would not be necessary and this would have had no implications for existing working code at the time (nothing says those codepoints must be used, they just give smart software options).Now, there is nothing one can do with I/ı that will make 'ı' eq 'ı'.uc.lc.uc.lc true without extra information. If codepoints existed, then such special casing and carrying around extra information would not have been necessary.
Also note:
# The letter Ö is not considered to be a variant of the letter O,
# and is a separate letter in the Swedish alphabet. The former
# character is, however, the accepted alternative in contexts where
# Ö cannot be used. Earlier practice substituted OE, which is no
# longer recommended but will still be encountered.
#
U+00F6 # LATIN SMALL LETTER O WITH DIAERESIS
It should not have been too hard to say "The letter İ is not considered to be a variant of the letter I" and vice versa for the lower case versions.I think its' because most people don't deal with it in big amounts. I heard a lot more complaints from people using android phones that didn't have jp fonts by default. At the third of fourth page they started to care, and once they noticed it frustration just stacked (it's just a matter of adding fonts, so not a big deal).
Otherwise writings are flexible enough for small variants to not be triggering (I mean, people can already read calligraphy...)
Noto Sans is basically Source Han, one of the best free JCK fonts. Not sure what you mean.
Without Han unification, this wouldn't really be a problem, but Han unification is to a large extent the same philosophy pursued with unification of Latin scripts (and other scripts)
Edit: I mean what is the expected (grammatically correct) way to do it if you were writing with pen on paper.
> Line breaking rules
This should link to W3C Requirements for CJK Text Layout [1]. The Wikipedia article alone doesn't fully describe the complexity of CJK typography.
CJK languages are common in that they all have classes of punctuations that can't be separated by a newline. But there is one more thing to consider for Korean: both word-based breaking and character-based breaking is possible depending on the context. The general rule is to use word-based breaking for larger texts and character-based breaking for smaller texts, but there is no clear threshold so you really want to consult Korean users for testing.
[1] https://www.w3.org/TR/clreq/ (Chinese), https://www.w3.org/TR/jlreq/ (Japanese), https://www.w3.org/TR/klreq/ (Korean)
> Messaging Apps: Do not directly hook to the Enter key to submit messages
This advice is also problematic. In pretty much all Japanese and most Chinese IMEs they should go through candidate windows so pressing Enter should not submit messages, but in some Chinese and virtually all Korean IMEs there is no automatic candidate window and pressing Enter should submit messages.
In the ideal world detecting a newline as suggested by the article should have solved this issue, but that got complicated by clueless pan-CJK IME implementations. They generally assume candidate windows even for Korean, so they do not commit texts on Enter and that's very inconvenient for Korean users. Therefore it is rather recommended to detect a newline by default, but also have an option to submit messages on Enter.
Emoji was added because Apple and Google had to deal with (then-)Japanese emails, and skin tones were not specified. It were implementations that impose certain skin tones (that do not even match the original Japanese emojis) and as a result Unicode had to introduce a mechanism to change skin tones and mandate the default emoji without that mechanism to be neutral.
RTL "modifiers" are actually formatting characters closely tied with the Unicode Bidirectional Algorithm [1]. Until then texts with both RTL and LTR fragments were handled incoherently, for example legacy character sets were still struggling with logical vs. visual order issues. So they are indeed Unicode inventions, but necessary ones that do not alter existing texts.
For CJK characters Unicode now provides ideographic variation selectors that select the exact glyph (or more accurately, a restricted glyphic subset of the base character). They do not disunify characters but they do provide a strong hint to display those characters in a specified way. In this way they do not cause an additional issue to existing Unicode systems (as they should already do normalization and collation in the Unicode way). The disunification by comparison would almost instantly break existing texts.
Adding skin tones was a choice, there was no immediate need for it.
You are incorrect. See the original design document [1] for skin tones and other diversity improvements, especially the "Sources of input" section.
[1] https://www.unicode.org/L2/L2014/14172r-emoji-enhancements.p...
So what do we do with 国 and 國? The first of those is always used in simplified Chinese and usually in Japanese, while the second is used in traditional Chinese and sometimes in Japanese (eg. names). Is this one, two or three characters?
More on the topic: https://en.wikipedia.org/wiki/Han_unification
I'm personally of the belief that the accented and cedilla characters should be exclusively stored as combining character pairs, even if modern keyboard mappings require only a single keypress. My own language stores every character as two bytes (at a minimum), so the storage aspect is a solved problem.
If all three had different codepoints and you replaced 国 with 國 a lot of people would realize, less so if you replaced 國 with 國.
To my understanding the only argument in favour for han unification was that it would have taken-up a lot of codepoints otherwise.
I can assure you that Han unification happened at the hand of native speakers.
Did you know that some languages distinguish dot less Iı and dotted İi ? English mixes them Ii and unicode needs to know exactly based on what language you might want to upper/lower case because it can't tell an english I appart from a Turkish dotless I.
The CJK variant issues under not specifying a lang are indeed present in Latin but to a smaller extent.
Looking at their doc [0] it seems they used their Adobe-Japan1 to wrap a much more wider set of characters than any single encoding standard, including ligatures, vintage encodings etc.
It seems to be a pretty big work and kinda fits with the image of PDF handling being such a monumental beast.
And part of the reason why I like PDF.
( Behind a Paywall ) https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e70...
Can't blame them much though, han unification is a huge mess and designed by someone who I can only posit to be entirely brainless. There aren't many characters that are affected, you aren't even saving any considerable amount of codepoints. It's just west-centricism and lack of knowledge on the subject.
Separate sets of Han characters cannot be encoded in the 16-bit space, but they could have been easily encoded in the current 32-bit space.
Nevertheless, I have never found this to be a problem in practice, because I have always taken care to have good separate typefaces for Japanese, Traditional Chinese and Simplified Chinese.
In documents that I create or modify, I apply styles with the appropriate typeface.
The only possible problems are with Web pages, but the good browsers allow you to configure typefaces for each language and I always configure the correct typefaces.
If the Web page does not specify correctly the language, it might be displayed wrongly, but this is only one of the many stupid things that can be done by a Web page designer that can make that page look ugly when rendered on other computers.
To avoid such configuration work when you prefer better looking typefaces instead of some standard system defaults would require a standardization of how to notify the applications about the association between certain typefaces and languages, e.g. by some environment variables or by some standard locations for the font files, depending on language.
While there were no existing legacy encodings allowing to write Chinese and Japanese at the same time, so there was nothing to keep compatibility with.
More likely they'll think the content was written by a non-native Japanese speaker, judge whether that makes you trustworthy or not (based on personal experience or stereotypes or prejudice, probably a bit of all three (we're all human)) and then not buy from you. A good example would be Amazon listings in Japanese that Japanese people can tell were almost certainly written by someone Chinese, and then decide not to buy.
If you want the cash, get a proper translation. Ironically, Japan is filled to the brim with incredibly poor English and abounds with stories of native English speakers' translations and corrections being disregarded because "it doesn't sound right"… to someone who can't string a legible English sentence together.
Follow the choice of user, please.
Question: was that a Japanese or Chinese character?
Answer: locale doesn’t help us here, unsolvable problem.
I think we clearly need an in-band solution. Some character that switches the Asian glyph variant, or separate characters altogether. The former would be annoying for non-variable sized Unicode because you’d lose the ability to use random access into a large corpus of text, because you’d need to scan the entire text to find out the current Asian glyph variant mode.. sigh
There's some similar issues with right to left, and left to right text. You can give people a good default, and try to be smart, but some cases will always be ugly.
On an iPhone using,
Chinese simplified keyboard: 刃
Japanese keyboard: 刃
So that didn’t go very well. When choosing the character on my Chinese keyboard it is displayed with correct Chinese strokes but turns into the Japanese version in the text box. I’m guessing for most of you reading, both will appear Chinese.
EDIT: Someone better than me at wrangling unicode can maybe try out the variation selectors, and print the correct variations in a comment. I think it would have been neat if my keyboard ime did it for me :)
https://en.wikipedia.org/wiki/Variation_Selectors_(Unicode_b...
> If the glyphs don’t exactly look like the Japanese result sample below, your code is displaying Japanese wrong.
Maybe exactly isn't the right term here; it doesn't need to be pixel-perfect, there are still different font faces just like with western languages, for example one that's supposed to make them look more natural or hand written and one for print, etc.
Also, afaict han unification was a mistake, but if you thought you only ever have 65535 code points available it might have been tempting.
Can this be fixed? New character code for ambiguous characters could be assigned, of course this would require manual conversion (with knowledge of the variant) for existing data, but at least it would make this issue go away moving forward (and unconverted legacy data would be "just as bad" as it used to be, so no loss).
0xxxxxxx
110xxxxx 10xxxxxx
1110xxxx 10xxxxxx 10xxxxxx
and
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
the only reason not to push those last bits and add
111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
11111110 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
and maybe even
11111111 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
is utf-32, they should have dropped it and solve the codepoint problem this way.
But no, there is no particular reason to introduce a longer encoding than the modern UTF-8 (which is actually shortened from the original one-to-six-byte encoding). The current set of 1,114,112 Unicode characters is sufficient for at least the foreseeable future, because any new assignment requires a demonstrable historic or current use. (Emojis are slightly different, but they still require that the underlying concept is widespread and do not significantly overlap with existing emojis. See [1].) Han characters are the largest source of new assignments to this date and they are yet to reach two out of 17 full planes (that would equate to 131K characters).
The other approach is to assign to each language distinct codepoints, but I guess the current approach is better for backward compatibility with pre-Unicode, and less redundancy in Latin script documents.
Checking the source, the page _is_ specifying language tags in the span, which I guess is supposed to help. My system just must not have fonts for those languages so I obviously can't even test them.
EDIT: Replaced with images.
At least for Chinese, the difference is font, they are all valid writing for the character. Different writing style can cause these minor difference as well: http://qiyuan.chaziwang.com/pic/ziyuanimg/E58883.png
If you look at right side of above image, you can tell how the same character is written in different writing style
It is less a problem for Chinese. Our brain has trained to read them, I would recognize the Japanese version, Simplified Chinese version, and Traditional Chinese version without noticing the difference. But I can imagine it can be a problem for Japanese, and other people do not read Simplified Chinese. Having the locale explicitly set to a country and load the correct font make a lot sense here.
The issue is - when there's no proper glyph it uses a fallback?
Or... The context of what glyph will be rendered depends on the document defined locale.
(If it's the latter. it means it's impossible to quote another Asian language text within the same paragraph?)
It is there in the text, but it’s almost hidden between the lines.
If I was a developer with no knowledge of Han characters or Han unification I would have to read two thirds of this article thinking I’m doing it right, so why am I reading this, e.g.: “but I am using the correct code point. It’s the character that the user entered!” or “I copy pasted it from a Japanese text, what do you mean I’m using the wrong character?” before reaching the “how to fix it” and even then I might not realize the root cause.
With that in my mind I might not even make it to the part about how to fix the problem and learn that I am using the right character/code point, but it is still displayed wrong.
For me this was very hard to wrap my head around the first time I encountered the problem. Maybe other people find it hard to understand in different ways.
I believe that Unicode even claims that distinctly looking characters are to have their own code points, but similarly looking characters should share a code point (e.g, there is no French a and English a, even though they are pronounced differently. And Danish ø and Swedish ö are pretty much the same pronunciation but differently written, so they don’t share a code point.)
I am also surprised at the support this problem now has. At least on this thread. Generally speaking Han Unification problem dont get much if any support on HN. Not even empathy. In the name of having Unicode becomes king they would much rather sacrifice the CJK language.
The answer or replies were always, it is "glyph" problem, not "code" problem. Stop asking Unicode to solve it.
patio11 aka Patrick McKenzie from Stripe has been the most vocal critics of Han Unification. Sums it up far better than I could, quote [3]:
>Reason the Han unification debate in Unicode got so acrimonious, and why lots of Japanese people carry a chip on their shoulder about it to this day.
>"Sorry, grandma, I know you've been sort of attached to your name for the last 80 years, but the white folks find it inconvenient for their computer systems. Don't worry, they promise they'll make something close for you."
>Many of the clients of my ex-day job are married to legacy encodings like Shift-JIS precisely because they do think that their customers and students have a "right" to having their names written correctly.
As mentioned in my other reply, Adobe gets lots of stick for its subscription and malware like Creative Cloud. But they do [4] spend huge amount of resources on CJK fonts, layout and encoding ( They have their own separate Encoding for each CJK language instead of using Unicode ). Part of the reason why I like PDF.
[1] https://news.ycombinator.com/item?id=3906253
[2] https://hn.algolia.com/?dateRange=all&page=8&prefix=false&qu...
[3] https://news.ycombinator.com/item?id=1438749
[4] https://ken-lunde.medium.com/my-28-years-of-adobelife-e97e70...
Is there a resource to read more about this? I don't get that vibe from things like:
We also have a website about a Japanese game using a Japanese font except for 0x9bd6. The font's 0x9bd6 is the CN variant and its 0xe001 is the JP variant of 0x9bd6. Fun times.
Like others said, on the web, you pretty much have to manually assign lang to every single thing. We just added support for CN/TW/KR text. I should come back and check 0x9bd6 in the other versions ...
However, this also meant that characters which differ in appearance across languages, such as
<span xml:lang="ja" lang="ja">刃</span>
and
<span xml:lang="zh-Hans" lang="zh-Hans">刃</span>
and
<span xml:lang="zh-Hant" lang="zh-Hant">刃</span>,
were given <strong>identical code points!</strong>
Yes but actuallly "α" is the Greek character alpha, whreas "a" is the Latin character "a". So if you displayed "α" as "a" to a Greek person that, too, would look αλλ ωρονγ.
- Simon from TGM :)
https://upload.wikimedia.org/wikipedia/commons/b/ba/JIS_and_...
better than unicode, you can't be helped.
Not that I like unicode much either -- amongst other things the idiotic arrangement of codepoints makes it basically impossible to do remotely efficient text processing; e.g. here's a graph of the automaton the RE2 uses to check if something is an uppercase character:
https://swtch.com/~rsc/regexp/cat_Lu.png
(For ascii there would exactly be a single arrow connecting two nodes).
Seems like you're right, at least as far as Firefox is concerned. Testing the data links below, it appears to guess the default language based on the encoding used. Neat!
data:text/html;charset=euc-jp;base64,PHA+RGVmYXVsdDogv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgv8/EvrOks9G5/Mb+PC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgv8/EvrOks9G5/Mb+PC9wPgoK
data:text/html;charset=euc-kr;base64,PHA+RGVmYXVsdDog7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIg7NPywfqtysfN6ez9PC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIg7NPywfqtysfN6ez9PC9wPgoK
data:text/html;charset=gb2312;base64,PHA+RGVmYXVsdDogyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgyNDWsbqjvce5x8jrPC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgyNDWsbqjvce5x8jrPC9wPgoK
data:text/html;charset=big5;base64,PHA+RGVmYXVsdDogpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9ImphIj5sYW5nPSJqYSIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9ImtvIj5sYW5nPSJrbyIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLWNuIj5sYW5nPSJ6aC1jbiIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLWhrIj5sYW5nPSJ6aC1oayIgpGKqva78qKSwqaRKPC9wPjxwIGxhbmc9InpoLXR3Ij5sYW5nPSJ6aC10dyIgpGKqva78qKSwqaRKPC9wPgoK