Your Code Displays Japanese Wrong
heistak.github.io
heistak.github.io
Every organisation I’ve worked at meets attempts to correct apostrophes and curly quotes in their English copy with such an intense mixture of indifference, confusion, and annoyance, that anybody who tries is quickly convinced that they are the insane ones to pay attention to such inconsequential nonsense. Use of tabs vs spaces is easier to turn into a legitimate problem with business impact by comparison.
The article doesn't focus on it as much, but it's also just as much an issue for Taiwan vs Mainland Chinese script. There's a lot of subtle font differences, and traditional Chinese rendered in a simplified Chinese font looks very much out of place, even if the right hanzi code points are used.
There’s good reason to stick to the basic multilingual plane.
That can happen in the BMP too. It's a well-known problem with well-known solutions. (E.g. you can ship fonts that support all the codepoints you use).
It looks like now Apple flipped the default to Simplified Chinese first maybe.
Some have "minor" different with the same codepoint.
I don't think that Han unification itself is too problematic: every variant is attested to some extent in handwriting in all these societies. Furthermore, the web designer conscientious enough will
1. For applications meant to be displayed in a particular language, use a font for that language, in which case the lang-untagged content always appears right
2. For multilingual content ensure that lang tags and fonts are set up right.
If the owner fails to do this, fallback font usage is going to be just as great a problem as variant issues.
Many Hanji/Kanji/Hanja characters have a long history of stylistic variation and simplification for the sake of aesthetics and convenience. For linguists and historians who are used to such variations, the direction of a minor stroke that doesn't alter the meaning would seem to be a purely stylistic choice, just as serifs on Latin characters don't alter the meaning. Some variations just happen to be more popular in some regions/countries/contexts and not others.
The general public, on the other hand, in each country is educated with the single "correct" variation favored by the government. Everything they read now uses the officially approved variation, so other variations look wrong. That stroke should absolutely not protrude to the other side, or you're a filthy barbarian!
The Unicode consortium seems to have listened more to the linguists and historians in this case. The academic stance doesn't always fit with public perception, which is often seasoned with a large pinch of nationalism.
The concrete term is the "normalization rules", which dictate how given arbitrary characters are transformed into domenstic variants. As far as I know most countries with significant Han character usages already had one before Unicode, and the ROK rules were (and still are being) developed alongside with Unicode.
Some background can be found in https://www.unicode.org/versions/Unicode1.0.0/V2ch02.pdf
(Note that normalization rules themselves are distinct from the eventual unification. It takes further works to actually decide whether the unification is possible or not.)
[1] https://appsrv.cse.cuhk.edu.hk/~irg/irg/irg53/IRGN2420_KRNor...
The lack of this causes so many problems and makes so many things that could be easy so much more difficult.
Why do English, Spanish, German and Italian all share the same homogylhs, even though German has for instance has ß among others, Spanish has ñ, and Italian completely lacks 5e letters JKWXY. Greek or Russian though? Completely different character sets despite containing many of the same glyphs. That to me seems like “Western European Elitism”
There are of course special cases where for instance the upper and lower characters don’t match across languages, but they can be resolved individually rather than as a full set.
It doesn't in practice. For example: 机 in Japanese means "table", while the same codepoint in Simplified Chinese means "machine" (i.e. 機 in Japanese).
Furthermore, the primary motivation for UniHan wasn't to 'clean up' similar orthographies to avoid encoding homoglyphs. It was to stay inside 16 bits of coding space. Unicode was fighting a civil war against UCS, which proposed a 32-bit codepoint space, which would mean 32-bit characters, which the entire industry NOPE'd out of. Turns out, UCS was right, 16 bits was not enough, and the result is that every system that jumped on the Unicode train early[1] is now permanently cursed with improperly handling less-common kanji and most emoji.
For the record, I consider both UniHan and the hypothetical "UniGreek" a mistake. What characters get encoded in Unicode should match what speakers of a given language would consider distinct characters, not what we can arbitrarily merge to fit under a given coding bitrate. The only viable long-term solution for internationalized text was 32-bit codepoints encoded using a variable length encoding compatible with ASCII. We wound up with 20-bit codepoints, but I suspect at some point we'll need to break UTF-16 some more.
[0] This is known as "yurei moji" in Japanese
[1] Windows, Java, and JavaScript[2] all mishandle codepoints outside the 16-bit basic multilingual plane, which were encoded as pairs of 16-bit surrogate values in a special range of non-codepoints. Modern UTF-16 handling is supposed to treat these as single astral characters and reject broken surrogates, but the systems in question retain Unicode 1.0 era quirks for backwards compatibility.
A few years after the Unicode/UCS wars, Ken Thompson and Rob Pike would propose UTF-8, implementing it in Plan 9. This encoding was and is superior to UTF-16 in every possible way - including support for 31-bit code points, which UTF-16 surrogates can't do. UTF-8 isn't mangled by anything except the above programs... and MySQL, which had to add a second "no seriously it's UTF-8 for real" encoding.
[2] ActionScript inclusive. Yes, Ruffle has its own wide string library because of this.
Won't combining Latin and Greek alphabets be horrible? Many formulas for starters, but also things like "β-radiation" (although I suppose b-radiation works too).
I don't really have strong opinions on Han unification as I don't speak those languages, but "Euro Unification" would seem like a right pain, and in general I feel Unicode is often too clever by half.
It's a much bigger issue on Linux, Android, and Windows prior to Windows 10, however.
I have no idea if there's any sane way to make this work. Since I only speak Japanese, my fix on Linux is to prepend IPA PGothic/IPA PMincho (which only contains Japanese glyph) to lang=ja, which worked quite well.
Every phone that was sold in the region I used to live (SEA) also lacks Japanese locale out-of-the-box. I had assumed this was still the case (locale availability tie to region being sold).
I haven't said much about my personal feeling in that discussion, but honestly speaking as a Korean, I do feel that some Japaneses weigh too much on exact glyphs. These variant glyphs are indeed "wrong", but so does many non-standard shorthands (e.g. ⿸广マ as a shorthand of 魔) or individual preferences (e.g. the first character of Yoshinoya 𠮷野家 is an unusual variant of 吉). Why some are okay while others aren't? I'm yet to hear plausible justifications for this disparity.
> , but many simplification attempts also had similar ambiguities.
No one asked to have our languages changed by Unicode
Not to mention sometimes you also get kana displaying in a Japanese font and kanji displaying in a stylistically different Chinese font and it looks awful.
start Greek, codepoint 'a' .... end Greek: 'a' wold be 'α' start English, codepoint 'a' ... end English: 'a' would be 'a'
Could even have the same codepoint for Cyrillic alphabets
Once upon a time Unicode had inline language tags [1].
[1] https://en.wikipedia.org/wiki/Tags_(Unicode_block)#Legacy_us...
> Those could even help unify some latin alphabets and save you even more codepoints, [...]
This is a common misunderstanding; they may look alike but can and often do have different rules (e.g. ordering) underneath, so they have to be disunified.
What should software do when it accepts user content? What if a user wants to make a comment containing one quotation in Japanese and one in Chinese?
It can also be problematic if you want to include japanese and chinese ideographs in the same document.
For example what is the probability of such character being rendered incorectly in some standart tex? lets say a wikipedia article.
Even more so the argumet that people don’t report this because they are "not speakers of English!” is just an assumption. Not to mention that translation applications are more than good enough for such a task.
Frankly, people have learned helplessness[0] about these oddities and don't think to report them when they see them, so the inference that something isn't serious just because it's not pointed out is weak.
In the first place, the proportion of software users who raise issues on GitHub/other is small, and when devs are a group of people who communicate in characters that are not used in their daily life, the translation apps they have at hand is not very encouraging.
[0]: https://en.wikipedia.org/wiki/Learned_helplessness
(Disclosure: I'm CJK native)
When copy pasted it here in this comment box (Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:120.0) Gecko/20100101 Firefox/120.0), I also got Simplified Chinese. What do you see?
刃直海角骨入
By the way, I do NOT have an Intel Mac, I have an Apple M1.
App localization (where you specifically select "Japanese" as the language you want an app to display in) is the main target of this article, I think.
By way of contrast, if you paste that into the search bar on https://yahoo.jp/ , you'll get the appropriate Japanese characters, as that site is explicitly Japanese.
When I paste it into VS Code it also renders Japanese, but when I click on the unicode code, it goes to https://symbl.cc/en/5203/ and https://symbl.cc/en/76F4/ which renders Chinese.
I am typing out the Chinese equivalent for reference: 刃直 but looks like they also turned into Japanese glyph for display past Chinese input method, in fact I can't type out the Chinese glyph in any apps on macOS.
It is not a big deal to me because I have no issue reading kanji as Chinese characters.
If someone enters a form with these characters how can a site know what language to show it as. It makes no sense to work this way.
The jump from 8 bits to 16 bits effectively doubled memory costs and still wasn't enough to achieve its goal.
Later it turned out that 16 bits wasn't enough anyway, but we are left with the UCS-2 mess regardless. UCS-2 and CJK unification are historical mistakes in the development of Unicode that we are stuck dealing with.
刃直海角骨入
in Slack. And it shows the simplified Chinese version, like HN.
What's the intended solution for encoding text that uses Chinese and Japanese at the same time? For example, in an article discussing the differences between said texts? Is there some kind of hacky workaround, emoji colour selection style, or are you doomed to using pictures?
This also reminds me a lot of the early "auto-detect encoding" tools long ago that tried to guess the encoding based on frequency distributions. There's a good reason to not do that to unified codepoints for CJK, but I wonder if there's a more sophisticated backwards-compatible solution here.
Edit: looking at https://en.wikipedia.org/wiki/Variant_Chinese_characters , some of these differences are really splitting hairs; if you gave me a text with some of the simplified characters swapped out with some of these variants, I would honestly not be able to tell the difference. There's probably more variation in the Chinese fonts out there than I can gleam from the article. I'm not convinced that some of these characters have a non-artistic difference that a fellow Chinese on the street will be able to tell apart.
For web documents, you can explicitly mark the language and let the browser pick the correct font. Can be applied to any element:
<span lang=“ja”>…</span>