Heck, most of us are effectively different people than we were in the past. And we very much feel some shame when we find out something bad we did back when.
Do you feel shame when you read about someone stealing a bike in a far country whose capital you wouldn’t know ? You must be in despair over all the crimes committed by humanity every hour
Do I also think it is arbitrary? Yeah. But I am also ok not feeling shame at stupid crap my immediate family has done. I just have no surprise that others do.
My guess is you are taking a narrower view of this than you intend? People don't necessarily take personal pride in what others have done. But identity is both personal and societal. And any accomplishment of identity is something that people are likely to feel pride in.
That is, pride is often tied to identity. And identity has a pretty wide brush.
On my phone I just added Japanese as a language and it now prefers those. Windows seems to support it fine too, as does Chrome on Linux.
I guess there are problems if you try to use two languages in the same application.
Is the text you're reading getting rendered in Japanese characters? Or is it getting written in similar-but-not-quite-the-same (and thus something you can use, but irritating to do so) Chinese characters?
Then again, it is a difficult taks. For example, there are three ways to write the first (topmost) stroke of 言: as a slanted dot (now PRC standard), as a short vertical (very common in pre-war prints, still used in Japan) or as a short horizontal (also very common). Sure, one can make three codepoints out of these, no big deal. But then 言 occurs in a gazillion of other characters, as in 說這語信 and so on; all of these characters would then need to get codepoints of their own (which btw is exactly what happened with 說 vs 説 (and incidentally 说)). It gets worse when you consider that another one or two components of a given character may also have two or three variants, then you might get 3x2=6 or 3x2x3=18 codepoints for what most people consider a single character. And did I mention we already have around 100'000 codepoints for CJK Ideographs? That number would multiply considerably.
Although I've worked extensively on these problems—what is a variant, what deserves a codepoint—all I've found that there can and should be guidelines, but there's no very hard and fast rule. And I'm afraid variant selectors are a cop-out, not a solution.
In practice in the pre-unicode world fonts were encoding-specific, so the encoding did help. When you viewed a Japanese document it would be in a Japanese encoding and you'd get a Japanese font, when you viewed a Chinese document it would be in a Chinese encoding and you'd get a Chinese font, and when you mixed up your encodings you'd get something that was obviously broken even to someone doesn't speak Chinese or Japanese.
> but those apply across all concerned languages and are not unique to Japanese
They're de facto unique to Japanese; Chinese support "won" in most international contexts, so "CJK" codepoints look fine in Chinese, and those codepoints are rarely used in Korean (mainly in historical documents) so having them rendered "wrong" there matters less in everyday life. Admittedly looking at another article that hit HN today I could believe Bulgaria gets equally shafted (and is probably, sadly, less capable of sustaining a "made in Bulgaria" software ecosystem than Japan is).
> Sure, one can make three codepoints out of these, no big deal. But then 言 occurs in a gazillion of other characters, as in 說這語信 and so on; all of these characters would then need to get codepoints of their own (which btw is exactly what happened with 說 vs 説 (and incidentally 说)). It gets worse when you consider that another one or two components of a given character may also have two or three variants, then you might get 3x2=6 or 3x2x3=18 codepoints for what most people consider a single character. And did I mention we already have around 100'000 codepoints for CJK Ideographs? That number would multiply considerably.
You don't have to add 18 variants, only 3, and that's the worst case. We don't need codepoints for letters that are half-Japanese half-Korean, or every other conceivable variant. But using Unicode shouldn't be a regression from using Shift-JIS.
Nobody wants to go back to a world with hundreds of relevant and thousands of legacy encodings.
> when you mixed up your encodings you'd get something that was obviously broken even to someone doesn't speak Chinese or Japanese.
Obviously much better when nobody can read anything than forcing some people to read texts where there's the occasional odd ('non-native' if you exaggerate a bit) font style choice.
> We don't need codepoints for letters that are half-Japanese half-Korean, or every other conceivable variant
Things are more complicated than that, I'm afraid. For one thing there's time, so setting the language in your HTML document or element to Japanese will give what is considered correct in Japanese today, not what was considered correct pre-war or any one time before that. There's no formal way to express that. Second, stylistic choices on one and language + region selector on the other are cross-cutting concerns.
I have to say that I find your criticism not very well founded. What you can do today is (taking HTML+CSS as the obvious choice) is proper font choice, `@font-face` declarations, proper downloadable fonts, proper configuration of OpenType font feature flags, if necessary on the basis of Unicode code point ranges. (There's also Unicode variation selectors, bit I've never worked with those so can't say how well those work.) This will get you a long, long way to achieving correct and typographically pleasing output for any text be it Chinese traditional or simplified, Japanese, or Korean. It is some work, but then it is a vexing and somewhat convoluted problem, too. I think one should give font designers more credit here because good CJK fonts typically go to great lengths to satisfy even the finickier type heads among us, enabling variations that I'd frankly just gloss over for heavens sake.
It's not like the pre-Unicode chaos of mutually incompatible encodings did anything to help you with any of these points. Rather, it locked you into a rather small (when compared to Unicode) set of codepoints, and you typically only could do one setting per document. So, no writing about Cuneiform or Hieroglyphs in Japanese. Today we can intermingle LTR and RTL scripts, and you can freely mix hundreds of scripts in a single document and a single encoding. None of that was possible in the good old days.
Japan does (assuming agreeing a unicode-like encoding that actually worked is off the table), for the simple reason that practical support for Japanese in "international" applications has regressed since those days.
> Obviously much better when nobody can read anything than forcing some people to read texts where there's the occasional odd ('non-native' if you exaggerate a bit) font style choice.
https://wiki.c2.com/?AlmostCorrect . Failing obviously and universally is better than failing subtly in edge cases.
> What you can do today is (taking HTML+CSS as the obvious choice) is proper font choice, `@font-face` declarations, proper downloadable fonts, proper configuration of OpenType font feature flags, if necessary on the basis of Unicode code point ranges. (There's also Unicode variation selectors, bit I've never worked with those so can't say how well those work.)
You can do a lot in HTML because HTML has a lot of specific support - something that would likely also be true in a world without unicode. (In the unicode world we have only one encoding per HTML file, but there's no reason a file format couldn't support multiple encodings in the same way as supporting multiple language spans etc.).
> It's not like the pre-Unicode chaos of mutually incompatible encodings did anything to help you with any of these points. Rather, it locked you into a rather small (when compared to Unicode) set of codepoints, and you typically only could do one setting per document.
The old system made the easy things easy. In those days effectively any program that could display French or Swedish documents could also display Japanese documents.
The unicode world makes some things that were previously hard easier, such as mixing languages (as long as the lanugages you want to mix aren't Japanese), but your file format has to go beyond "plain" text for the most basic functionality of properly displaying monolingual Japanese documents.
In the old days if you didn't think about languages at all your program would be broken in every non-English language. Now it's broken only in Japan. That's an improvement for most places, but it's a huge regression in Japan, because previously most programs made some effort to deal with internationalisation issues and now they don't, they just say "it's unicode so we don't have to think about it".
So you write, You can do a lot in HTML because HTML has a lot of specific support - something that would likely also be true in a world without unicode. (In the unicode world we have only one encoding per HTML file, but there's no reason a file format couldn't support multiple encodings in the same way as supporting multiple language spans etc.). Basically you reject a format—HTML+Unicode—that has seen tremendous global acceptance, that has brought us 'rich text' (to quote Microsoft), near-complete universal language support for hundreds of previously entirely unencoded writing systems, ability to display ruby text (i.e. small annotations, parallel version of text; btw. not limited to Japanese at all) and want to replace that with 'a file format [that] support[s] multiple encodings'. I agree that it would sometimes be nice (i.e. nice-to-have) inline encoding switches in HTML. You could totally write a JS method to do exactly that, right now, so go ahead! But to replace HTML+Unicode with 'some unspecified file formats' that 'could have' multiple encodings in a single file? Doesn't sound like I want to go down that particular cul de sac.
OK now on to that other point: you write your file format has to go beyond "plain" text for the most basic functionality of properly displaying [some features for some linguistic situations], if I may be so bold and generalize to avoid that fixation on "boo it's Japanese again and only Japanese that has been put at a disadvantage". So where I agree with you is that historically the Unicode consortium was too reluctant to acknowledge responsibility for some features of written text that should be expressible in the very text, not in a 'side channel' (e.g. in HTML tags). This would IMO include markers for start and end of stretches of a language, so you could write "good day, {lang=fr}Messieur{/lang}" where the curlies symbolize where Unicode control character sequences (not unlike those already used for flags) should appear. There's more stuff like that like e.g. indicating position of text relative to other text; we sort of have it in the form of CJK Ideographic Description Characters and Control Character for Ancient Egyptian, but nothing for the general purpose.
It's not rosy glasses lol, I have to deal with encoding differences every day. I'd like nothing more than to be able to forget about them if there was something like unicode that actually worked. Instead I watch support for the language I'm using degrade day by day because everyone high-fives each other that they've got unicode now so they've solved all the problems (and they have! As long as you're not using Japanese), and then they add insult to injury by trying to tell me it's better.
Sure when you have unlimited time and unlimited resources why accept anything that is even a hair's width not quite like what you intended—perfection?
Meanwhile we have Unicode with all its flaws. Sick!
I wonder why they couldn't have done a composition system like Hangeul?
We don't live in a world where any part of this holds water, at least I don't. Care to elucidate where exactly HK and Taiwan got their full set but Japan didn't?
> I wonder why they couldn't have done a composition system like Hangeul?
The answer to that question is because it's too difficult. I worked a bit about automatic CJK character generation ion the early nineties, around the time Unicode v1 and then v2 came out. Nothing at the time resembled anything anywhere close to aesthetically pleasing output. Quite a number of people tried over decades, not a single paper demonstrates acceptable character shapes, only 'legible' and 'bearable' ones at most. Given recent advancements in AI text-to-image, it wouldn't surprise me if someone has already tried their hand at providing a model for CJK generation. But even then you'll be computationally much cheaper to just generate images for given code points plus those few unencoded characters that have to be generated ad-hoc.
Characters from "Traditional Chinese" and "Simplified Chinese" have different codepoints even for "the same" character, whereas characters from Japanese (or old Korean) are made to share Chinese characters' codepoints even for visually different characters.
You'll find all the gory details of what changed between JIS encoding editions in this fine book: 大熊肇: 文字の骨組[1], esp. pp164—178 where all the fine details between different editions of the JIS encoding are listed. It's mind-boggling! It's almost as though the Japanese standards body itself has been having slightly different opinions about their own writing system over the decades. Now if you're intent on getting all the fine details just right for your print edition you can either apply a newer or older JIS encoding to your document, OR invent that file format you're talking about where you can switch encodings mid-way, OR access OpenType font features from your document somehow... mmmh... how to do that... mmmh...
Oh I know! Let's use Unicode and HTML and CSS and get full access to OpenType font features for any point in the document on a per-codepoint basis.
For those folks who really pine for the bad old days when all we had was US ASCII plus an insane number of other encodings, please do read this single Wikipedia article: https://en.wikipedia.org/wiki/ISO/IEC_2022 it will change your mind.
* [1] https://www.amazon.co.jp/%E6%96%87%E5%AD%97%E3%81%AE%E9%AA%A...
No it doesn't. Unicode has no problem with including multiple codepoints for characters that are (arguably) not visually distinct, and has already done so, many times over, not least for simplified vs traditional Chinese (but also for e.g. Cyrillic vs Latin). Just... not for Japanese.
> It's almost as though the Japanese standards body itself has been having slightly different opinions about their own writing system over the decades.
Slightly different, sure. Doesn't mean "just write it all in Chinese" is acceptable. I'm not asking for pixel-perfect everything, just for the ability to write the native everyday language of over 100 million people without having to use some half-assed out-of-band font setup that always has some funny edge case where it breaks.