A spectre is haunting Unicode
dampfkraft.com
dampfkraft.com
He’s done awesome work in the Japanese NLP space over the last decade which has really helped me in my language learning projects.
He maintains a mecab (Japanese tokenizer) wrapper for Python [0], has a book on Japanese NLP written for English speakers [1] and also worked on Spacy at one point [2].
[0] https://github.com/polm/fugashi [1] https://www.japanesenlp.com/ [2] https://spacy.io/
á (á) is also an 'a' with a diacritical acute accent. If you mean that ÿ should not have a precomposition in Unicode, well, why not, especially if it really is used in "a handful of proper nouns in French and Hungarian".
Remember, the reason we have combining marks is that that is in fact how many of these characters were composed in actual use, especially with typewriters. Heck, 1966 US-ASCII / ECMA-6 (1985), section 5, describes the use of backspace/overstrike in US-ASCII for composition of such characters! That comes from typewriter use. And that's where compose-key sequences generally come from, too.
So it's not at all surprising that given that ÿ has _some_ use, therefore a) it is a valid glyph to construct with combining diaeresis, and b) that it has a precomposed codepoint in Unicode.
(Not sure how the latter will render in your browser).
In fact, (semantics aside, from a technical perspective) the preference should always be for modifiers rather than standalone characters because the chances of being supported by the viewer’s font are much greater: it doesn’t need a separate glyph explicitly drawn and add to the font file for the code point. Difficulties in entering it or typing it out should be mitigated with client-side affordances in the UI, shortcuts, etc.
Yes. But the font has to be designed to allow this.
That means that a) lower-case letters must be small enough to allow "overstrike" with diacritical marks to render correctly, b) diacritical marks must be small enough too, c) if you want capitals to also render correctly then the font must have either a set of smaller capitals, or smaller/higher diacritics, and the renderer must scale the capitals and diacritics to fit, or change line spacing, etc.
Also, the 'semantics' for the _human_ reader are the same whether you use pre-composed or decomposed codepoint sequences -- the semantics for the human are about the glyph as rendered and not the details of how that glyph was obtained.
And to be super-pedantic (sorry!), what you call 'modifiers' are called combining marks in Unicode, and what you call 'standalone characters' are called precompositions in Unicode. And it's not necessarily true that the rendering will _in practice_ work better with the former than the latter, but in theory absolutely it is, and in practice it almost always is for _browsers_.
> Difficulties in entering it or typing it out should be mitigated with client-side affordances in the UI, shortcuts, etc.
I really wish Windows would adopt X11-style compose key sequences. Those are incredibly natural for all glyphs that can notionally be constructed via 'overstriking', and historically that is exactly how people did construct those with typewriters. (I don't know, but I suspect that for typesetting it was necessary to have a type for each modified character because having decomposed movable type would probably not have been robust enough.)
My understanding is that decompositions for Latin scripts was just natural typewriter-style constructions, while precompositions for Latin scripts was also natural to simplify table-driven transcoding between Unicode and ISO-8859.
Similar considerations probably applied in the case of Hiragana (I'm guessing) and other scripts.
Besides, combining marks (decomposition) allow for creating new glyphs based on existing ones even where Unicode does not define them.
Once two or more ways existed to write any given glyph the canonical equivalence problem immediately arose, and the UC was aware of it immediately, thus we get two basic NFs (NFD, NFC).
When it comes to the semantics of glyphs, whatever the UC intended is one thing, but how natural language evolves to use those glyphs is another. So to some degree what the UC intended is a footnote, and what matters is how people use Unicode.
I suspect the reason it renders fine is that 'n' in the font I'm using is small enough that the combining mark can be rendered by "overstriking" a diaeresis.
Surely those aforementioned non-initial cases would sometimes find themselves in a piece of all-uppercase text? You'd have found such things in print media even before the advent of computers.
At least until 2017, I imagine there are others though.
> In 2017, the Council for German Orthography officially adopted a capital form ⟨ẞ⟩ as an acceptable variant, ending a long debate. https://en.wikipedia.org/wiki/%C3%9F
Huh? How do you pronounce 切?
> And it hints at meaning: bowed weapon, sound, elder brother; but your guess is as good as mine!
Why is "elder brother" a meaning hint if you've already assumed that 哥 is the phonophore?
Also... there is no "gather" or "clog" sense of 沌? There's a "gather" sense of 屯...
There are other meanings to it.
I admit I had a hard time finding an example.
Where did you get the original 'example' of 沌 / 屯 from? Most people prefer to use examples that they have some kind of personal knowledge of.
> There are other meanings to it. [沌]
Not as many as you might expect. It's (part of) a mythological term referring to the state of the universe before it took the form we can observe today. It has a couple of other metaphorical meanings coming from that, or from traditional sayings related to that.
I see that wiktionary lists 沌がる as an alternate spelling of 塞がる "be blocked" [塞 - obstruct / block / plug], which looks like the closest thing we're likely to find to your earlier claim. Obvious problems using this to support your comment are:
- There's still no "gathering" or "accumulating" sense. (Nor is there an "obstruct" sense of 屯.)
- This is a meaning assigned long after the character came into existence, which means it cannot have informed the construction of the character.
The dictionary link I posted had あつまる、水があつまる in the first definition. I don't know if that's a reliable dictionary. Maybe you think it's not?
> - This is a meaning assigned long after the character came into existence, which means it > cannot have informed the construction of the character.
I don't know anything about which meanings were used first and which were added later, but what you say sounds believable.
> - There's still no "gathering" or "accumulating" sense. (Nor is there an "obstruct" sense of 屯.)
Let me clarify. Even without the first sense for 沌 in the above dictionary (あつまる、水があつまる) some of the other definitions for 沌 had "gather" included in it conceptually, namely clog and joined together without distinction.
This phenomenon we are talking about is called 会意兼形声文字 and this site seems to have a lot of accurate examples: https://okjiten.jp/20-kaiikenkeiseimoji.html#google_vignette
I confirmed a few, and they seem correct. I thought 界 was the most easily understood example. 介 means being between two things and contributes the pronunciation. 界 means region, span, separate, division.
There are no kana in that word, but if you search baidu for it you'll find a total of zero uses. It's not difficult to understand what the phrase means - "a character that is simultaneously 会意 and 形声" - but it's not an existing term. It looks more like a joke, the way English speakers will sometimes talk about "autoantonyms".
There is a list of six traditional categories of character construction, and we still understand what five of them mean:
象形 - a character that literally pictures its referent, as 日 depicts the sun or 木 depicts a tree.
指事 - a character that metaphorically pictures its referent, as 刃 indicates the edge of a blade by placing a mark next to 刀 ["knife"], or 本 indicates roots by placing a mark at the bottom of 木 ["tree"].
会意 - a character whose meaning derives from the interaction of two semantic components, as 明 ["bright"] pictures the sun and the moon, or 休 ["rest"] pictures a man next to a tree.
形声 - a character with one component indicating the meaning and another component indicating the pronunciation. The vast majority of characters are in this class, but we may use the example 河 ["river" or specifically "the Yellow River", today pronounced hé], in which the semantophore is 氵["water"] and the phonophore is 可 [a modal verb having to do with permission or ability, today pronounced kě].
假借 - a character that is borrowed from some other word (because it shares the same pronunciation). These have tended to be "corrected" over time, but an example would be the tendency in ancient texts to write 女 ["female"] for the word that is today written 汝 ["you"].
(The other category is 转注. We don't know what it means, but we are given an example - it means whatever the relationship between 考 and 老 is.)
> I confirmed a few, and they seem correct.
Really?
Really?
--- EDIT - my discussion of 与 is flawed. There is an ancient 与, and it is given as sharing its pronunciation with 與, while 與 is said to be 会意 with 与 as one of the components. 與 could be fairly called "both 会意 and 形声", though 与 can't. ---
The first example on the page is 与. This is a simplified form (from 與) and it doesn't carry any phonetic or semantic weight. It's a representation of that bit in the top middle of the older form. I can be sure that the page meant to list the simplified form, though, because it's in the category of "three strokes".
----------------------
Moving to the "four stroke" category, we can see 切 ["cut", today qiē]. This is as clear as they come: it has a phonetic component 七 [today qī], and a semantic component 刀 ["knife"]. There is absolutely no possibility that it could be interpreted as 会意, because the meaning of the phonetic component is "seven".
円 is another simplified form. It means "round" (like a circle). The older form is 圓, which does have two components: the semantic component 囗, and the phonetic component 員. The meaning of the phonetic component is "staff; personnel". I'm not seeing the case for how that contributes to expressing roundness.
攴 ["strike; beat"] is another 形声 character with a phonetic component 卜 and a semantic component 又 ["again" - the connection here isn't obvious to me]. The phonetic component on its own refers to divining the future, a concept unrelated to 攴.
(仁 appears to be a fair call. It is identified as a 会意 character for what I assume are good mystical philosophical reasons. But equally it's true that 仁 shares its pronunciation with its lefthand component 人 (and this was also true in the past).)
In the "five stroke" category, we find 氷 ["ice"]. This character doesn't have two components and therefore cannot be 会意 or 形声. Today it is more commonly represented as 冰, which does... sort of... have two components. However, it's a weird case, because the component on the left, 冫, is usually understood as the combining form of 冰 itself, making this character infinitely recursive. The ancient form of the character had 仌 on the left, but I haven't been able to determine to my own satisfaction what that signified. As best I've been able to tell, 仌 by itself is now considered an archaic variant of 冰, and it means "ice", making the ancient character self-recursive in the manner of the modern one. 冰 (or rather the older form 仌水) is identified as a 会意 character, and I guess you can see it that way - you have a character meaning "ice" built from components meaning "ice" and "water" - but since it seems to be identical with its own left component I have difficulty calling it 形声.
At this point, I really don't see any value in looking further into the page.
Originally, 七 meant to cut vertically and horizontally (confirmed on a few sources).
Regarding 與 and 与, the site shows the breakdown for the former and shows how it was the older form for the latter. I guess you didn't actually click on the characters and look for the explanation. The analysis shows it comes from 牙+口+舁. The last is the meaning as well as pronunciation, the meaning being "hold up, carry together"
> > I confirmed a few, and they seem correct.
> Really?
Yes I did click through and read a couple of explanations.
As for 圓, according to the site the inner portion actually represents a picture of a 鼎 (ding). It is surprising to me, and a lot of this is theoretical and there will be more than one opinion. That same site doesn't say 員 itself is related to ding.
> if you search baidu for it you'll find a total of zero uses
There is a page on it right here. I cannot read Chinese but putting the first section into Google translate there is nothing surprising here. It also has examples. https://baike.baidu.com/item/%E4%BC%9A%E6%84%8F%E5%85%BC%E5%...
Indeed, the source link is about exactly this: a crappy scan appears to have 彁 but a better copy reveals it was 彊.
[0] https://emojipedia.org/pregnant-man
[1] https://www.reddit.com/r/technology/comments/u7x3l9/comment/...
And yet we have the phrase "It has a a certain je ne sais quoi."
* https://lingoculture.com/blog/culture/je-ne-sais-quoi/
You can recognize/know that something is special, but cannot quite put your finger on why.
Perhaps the meaning of this character is sealed in some vault
Major previous discussions:
110 comments: https://news.ycombinator.com/item?id=17637375
130 comments: https://news.ycombinator.com/item?id=24951130
180 comments: https://news.ycombinator.com/item?id=32095502
The peculiar properties of CJK characters and the philosophy (apparently the Japanese did not like Unicode's tendencies towards Aristotelian essentialism) under which they were implemented in Unicode probably singlehandedly forced unicode to expand beyond the BMP....
Do you mean Han Unification? I.e. that conceptually equivalent characters which are written differently in Japan and China received only a single unicode code-point, and are rendered the Chinese way by default on most computers?
I think some characters got different code points, while others were merged. And apparently the Japanese complained bitterly over the ones that were merged. If you had read any articles about that, this is probably what you have in mind right now.
And I'm also a bit tilted by the ones that had different code points, because when processing CJK text now we have to deal with characters that are (in my native Cantonese) essentially the same, looks similar (to my undiscriminating eyes), yet having different code points so that things like text search sometimes don't work.
Of course I'm not "blaming" the Japanese, if anything the simplified vs traditional Chinese thing is much more of a practical problem, and the conflicting code points I deal with on a routine basis are more of a Hong Kong vs Taiwan thing, but I was told that the Unified CJK thing adopted a different "philosophy" from the rest of Unicode (which I think really is some kind of Aristotelian essentialism...) mostly due to vocal objections from the Japanese.
I have no idea what "Aristotelian essentialism" is supposed to mean, or if you are saying that the unification was that.
> And apparently the Japanese complained bitterly over the ones that were merged.
and
> but I was told that the Unified CJK thing adopted a different "philosophy" from the rest of Unicode [...] mostly due to vocal objections from the Japanese.
Seems to contradict each other.
[1] https://en.wikipedia.org/wiki/Ideographic_Research_Group
Page 55:
""" there are characters with no glyphs. glyphs that can correspond to a number of different characters according to context. Glyphs that correspond to multiple characters at the same time (with weightings assigned to each), and even more possibilities.
The problem of glyphs and characters is so complex that it has gone beyond the realm of computer specialists and has come to be of interest even to philosophers. For example, the Japanese philosopher Shigeki Moro, who has worked with ideographic characters in Buddhist documents, goes so far in his article Surface or Essence: Beyond Character Model Set [274] as to say that Unicode's approach is Aristotelian essentialist and to recommend supplanting it by an approach inspired by Jacques Derrida's theory of writing [114, 115]. The reader interested in the philosophical aspects of the issue is invited to consult [165,156], in addition to the works cited above. """
I think "essentialist" is probably a good description of the philosophy of how Unicode defines characters as opposed to fonts and glyphs, so I adopted it.
So I’m not sure if there are glyphs with multiple characters but there _are_ code points with multiple characters. E.g, you can enter the correct code point and still misspell a Chinese or Japanese word. This is because for Chinese Korean and Japanese characters in Unicode it is not enough to choose the correct code point but also markup the code point with the correct language.
That's my interpretation, disclaimer I'm not an expert in this stuff.
(Arguably even in English we run into the fact that there is not one "A" - there are many "A"s that sound and act completely different, and we collapsed them into one representation.)
Which is always fun when you think that clearly nobody thinks of "o or o with a leg" (O vs Q).
Many languages are strictly phonetic and their orthography has a very strict 1:1 ratio of glyph:sound. English is not one of those!
So in English, while we have five glyphs for the main vowels, those five glyphs have extremely versatile interpretations when actually converting into the spoken words.
What would that philosophy be about? Sounds apocryphal. Unicode has never done "unification" like that for other languages/scripts?
i/ı/i, ö/ø/ø̈/oͤ: Same same, different codepoints.
Search and sorting is a mess everywhere. Depending on your locale, ö sorts either after o or after z. Sometimes it's semantically and phonetically equivalent to o wrt search but moreoften not. https://en.wikipedia.org/wiki/%C3%96
Conceptually we long debated unifying everything and in an ideal perfect world we would have done it. The reason was one primary goal for a new standard was to make it easily parseable and having unique rather than repeated codes was key to that. Sadly in the end we did not unify everything only to get buy-in from all major countries to support. That’s even why you see the roman/asciii characters repeated within Unicode itself—like as romaji. This was all well in good until we came to CJK and the number characters with semantical overlap was huge that this was more seriously considered—infact we started investigating this at Xerox before even thinking about Unicode and that work predated and influenced and leveraged the work done later.
Japanese users are angry when system chooses Chinese fonts over Japanese, Chinese users apparently suffer with a font mishmash of Chinese and Japanese all the time like a mid-word capitalization. The official sanctioned solution is to just commit a genocide and nuke the offending language out of the system or to attach IVS to every Japanese characters which is like zero padding every single characters by a byte or two.
It's really putting everyone in pain and making absolutely no one happy, which sounds like what a good compromise tends to be, but LLMs seem to be struggling with there being multiple completely separate language families sharing codepoints but not character shapes or syntax. I've seen a smaller LLM model use made up verbs spilling over from the other language. I believe ML guys don't like datasets that do that kind of things. Same apply to image models.
There has to be a reason why it happened(as to why Unicode suddenly started insisting it has to fit inside a 2^16 total chars or whatever).
Cross language search seems like a hack to me. Searching in Chinese should find Chinese words and searching in Japanese should find Japanese words. Being able to search in Japanese and get Chinese results is not what most users want, unless they don't have a proper keyboard.
I think you underestimate the similarities. For instance Japanese names are not translated to Chinese. Theyre just read with Chinese pronunciations. So Chinese will regularly interact with Japanese content (im guess it happens the other way around too, but i dont have the personal experience)
Yes. If you search gâteau it shouldn't show results for cake by default.
>Japanese names are not translated to Chinese
Then it makes sense to search for the names in Japanese. Either they are translated to Chinese so you should be able to search with Chinese, or they aren't and you should search with Japanese.
The point is that when they look identical most of the time, you have no idea which language "mode" the text is in.
Tokyo is 東京 in both languages. They're not visually distinguishable. Maybe in a long list some particular characters are written slightly differently, but you'd have to really inspect the list and hope that distinguishing characters show up.
From the outside this may look weird, but to people that are around Chinese characters having multiple ways to write a character is just a normal fact of life. They look different in classical writing, seal scripts and cursive scripts. Trades people will also use shorthands. You have analogous situations with Simplified and Traditional Character - where some characters are simplified to fewer strokes and others are not. But as a reader you don't really care if it's 吃/喫 or 臺灣/台灣/台湾. There is basically no situation where you want to find 吃 but not 喫.
I get the desire to preserve native Japanese forms of the Chinese characters. But that seems like something mostly resolved with a font
If you want to mix Chinese and Japanese characters, then you are left having to mix fonts - which is a bit ugly I guess. It's not the most ergonomic solution, but this is the edge case. Most of the time you want different forms to search the same. If you search 臺灣 and your word processor skipped 台灣, then you'd be rightfully annoyed.
> As far as I understood it, the result was an incoherent mess.
Do you have any specific examples? I never heard this before.It looks like the CJK unified space is over 20,000 characters, so that's a real technical magnitude distinction compared to Latin and Cyrillic. "Asian languages be damned" seems like a bad faith read, compared to "Java and Windows char is 16 bits and that will never change realistically" (and in fact they still haven't, even in 2026 things which rely on UTF16 instead of UCS2 are still commonly bugged unfortunately)
Plus, these characters were only added 20+ years later, and in the supplemental planes, not the BMP.
The fullwidth/halfwidth stuff is a bit of a mess, but you could argue that these are actually CJK characters that just happen to resemble Latin characters (much like how Greek and Cyrillic both happen to have letters that look a lot like the Latin "A"), since they only exist for compatibility with older CJK encodings. This wouldn't be a very good argument though :)
And then some examples of non-unified Chinese chars: https://en.wikipedia.org/wiki/Han_unification#Examples_of_so...
Note that Han Unification has a history predating Unicode, and the Unicode work very much followed on from that.
The most notable is CCCII, developed in Taiwan in the early 80s, and then standardised in various places.
That's not to say that it hasn't been controversial (it has!), nor to say that it hasn't caused problems (it has!), but it's also unfair to say that it's just Unicode doing its own special thing.
I've seen YouTube videos on this topic before.
> The original character (𡚴) was not added to JIS or Unicode until much later and doesn't display on most sites for me.
It's from Unicode version 3.1 (published 2001) so this is surprising.
Why didn't they simly replace the original bad one?
> nine hundred pages. Imagine tracking down a single character without a page reference
Not that hard to imagine, OCR existed back then?
Starting with version 2.1 they made it so that code points can't ever be unassigned. People were getting really upset with the removals, since they can result in old data being misinterpreted. (Er, nothing actually was removed in 2.1, but it probably didn't become official policy until 3.0… ?)
Besides, it's useful to have the bad characters in there for people who want to write posts like this one discussing them.
... And they can, since a code point was assigned for it; so what's the problem?
> was not added to JIS or Unicode until much later
then TFA is simply inaccurate; 𡚴 has been in Unicode since 2001. More importantly, replacing the incorrect character (instead of supplementing it) wouldn't realistically have made it possible to add it sooner. They didn't really know what they were doing back at the start.
How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.
I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.
For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).
IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.
[0]: Aside from the "Overview of National Administrative Districts" book mentioned in the article
I appreciate the article nevertheless, of course, but I do feel that it would probably be more meaningful to someone that has at least a basic understanding of Japanese.
<Multi_key> <c> <c> <c> <p> : "\xe2\x98\xad" #symbol representing proletarian solidarity between agricultural and industrial workers
Never used, but I laugh every time I see it there.As a slightly related tangent, The compose key mnemonic interface for rarely used characters is pretty great, beats trying to remember alt codes.
never mind i pressed caps lock everything is better now