‘Ghost kanji’ lurk in the Japanese lexicon
japantimes.co.jp
japantimes.co.jp
https://news.ycombinator.com/item?id=17637375
A spectre is haunting unicode:
> The fact is, most ghost characters have yet to be fully understood. One such example is 彁
while the blogpost says
> In the end only one character had neither a clear source nor any historical precedent: 彁.
According to the table in Wikipedia[1], of the 28 characters which were originally considered ghosts, 16 were found to be legitimately used to write place names, 8 occurred in earlier dictionaries, 3 (including 妛) were similar to (and probably a mis-read form of) some dictionary entry, and 彁 is of unclear origin.
[1] https://ja.wikipedia.org/wiki/%E5%B9%BD%E9%9C%8A%E6%96%87%E5...
Unicode is full of such mistakes, now stabilized thanks to the Stability Policy [3]. Thankfully having some character mistakenly encoded and/or named in Unicode seems to hurt virtually no one.
[1] https://blogs.adobe.com/CCJKType/2016/03/bahts-is-parts.html
100 บาท (in the local language)
100 B
100 ฿
100 BAHT
100 BATH (disturbingly common)
None of it matters of course as the meaning is totally evident from the context, but it's interesting coming from countries where symbols like $ or ¥ are totally dominant. ฿ is probably the least common of all of them (yes, less common than "bath") and it's kind of easy to see why - there's only one conceivable meaning a capital B could have coming after a number in Thailand, so why even bother learning the keystrokes to bring up the unicode?It is read "ka" or "sei".
Interestingly, the dictionary also gives them readings. I wonder how they come up with them?
[1] http://www.asahi.com/special/kotoba/archive2015/moji/2011081...