A Spectre Is Haunting Unicode (2018)
dampfkraft.com
dampfkraft.com
In order to record deeds and property transfers, we needed to enter people's names and official registered addresses into the computer system. The problem was that some people used non-traditional writing variants for their names, and some of their birthplaces were tiny places in China with weird names.
Someone might write their name with a two-dot water radical instead of three-dot radical. We would print it out in the normal font, and the people would lose their minds, saying that it was wrong. Chinese people can be superstitious about the number of strokes in their name, so adding a stroke might make it unlucky, so they would not buy the property.
The customer went to the agency responsible for managing the big character set, https://en.wikipedia.org/wiki/CNS_11643 Despite having more characters than anything else on earth, it didn't have those variants. The agency said they would not encode them, because they were not real characters, just printing differences.
The solution was for the staff in the office to use a "font maker" program to create a custom font with these characters. Then they could print out the deeds using a Chinese variant of Adobe Acrobat, and everyone was happy.
They had that, but needed a font that used a different number of strokes for the characters because of the superstition.
On further review, I think this is als similar to #12 & #13 on the list: "names are case-sensitive," and "names are not case-sensitive." To generalize that to include non-Western alphabets: display variations of the same character are significant, and display variations of the same character are not significant.
This of course goes back to the evergreen philosophical question "what even is a character, anyways?" Since we've found a case where two characters which are the same character are not the same character. Are they distinct characters or typographical variants? Yesn't: one would want them unified for searching, but distinct for printing.
But regardless of what they are, these characters/variants only show up in names. Names tend to retain archaic (or extinct) language variations longer than speech, which is the reason for rule #11, which is at least part of the problem.
I think these language examples are so good, as examples, because all aspects of them are clear and easy to follow. I think computerization of business and society and the systems that make them work, causes immense amounts of this kind of friction and pain all the time, in ways that are much harder to understand, explain, or catalog (which is precisely why it's such a big problem, though as far as I know it's received little attention)
[EDIT] To distill it, I think that trying to make a computer a "source of truth" rather than a tool, tends to do substantial violence to the "truth".
Language and writing exist to communicate, using patterns of signals that have shared meaning and recognition; things like alphabets and vocabularies are effectively (loose, overlapping, diasporic) consensus-state autoencoding models. They only work to compress meaning, when there are rules for said compression that generalize, and which don't have as many exceptions with their own separate symbols as there are words/names needing to be encoded.
Most countries don't allow you to just make up your own novel graphemes when writing a name on a birth certificate. And nobody is asking for that, either. (Presumably because living in a world where that was allowed would be horrible: you'd no longer being able to error-correct when reading, because any given mysterious squiggle in the middle of a word or name, might be exactly what some unknown-to-you-or-anyone-other-than-the-author character is supposed to look like. Is that "o with a curlicue" written here just a semi-cursive attempt at writing an "o" — or is it an "o" with a novel accent marker, one that appears nowhere else, but which must be preserved nevertheless to properly record this person's name?)
Instead, legal names are (in every country I'm aware of) required to be spelled using the character-set of the country you're entering a legal relationship with by being born / immigrating / etc. America? Legal names using the Latin alphabet. Japan? Legal names using characters from this set: https://en.wikipedia.org/wiki/Jinmeiy%C5%8D_kanji
Note, though, that legal names are representations of names. They aren't encodings of names. Your legal name is a distinct thing from your name, just as your credit-card number is a distinct thing from your name. It's an applied-for + registered + assigned systematic identifier for you — a bit like a domain name, or a vanity license-plate number. Which means that your legal name is not a lossy or lossless encoding of your name. It's, per se, a nickname. It doesn't have to have anything to do with your name. (And it often doesn't; immigrants often choose legal names entirely distinct from what they / their home country thinks of as their name.)
> do substantial violence to the "truth"
I don't think this kind of wild escalatory rhetoric is at all helpful to the (presumably) good cause intended. Probably the opposite actually.
Most of the time people are not in fact having substantial harm inflicted upon them by others out to do violence. That mindset seems incredibly fragile, paranoid, and divisive.
The reality is people are just working away to improve things and sometimes they make mistakes or don't have perfect information ahead of time, and other times making things better simply necessitates that not everybody's last whim can always be accommodated. The healthy mindset to have is that not every real or perceived slight against you is done because you are being persecuted by violent hatemongers, and that people should be more accommodating and accepting of the reality that systems and procedures designed for the benefit of everybody may just not be able to accommodate every unique request they have.
Webster's 1913:
> 2. Injury done to that which is entitled to respect, reverence, or observance; profanation; infringement; unjust force; outrage; assault.
> We can not, without offering violence to all records, divine and human, deny an universal deluge. - T. Burnet.
It's a bit more poetic a use, but it's not escalatory in the way you suggest.
[EDIT] Further, my point is (and I think that was clear?) that computerization necessarily does these things if you treat the computer as correct and humans as suspect, not that anyone's doing this on purpose.
"I'm going to shoot you" is escalatory. It doesn't matter that I could have been talking about photographing you.
Heck, my initials are totally non-standard.
https://nymag.com/intelligencer/2016/04/princes-legendary-fl...
"Freur", or, "The squiggle we chose as the name for a band but that CBS Records insisted should at least have a pronunciation".
I see it is not in Unicode (well, you can never really know if you do not try), nor I can find pieces to reconstruct it.
The "freur" in foreground: https://d4q8jbdc3dbnf.cloudfront.net/user/6885/edb290c6183ac...
https://news.ycombinator.com/item?id=226853 (18 comments)
As a contrived example if you had a symbol for 'happy' you want to be very cautious that it doesn't get converted to 'gay' because in your language gay and happy mean the same thing, in some repressive regime it means the leadership gets to execute you with the approval of the law.
Edit: weirdly HN refuses to display that emoji.
The important thing for engineers to note is a technical shortcoming caused a tragic misunderstanding. Focusing instead on the well-known fact that some people have poor impulse control, knowing full well that is a non-controllable input, instead makes an excuse for poor engineering and implicitly expresses powerlessness to do anything about the problem.
But yes, misunderstanding or not, we should not kill people.
The story in the sibling comment is about a man attacking his daughter's ex because the ex came to apologize about a confusion over the Turkish dotless I. That's still a violent attack that the father could have kept his emotions in check. I don't condone calling the daughter names, even accidentally, but it is not a crime and the right response is not attempted murder.
I don't know who you're arguing with, but it isn't me. Nobody is saying it was.
I'm saying it is an irrelevant non sequitur.
Imagine that Dad instead misunderstood an instruction related to a financial transaction and lost a ton of money. Would you now be discounting the technical problem that caused the misunderstanding and berating Dad for being foolish?
If I were on a code review and I spotted an issue affecting Turkish dotless I, I assure you I would rant about it more than is reasonable.
Maybe, though it's still halfway the same word.
> maté
Not a change, both spellings are valid.
Malé parties are a lot of fun.
Those are some pretty lamé runners.
https://en.wikipedia.org/wiki/Yerba_mate#Name_and_pronunciat...
Why am I not surprised in the slightest?
Later versions of Unicode support "Variation Forms" of Han characters as a way to be able to encode different variations. They are encoded as a Variation Selector code (U+E01000 and up) after the Han character. The forms are listed separate from Unicode versions in the "Ideographic Variation Database" <https://www.unicode.org/ivd/>. So far, it contains characters from a couple of Japanese dictionaries, a Korean and one from Macao/Hong Kong.
In fact going any place with her very nearly became an “are we living in a simulation” crisis for me because the number of times she would say her name and the other person would say it back incorrectly was… upsetting. The degree to which some people butchered her name, especially combining half of her first and last name into a completely different name, made us joke about buggy NPCs.
I could imagine how in some cultures writing it incorrectly hurts as much as pronouncing it incorrectly. Or possibly moreso in places where multiple plausible pronunciations have to be negotiated via an introduction, which is the case in China, is it not?
My name got changed when I moved to Spain and it never bothered me, while I have met people who took great offence at the use of standard nicks that they had not explicitly sanctioned in advance. I know a guy who makes a new name up for everyone he meets. Like or lump it. If you are too sensitive about your name, you risk people not using it at all.
I say this while fully aware of my own butchering.
He was in Argentina for the Chess Olympiad in 1939 when WW2 started, so he stayed and got stuck there. It's unlikely he'd be known as "Miguel" now if this had never happened.
EDIT: Aha! this website has a guide on these names, and even dispells the Ellis Island explanation I was told as a kid: https://pgsctne.org/changed-surname-list/
They did it the same way you did.
Nevertheless there have been countless times where people automatically substitute the more common name, or even worse in text messages manage to misread it and reply incorrectly.
It sometimes upsets her. The npc analogy is very apt, i guess many people are just very preoccupied?!
Overzealous autocorrect can happen to names, too. There's a whole thing about Asian names not being in computer spellcheck dictionaries: https://www.abbynews.com/news/youre-not-a-mistake-b-c-group-...
Unhelpful, though luckily found funny when I did it.
There's a trick in Chinese to explain which character you are referring to when it could be any number of homophones- you repeat it as part of another well-known word. To use a crap analogy, you might say "Je like Jeep" or Ge like geography" if your name is Jennifer or Gene. The latter is a bad example since Gene is itself a single syllable, but hopefully it is illustrative enough. On the other hand, if you have a special stroke in a common character (or remove a stroke, as in the post above), I am guessing it's harder to explain.
On a perhaps related note, I was once accosted in the street by a Thai tuk tuk driver who wanted to know which football team to bet on ("just ask any white bloke you find" probably isn't the best strategy but I digress). One of the teams was Portsmouth, but I told him unless I knew who was playing I had no idea of the strength of the team. He held up a newspaper written exclusively in Thai, which does not use a roman alphabet or anything like it, and proceeded to read out the names of the players almost perfectly. I'm 99% sure Thai has all of the sounds made in English, but even knowing that it still blows my mind to this day.
You’ve added to it, as custom fonts wasn’t one covered.
I think it’s this thread: https://news.ycombinator.com/item?id=18567548
Edit: and it’s there, #11.
The Kangxi dictionary (1716), an authoritative dictionary of Chinese characters, contains definitions for 47035 characters, even though only a couple thousand are in common use. Quoting from Wikipedia: "The dictionary was the largest of the traditional dictionaries, containing 47,035 characters. Some 40% of them are graphic variants, however, while others are dead, archaic, or found only once. Fewer than a quarter of the characters it contains are now in common use."
All of these archaic (or even bogus in some cases) characters found in the dictionary are now part of the Unicode standard, of course :) The unihan database even has a field that shows the page number where the character appears in the Kangxi dictionary. If you're wondering why 65536 characters isn't enough for everyone, the junk in Kangxi dictionary is a significant contribution.
Anyone can invent characters whenever they want, and it's only a question of them sticking or not.
I think this is also one of the reasons for the Chinese tendency to push for unification and uniformity.
Sometimes it’s mashing two previously unrelated ‘words’ together (aka the tons of compound characters in Chinese), other times it’s coming up with something completely new.
Same rules apply though, if it doesn’t add value worth the trouble (or get mandated by the powers that be), it’ll eventually just die out or be a curiosity.
Also, to keep it tech related:
RISC = English CISC/VLIW = Chinese?
More complex characters require more education to understand is my guess. Some of the traditional ones are….. obscure, and crazy complex.
A good historical example is all the strangly specific words for groups of animals. A history I read of this indicated these terms were first found in books sold to nobility, and they were just made up. But you weren't hip if you weren't reading that literature.
One of my favorite features of the language is the fact that every element in the periodic table gets its own character - and the characters have radicals that indicate their usual state of matter (钅 for metal, 石 for non-metal solids, 氵 for liquids, 气 for gases) - see https://en.wikipedia.org/wiki/Chemical_elements_in_East_Asia....
There are also lots of specific words for species of trees (https://en.wikipedia.org/wiki/Radical_75) and other plants (https://en.wikipedia.org/wiki/Radical_140), but characters under those radicals also include botanical terms, medicines, things made of wood and plants, etc.
Yep! For example, the most common Chinese term for "Internet" is 因特网. This is composed of three characters:
互: "mutual"
联: "join", "coupled", "allied"
网: "net" -- carrying both the meaning of a woven net and a computer network
Despite all it’s problems, the fact these two messages ‘just worked’ is really awesome. I heart Unicode, despite all that.
That's not how it works. Most Chinese characters stem from a character C having a pronunciation A referring to a meaning M being used to note another word of meaning M' with same pronunciation A (sometimes slightly different A'). This of course doesn't scale really well, hence the existence of determiners in logographic scripts, which are words used without their pronunciations placed before or after another to give a semantic clue. The innovation of Chinese (which I think is why it's still an efficient script today) was to incorporate the determiner in the character itself to give birth to a character C' where a part refer to the pronunciation and another acts as the determiner, instead of padding the main text with (a lot of) determiners.
I think character proliferation in CJK languages are a result of each word having its own character. The proliferation isn't fundamentally a proliferation of characters, it's a proliferation of words, which happens all the time in all languages. But only in certain languages does this proliferation of words result in additional characters being added to the language.
There is a fundamental difference between pictographic languages where glyphs have intrinsic meaning, and alphabetic languages where letters reflect sounds.
Old English was much better spelled than current English because it didn't have a spelling. People wrote what they heard. The current mess is because we have 5 centuries of bad standards that can render ghoti as fish. I think you'll agree that the digraph ti, as in nation, is just as nonsensical as sh for the same sound and we'd be much better served by a single glyph for both.
We in fact have that already: https://en.wikipedia.org/wiki/International_Phonetic_Alphabe... English uses somewhere around 45 of those sounds depending on accent. Th for example renders two distinct sounds: ð and θ. þ is not any better than th, apart from brevity, because it also rendered to ð or θ when spoken depending on context.
Chinese is of course as much a pictographic language as English is an alphabetic one. A substantial number of glyphs come from combinations of simpler glyphs which have the same sound as the word you're trying to write.
12K characters in common use is equally impressing for me as a non-Asian.
Native speakers can probably recognise at least 3-4k kanji if not more but can probably only write around 2k from memory, depending on how well-read they are.
嘘 (lie) is the best example of an incredibly common word whose kanji form (which is used fairly often) is not in any official government list.
That means that if you know 4800 characters, and you read a text that is 1000 characters (equivalent to around 700 words) long, there's likely one character you won't recognize.
The funny thing is, if you recognize only the top six characters, you already know 10% of the characters in a typical text. The distribution is very top-heavy, but with a long tail that you do have to learn to become literate.
0. https://lingua.mtsu.edu/chinese-computing/statistics/char/li...
By the way, those are the 17th, 125th and 370th most common characters in modern written Chinese.
0. https://zh.m.wikipedia.org/zh/%E5%A4%A7%E4%B8%89%E5%85%83
Considering that the majority of code is written by people who don't know Chinese characters, it would result in never-ending issues, pretty much everywhere.
Korean actually has a two-way system in Unicode. Every conceivable character (= syllable) possible in modern Korean has its own codepoint, which allows most software to display them correctly: from their point of view, it's just another CJK character.
On the other hand, there is a Unicode area containing Korean sub-blocks ("jamo") that were used historically. In theory, you can combine them and get some pretty funky archaic syllables. Almost no software renders them right.
There are somewhat more sophisticated systems which define both the rendering and stroke decomposition of characters (e.g. CDL: http://guide.wenlininstitute.org/wenlin4.3/Character_Descrip...). The general workaround for characters that aren’t on Unicode would be to use one of these stroke description systems to create the character, then render it to an image and insert it.
Currently stroke-based systems are used for calligraphic effect. You could generate new font types, e.g. bold., but controlling the shape of strokes.
Stroke systems are important for teaching character writing because the drawing order is rigorously prescribed. Once you learn the first couple hundred, you can pretty much guess future characters. Wrong order characters often look bad and suggest a non-Chinese speaker mis-copied them. (e.g. some tattoos)
I dread to think what an enormous mess would result if every character was represented as a build-it-yourself instruction manual rather than allowing font authors to correctly represent the characters. This is also ignoring that (depending on the font style), the apparent strokes for a character can change between fonts in the same language (this is because the computer font stroke style and the written font stroke style can be different) -- by putting stroke decisions in the encoding you're introducing a layering violation since fonts should be deciding how characters are styled, not encoding format committees.
Also nobody in China, Japan, nor Korea would switch to an encoding system so incredibly inefficient that more strokes results in more bytes being necessary to store the character (they already compromised with having 3-byte UTF-8 characters when JIS, GB, and Big5 all only required 2 -- and Japan was basically forced to compromise on Han Unification). This would've resulted in the failure of Unicode's mission to be the One True Encoding Format.
It is the distilled essence of the idea that you need to be inclusive of everyone along with a fundamental ignorance of what anyone who isn't American does.
The idea that Chinese characters are glyphs in the same sense of Latin characters can only have come from someone who has never written Chinese.
It is as stupid as demanding a glyph point for each possible integral, e.g. https://quicklatex.com/cache3/8c/ql_9739884527bd893429657272... and https://quicklatex.com/cache3/18/ql_c51509950f58a52253c696a4....
Solutions for English are not solutions for all languages. You can tell because the solutions that were natively invented by people who spoke those languages were _not_ unicode.
Han Unification was in my view problematic but was driven by technical limitations (then again, if Simplified Chinese characters had also been unified I suspect there would've been more pushback to come up with a better solution, but ultimately Japanese was stuck with being the only one making a major compromise on that front).
I don't think a stroke based or combination system would've been better for many reasons: https://news.ycombinator.com/item?id=32102093. And if you don't trust Americans who at least tried to learn about the subject matter, how much do you trust any other programmer (who has no interest in other languages) to be able to handle a more complicated system for representing and rendering 漢字?
But copying how every Chinese dictionary renders characters as trees of simpler characters seems like a much better approach than the arbitrary letter to number mapping of unicode. That it works well in English is only due to the fact it has no accents.
I can confidently say that native solutions for scripts with accents were _not_ unicode like but overstrike. My grandfather has the source code, in Romanian, of a 1960s computer payment system he worked on which had to deal with both Romanian and Hungarian names.
The combinatorial explosion of possible letters and accents made unicode like encodings an obvious non-starter. Historically names could pick up any accent (some times more than one) on any letter. When you have 26 base letters and 6 possible accents you'd need 26 + 26 * 6 (182) unique representations for single accented letters and 26 + 26 * 6 + 26 ^ 2 (1118) for double accented letters. That makes a language which is fundamentally alphabet based unusable on a keyboard. Something that Unicode is still sweeping under the rug.
Then why does every natively-developed encoding system in a 漢字-using country not do it that way? For one thing, how would you handle the fact that 食反=飯 but 食耳=餌 (and the correct rendering depends on the language and locale of the text being rendered)? How about 辵 (aka 辶, ⻍, ⻌)? There are many issues on top of this one, but this is among the most obvious. In the end you would end up with having your encoding format look like Ideographic Description Sequences (which exist in Unicode) but every rendering library would need to have its own lookup table anyway to produce the correct character. Overlaying accents on top of latin characters (in most European languages) is nowhere near as complicated as combining components to form 漢字.
> I can confidently say that native solutions for scripts with accents were _not_ unicode but overstrike.
Unicode supports combining characters for this reason, though there are separate problems with this approach (some characters look almost identical but semantically should be treated differently -- maybe that is something fonts could deal with, but I suspect "Latin Unification" would've gotten more pushback than Han Unification did). If we want computer systems from different languages and cultures to interoperate there are going to be a few rough edges.
No, you don't. Only the most common combinations have their own Unicode number. Most combinations can simply be combined by base and accent ("Mark") numbers. Unicode is not that stupid.
https://en.wikipedia.org/wiki/List_of_Unicode_characters#Lat...
The most common being literally all of them.
Between Latin-1 Supplement, Latin Extended-A, Latin Extended-B and Latin Extended Additional you have some 700 extra characters of which half are some type of accented letter. I only said you'd need 182 for the six most common European accents. Unicode somehow ends up using 300.
The only people who defend unicode are people who have never looked into the spec.
Even typewriters worked with a giant table: https://en.wikipedia.org/wiki/Chinese_typewriter
JIS did this in 1978 for instance.
It does appear that computerizing CJK languages has made them very different from handwriting them; native Chinese speakers now constantly forget how to write hanzi. But they did this to themselves.
This is separate to the question of encoding -- phonetic-based input systems (IMEs) are a far more likely cause (there are less-widely-used shape-based IMEs which still spit out a Unicode codepoint). The same is happening to Japanese natives, though it should be noted that it's not the case that they cannot write 漢字 normally, they just might forget how to write a relatively rare one (just like how you might forget how to spell a word in English because of a dependence on autocorrect and spellcheck).
It gets pretty trippy, pretty quick.
As in "We don't have a clear idea what this rune was for, or what it means, but we see it in documents and so added it to Unicode."
Documents? I had the strong impression that there are no documents written in runes. A rune we only know by its occurrence in documents would be far more interesting for the existence of a document than it would be for its own sake!
Compare what the page about Anglo-Saxon runes says about the corpus:
> The Old English and Old Frisian Runic Inscriptions database project at the Catholic University of Eichstätt-Ingolstadt, Germany aims at collecting the genuine corpus of Old English inscriptions containing more than two runes in its paper edition, while the electronic edition aims at including both genuine and doubtful inscriptions down to single-rune inscriptions.
> The corpus of the paper edition encompasses about one hundred objects (including stone slabs, stone crosses, bones, rings, brooches, weapons, urns, a writing tablet, tweezers, a sun-dial,[clarification needed] comb, bracteates, caskets, a font, dishes, and graffiti). The database includes, in addition, 16 inscriptions containing a single rune, several runic coins, and 8 cases of dubious runic characters (runelike signs, possible Latin characters, weathered characters). Comprising fewer than 200 inscriptions, the corpus is slightly larger than that of Continental Elder Futhark (about 80 inscriptions, c. 400–700), but slightly smaller than that of the Scandinavian Elder Futhark (about 260 inscriptions, c. 200–800).
So across every runic system we know, we have under 600 texts, all of those texts are short inscriptions, and even to reach that number of samples we need to include texts that we aren't even sure contain any runes.
One of the original goals of Unicode was to be able to computerize every document. I still have some old linguistics books in which characters have been handwritten into typed or even typeset text. So these are the types of documents being referred to: academic papers.
Some fancy books have photographs of ancient writing; I’m not sure if Unicode tries to encode such sources and I pretty much doubt it (how would you even know what to call the symbols? You touch on this in your comment). However often they are attached to treatises that order the characters in some way (I.e. index an alphabet) in which case the first case above would apply.
In other words: thanks to some scholars who wrote down and ordered runic alphabets, you can now discuss runes with your friends and colleagues through email.
That's a weird goal for Unicode to have. We've already accomplished that; a PDF file does the job better (note: PDF documents already support every character existing in the past, present, or future!) while being less complex.
Separately, PDF felt like a step backwards on the day it was announced and sadly nothing since then has changed that.
Of course, that sucks, so I've programmed a nearby key to act as l-ctrl+l-shift+u.
Several characters can also be typed with Compose Key.
For characters I use regularly (in my case, generally the elder and younger futharks), I've created a keyboard out of an Elgato StreamDeck XL so I can type any of these runes with a single button press.
> How do you search for *non-unicode* characters in a pdf document
I guess math has a similar representation in unicode as well.
All that said, I think people use runes to express magic and spells (even to this day). I don’t think all the magical runes are expressed in unicode (and perhaps they shouldn’t). If you want to use a rune in that way, you might have to draw it out in SVG or something and then email it to your friends.
It's an ongoing project. As you seem to have guessed, Unicode math symbols are just about as useless for representing math as Unicode music symbols are for representing music. Producing mathematical documents is done using dedicated software, generally LaTeX.
(And what you get is a PDF, because, as I noted in another comment, PDFs already support every notation there is, was, or ever will be.)
I now have a perverse urge to invent a sculptural writing system just so I can break this completely reasonable claim.
If a clay tablet counts, why not a runestone?
One of the biggest problems in the study of these cultures is that they left no written records. We know they had a writing system, the runes, but as far as we can tell they almost never used it for anything. Quite the opposite is true of Mesopotamian cultures, where we're buried in more records than we have the manpower to translate.
There are. Such documents are called runestones and thousands survive to this day, most in Sweden.
> Documents are also distinguished from "realia", which are three-dimensional objects that would otherwise satisfy the definition of "document" because they memorialize or represent thought; documents are considered more as 2-dimensional representations.
I think "realia" - a term I had never heard before - describes runestones better than "document".
This bug will be fixed in Unicode 15.
https://en.m.wikipedia.org/wiki/Runic_(Unicode_block)#cite_n...
https://dl.ndl.go.jp/info:ndljp/pid/1312837?itemId=info%3And...
https://philamuseum.org/collection/object/84871
Googling for Tsukioka Yoshitoshi brings up so much SEO that it is hard to find information in English. If anyone knows anything about it, I'd be appreciative for a pointer about its content/subject!
So, that's a really interesting thought. Perhaps our solution to a permanent reminder of nuclear destruction[1] could be hidden inside a plane of Unicode.
[1] https://en.wikipedia.org/wiki/Long-term_nuclear_waste_warnin...
1. Maintain humanity under 500,000,000 in perpetual balance with nature. 2. Guide reproduction wisely – improving fitness and diversity. 3. Unite humanity with a living new language. 4. Rule passion – faith – tradition – and all things with tempered reason. 5. Protect people and nations with fair laws and just courts. 6. Let all nations rule internally resolving external disputes in a world court. 7. Avoid petty laws and useless officials. 8. Balance personal rights with social duties. 9. Prize truth – beauty – love – seeking harmony with the infinite. 10. Be not a cancer on the Earth – Leave room for nature – Leave room for nature.
source: https://en.wikipedia.org/wiki/Georgia_Guidestones#Inscriptio...
> 1. Maintain humanity under 500,000,000 in perpetual balance with nature.
> 2. Guide reproduction wisely – improving fitness and diversity.
#2 is actually quite unique I think. I mean you could say it's eugenics, but eugenicists are rarely in favor of "diversity".
> This Unicode range is not a place of honor. No highly-esteemed symbol is registered here.
> What was here represented cultural signs that were considered powerful in our time.
https://www.unicode.org/mail-arch/unicode-ml/y2020-m02/0018....
Also if you google search for 彁 one of the results will be this video [!!!!seizure warning!!!!] https://www.youtube.com/watch?v=EsOU0V2kpUI that seems to borrow on the theme of a computer ghost character.
These days, we scroll though the Unicode standard and find rarely used characters that were accidentally added and imbue them with new meaning. (yes, this is seriously a thing)
Or you just make something up. If you’re coining a new character, you probably don’t care about whether the pronunciation is already known.
There's also 奭 https://en.wiktionary.org/wiki/%E5%A5%AD which is occasionally used as a censorship workaround to mock one of Xi Jinping's gaffes in an early 2000's TV interview where he bluffed about being able to carry two hundred "catty" (~100kg)'s worth of wheat on rural mountain roads. The character is composed of two 百 ("hundred") and one 人 ("human/person/people") which is a pitoral euphemism to that line he said on TV. I can't find any sources about this one that's in English so please bear with my half-assed explanation.
The unihan database records pronunciations as well but I don’t think anyone takes it seriously. I’m not aware of any software that “prescribes” specific meaning or pronunciations to characters, so basically if you want you can pick a character and assign it a meaning and convince everyone else to use it. For languages that don’t have a fully standardized writing system , it’s something you can do and people have done it. (There’s also a bunch of people trying to convince people to not use the defacto standardized characters in favor of archaic ones that they claim are more “authentic”… but that’s a story for another day
"In the end only one character had neither a clear source nor any historical precedent: 彁."
my instinct was that this character could be retconned to mean "character whose meaning has been lost", thus creating a self-referential paradox.
Presumably someone would have to then separately come up with a pronunciation for it. Perhaps pronouncing it "duangu" would solve another problem:
https://coconuts.co/hongkong/lifestyle/duang-jackie-chan-ins...
The previous owner had both highlighted and circled the word "spectre" and wrote "ghost?" in the margins. The rest of the text was similarly marked up.
Every time I hear the word "spectre" I see "ghost?" in my mind's eye.
Nothing else in Unicode acts like that. You can't properly parse complex emoji glyphs from a random starting point because you need previous context to know how to interpret following codepoints. With combining characters you just skip ahead to the next non-combiner.
Even a perfect unicode standard wouldn't be able to mitigate the arrogance of a programmer.
Unicode contains even some ancient and long forgotten scripts so historians can keep proper records of them.
Some Japanese characters that aren't real got accidentally built into the unicode code table. This is NOT related to speculative execution attacks at all. It's just "whoa, these kanji don't mean anything, how the hell are they here?!"
Other than that caveat (not security related), this is a fascinating article, especially if you've studied Japanese as a (foreign) language.
edit: ah the character (hammer and sickle) does not show up
The only circumstance I can imagine is where a Latin character has been erroneously encoded with an unused diacritic, for instance a T with a diaeresis.