A Spectre Is Haunting Unicode (2018)
dampfkraft.com
dampfkraft.com
Both those characters made it into Unicode as there was some use in historic script. However even completely absurd variations made its way into unicode. The "many eyed seraphim" is written as серафими многоꙮчитїи. So if you need to write something with a lot of eyes, you can use ꙮ.
https://en.wikipedia.org/wiki/Go_(game)#Notation_and_recordi...
Relevant thread on the Unicode mailing list, with Subject "Purpose of and rationale behind Go Markers U+2686 to U+2689" is here: https://unicode.org/mail-arch/unicode-ml/y2016-m03/thread.ht...
I'm curious about this. Was your post removed for containing certain unicode characters?
There's no way to know one way or the other.
Maybe it'll become in-vogue over the next few years? ;)
Interesting. That "ꙩко" looks phonetically (if that's the right word, I'm not well up on linguistics (if that's the right word again, ha ha)) a bit like the Hindi word "aankh" for "eye". The "n" sound in "aankh" is emphasized less, for lack of a better term. Actually, in Hindi, it is shown as a dot on top of one of the other letters, to show that.
Also reminded by this, via George Borrow's novel Lavengro[1] (a story about gypsies), that the gypsy and Hindi words for "nose" are similar, "nak".
One theory is that the gypsies (Roma(ni)[2]) migrated from northwestern parts of India to other parts of the world, such as North Africa and Europe.
Aankh (pronounced almost like aak) and nak, get it? :)
https://www.etymonline.com/word/*okw-?ref=etymonline_crossre...
Nose is similar:
Much better indicators are the grammar, the pronunciation, etc.
Sometimes an original nasal consonant will reduce to a vowel that remembers the original consonant only by releasing air through the nose. (Where ordinarily the air would come out of the mouth.)
https://en.wikipedia.org/wiki/Nasal_vowel
This is a big thing in Portuguese and French. (At this point, the consonants are long since gone and the nasal vowel is correct Portuguese/French. But the change would have originated in people speaking something closer to Latin, which doesn't use nasal vowels, and being "careless" with their pronunciation.)
Is this what you're talking about?
It's actually a little bit weirder than that; the vowel before -m also disappears. But it's certainly plausible for some nasalization to remain anyway.
> and/or lengthening of the preceding vowel
You're referring to -ns- / -nf-? You're also right there. As far as I'm aware, this doesn't happen for -nd- / -nt-, though.
> Also youtuber ScorpioMartianus has invested some time into training himself into using reconstructed pronunciation and has talked extensively about that
While that sounds like a cool project, I don't think it necessarily has a lot to tell us about the historical pronunciation. I think you could develop a pronunciation system that matched nearly every documented feature of a dead language while failing to match a large number of undocumented features.
You could say the same about pronunciation research published in linguistic journals. Let me use an analogy:
Imagine looking at the source code of a game. It's technically possible for a reader to technically understand what the program is doing and understand what the game is about, how it works, it's rules and goals.
However, if you pass the sources through a compiler (whose behaviour you also can well understand) what you end up with is a game you can run and experience.
Reconstructed pronounciations are a bit like that. You get to "experience" rules that are otherwise coded in an abstract language. The effort of translating those rules into something you experience actually requires a lot of effort and expertise. You can in theory become a "compiler" and learn how to do it yourself (aloud or in your head) but it's hard; what's wrong with outsourcing it?
> Imagine looking at the source code of a game. It's technically possible for a reader to technically understand what the program is doing
This is already well beyond what's possible for a dead language. It's not even possible for living languages, although in that case we can draw empiric conclusions.
I've been interested for a long time in the question of how we can determine how a language divides up the space of possible sounds. For example, English [θ] (the sound at the beginning of "thick") is perceived by Mandarin speakers as being the sound [s] (as in "sick"). It is perceived by Cantonese speakers as being [f] (as in "fickle").
The sounds [s] and [f] are both phonemic in both Mandarin and Cantonese. But something about the phonology of each pushes the sound [θ] into one category or the other. The choice is not arbitrary; it is quite consistent across speakers of each language.
To the best of my knowledge, we have no way to answer the question "how would language X categorize sound Y?" other than experimentation, which is impossible with a dead language. But it is a fact about the language, and in principle the question can be answered solely by looking at the pronunciation of sounds within the language -- in the ordinary course of events, a Chinese speaker would go their entire life without being exposed to the sound [θ], and yet they would largely agree with each other on what the sound was if they did hear it.
I say that this categorization question draws upon rules of pronunciation which we don't presently have a good idea of how to describe or characterize at all.
So I say reenactment of a dead language is an interesting project, but you're inevitably going to make choices that are wildly different from the language as it existed in the past. Pronunciation reconstruction is on much firmer ground -- and it gets there by not addressing most questions. But a reenactment cannot avoid addressing every possibility, and it's going to get most of them wrong.
YMMV. I once watched a short video by an accent coach teaching how to make an Irish accent, a Scottish accent, an Australian accent etc. He talked about place of articulation and made pretty decent (although clearly not native) approximations of the pronounciations. I found his attempts at actively voicing things out quite helpful. I'm fully aware this is just an approximation, but in a way I found that teacher to be more effective at conveying what makes a given accent peculiar, more than what just listening to a native speaker would. Probably it all depends on what you're interested in.
>Is this what you're talking about?
I'm not quite sure, since I don't know much about phonetics / linguistics, as I said above.
Something like this, the French sound from near the top of your link above:
https://upload.wikimedia.org/wikipedia/commons/0/0e/Fr-en.og...
But that nasal sound is not exactly the same as the nasal sound in aankh (at least to my untrained ear).
Can't describe it better than that, sorry.
The one I've stumbled across recently are the "Negative Squared Latin Capital Letter" (1F170 - 1F189 [1]). The 26 letters are there, but "A", "B", "O", and "P" are special. Blood types and a parking symbol. Sure, but why? Why was it decided we needed all of the Latin letters in inverted squares, except those? Why aren't they their own symbol? I'm not complaining here, I just want to know the history.
I'd also love to know the history of the other Latin letter ranges. It feels super odd to me to have, say the "Mathematical Bold Fraktur" range of characters (1D56C - 1D59F [2]). What's the history of including a font in Unicode? Why did we stop with the few that are in there?
This is probably somewhere online, but I can't find it.
[1] https://unicode-search.net/unicode-namesearch.pl?term=negati...
[2] https://unicode-search.net/unicode-namesearch.pl?term=mathem...
Some OSes display some unicode characters as emoji-default. These characters were selected for that for exactly the reason you surmise. There’s a [text presentation selector][] to enforce the text variation. This is also useful for making (yellow triangle with exclamation point inside) become (single color triangle outline with same-color exclamation point inside).
ETA: Ah, OK, here’s a post with pictures demonstrating this: https://stackoverflow.com/questions/48534667/how-to-display-...
[text presentation selector]: http://www.unicode.org/reports/tr51/#def_text_presentation_s...
Those characters have the Emoji_Presentation Unicode property, as listed in https://www.unicode.org/Public/13.0.0/ucd/emoji/emoji-data.t...
If a site wished to exclude garish blobs from comments while permitting textual symbols, that would be the right way to do it.
[0] https://en.wikipedia.org/wiki/Fraktur#After_1941
[1] https://mathoverflow.net/questions/87627/fraktur-symbols-for...
Despite what it may seem like, I'm really not trying to mis-parse the reasons here, I'm honestly trying to figure out where the line is, and why it's there.
Though, as I understand it, the full-width characters are there not for any modern use cases, but for historical reasons having to deal with older character sets.
Still interesting.
A clean grid would be desirable in formal use, but formal use means trying to avoid Latin characters as much as possible. It's generally possible. Plaques and the like are much more likely to say e.g. 二〇二〇年 than to say 2020年.
And I doubt you'd want to eliminate all formal uses of Latin characters. E.g. a plaque about a person would likely want to use their preferred name, which might be in Latin characters.
V A P O R W A V E
AESTHETIC
But, note that the identical process, much earlier, is how we got separate capital and lowercase forms. Writing systems never do that when they're developed.
The screenplay is still being written in letters. Mathematical ℝ is more accurately thought of as an ideogram than a letter. If you were to write "let r be a member of ℝ", the "ℝ" would be structurally parallel to the full word "member", not to the "r" within it.
Courier for screenplays is a choice you make at the document level; blackboard bold for mathematical entities is not. ℝ is always ℝ no matter what styles apply to your document.
In which way are they special? As far as unicode is concerned they are all the same. It's just that some have an emoji rendering variant for the blood type reasons. The rendering can be picked through representation characters: http://www.unicode.org/reports/tr51/#def_text_presentation_s...
I thought it was suggested to render them like that from Unicode somewhere, but it's 100% possible I'm making that part up.
Some implementations are just wrong.
These all have Emoji_Presentation=No as specified in emoji-data.txt, for precisely this reason (to avoid discrepancies in rendering), but most platforms don't respect these defaults, as it's common to get strings containing emoji from mobile devices, which usually default to emoji presentation. I talk about this a little in my recent "Text layout is a loose hierarchy of segmentation" blog post.
(And yes, they included not only 𝔸𝔹ℂ𝔻𝔼𝔽𝔾ℍ𝕀𝕁𝕂𝕃𝕄ℕ𝕆ℙℚℝ𝕊𝕋𝕌𝕍𝕎𝕏𝕐ℤ, but also 𝕒𝕓𝕔𝕕𝕖𝕗𝕘𝕙𝕚𝕛𝕜𝕝𝕞𝕟𝕠𝕡𝕢𝕣𝕤𝕥𝕦𝕧𝕨𝕩𝕪𝕫 and 𝟘𝟙𝟚𝟛𝟜𝟝𝟞𝟟𝟠𝟡 in Unicode. In iOS Safari, using the font for entering HN comments, ℂℍℙℚ renders a bit taller and bolder for me. ℝ and ℤ render taller, but not bolder. Why beats me.)
I agree with them on ℂ, ℕ, and ℝ. I also think that makes it hard to disagree with them on the Fraktur letters.
This Wikipedia page has a good list of why various ones were added, in addition to those from the Subscripts and Superscripts block: https://en.wikipedia.org/wiki/Unicode_subscripts_and_supersc...
Font fallback. HN specifies Verdana, Geneva, sans-serif. Geneva has the six characters mentioned. The rest will be rendered using a different font that covers more parts of Unicode, or that covers Unicode math specifically. With the built-in system fonts, that would be Cambria Math on Windows and one of the STIX fonts on macOS.
(Source: Firefox → Inspect Element → Fonts tab of the right Inspector pane. Chrome can do that too in a slightly different place, macOS Safari is too user-hostile to have this feature.)
(And, by the way, iOS doesn’t have Geneva, but it has a version of Verdana)
- Complex
- Quaternion (H for Hamilton)
- Natural
- Prime
- Rational (Q for quotient)
- Real
- Zinteger
just kidding. i think Z stands for something in German.
- ℙ used to denote a probability distribution
- 𝔼[X] as the expected value of a random variable X
- 𝟙_S as an indicator function for a set S (which is denoted as a subscript)
- 𝕕 as the differential and integral operator denouncing which variable to integrate over 𝕕f/𝕕x f(x) or ∫ f(x) 𝕕x
I think I've seen 𝕀 and 𝕚 for identity matrices and suchlike, and I'm sure 𝟘 also gets some use.
As to the mathematical uses, this thought is correct. Niche, obviously, but Unicode has no trouble with niche characters.
It's hard to justify the full alphabet on those lines, though.
There's, imo, a reasonable expectation that if some of those characters were used then some other might be later, and it's easier to have them already in the standard than reserving space for them and then having to backfill it, I suppose.
[0] https://en.wikipedia.org/wiki/Blood_type_personality_theory
Last year my Steam name showed properly on MS Win10 and on Kubuntu. About 6 months ago Win 10 started showing it with the O as red. Last week the red O doesn't show at all in one part of Steam but shows white in another part.
I copied the symbols to a Unicode decoder and it showed the O name with a {blood red} modifier, something like that.
I searched briefly but couldn't find the symbol without the red colour, but it only shows red in some places: I'd pasted it into the search bar in Firefox (on Kubuntu), it showed as a number-square (1Fxxx) but then after searching it showed as the red-O character.
Utf8 is weird.
Your issue here is with the characters and their representation, not with the specific encoding. Hence, what you wanted to say is: "Unicode is weird" ;)
o = "\U0001F17E"
black = o + "\ufe0e"
red = o + "\ufe0f"
print(black, red)
HN seems to strip the special sequences, here it is on a pastebin instead: https://paste.debian.net/1169485/More details: http://unicode.org/faq/vs.html http://www.unicode.org/Public/emoji/5.0/emoji-variation-sequ...
IIRC, it also came from a Japanese code page and was probably used for baseball scores, although the exact origin and usage remains a bit of a mystery.
Definitely the legend of this area is Korean[1]. Almost 95% (or even more) of characters have never been used and no one know how to pronounce it. Even from the first line of the blocks, I could find weird characters like 갅, 갌, 갍, 갎.
This is absolutely false. KS X 1001, the primary character set for Hangul, contains 2,350 out of 11,172 (modern) syllables, which is definitely much larger than 5%. And that's not enough (KS X 1001 itself was heavily criticized for this), there are multiple secondary character sets got into Unicode; Unicode 1.1 had 6,656 arbitrarily ordered characters that were finally replaced by 11,172 neatly ordered characters by 2.0. They are real characters people use, my very old personal research [1] indicates that at least a half of them has legitimately appeared in the chat log for example.
Pronunciation-wise, multiple consonant clusters in the final position indicate conditional pronunciation. "갅", for example, comprises of ㄱ g + ㅏ ah + ㄴ n + ㅈ j which is normally pronounced "gahn" by its own but "gahn-j" when followed by a vowel (e.g. "갅이" is pronounced "간지" gahn-ji). Phonetically Korean only has 7 possible codas and other final consonants exist because Korean orthography is a compromise between phoneticism and ideographicism and some words had to be adjusted to show their morphological roots.
[1] https://j.mearie.org/post/24348147729/hangeul-usage-in-irc-c... (in Korean)
There was an album put out last year named 彁, which is one of the more well-known ghost kanji: http://www.rd-sounds.com/c97.html
That's a similar problem as Prince's name during the time he had a symbol as name for which the workaround was just to call it "symbol".
More broadly speaking, there are some true facts cited in this post, but the basic premise is horseshit, and the whole of it should be retracted.
As Andrew West has pointed out, these characters were not "made up" as they are real characters attested in Chinese texts. From a Unicode perspective, none of them is a "ghost character" at all.
[1] https://ja.wikipedia.org/wiki/%E5%B9%BD%E9%9C%8A%E6%96%87%E5...
What a language.
This is the exact kind of place lost glyphs would find a home. The UNIHAN effort (which spans tons of sources) is very dedicated to cataloging these efforts.
I run a few projects based around UNIHAN, which is geared toward cataloging glyphs that are variants, many of which are archaic and no longer used.
Write up about UNIHAN: https://unihan-etl.git-pull.com/unihan.html
I have a tool to export Unicode's database of CJK characters to CSV, JSON, and so on: https://unihan-etl.git-pull.com/
Python library: https://cihai.git-pull.com/
CLI tool (compare to cjklib [1]): https://cihai-cli.git-pull.com/
No knife / character decomposition. I'm working on it but the data is GPL'd: https://github.com/cjkvi/cjkvi-ids/issues/65
("ou" without an accent means "or", sql would be funny in French : SELECTIONNE * DE matable OÙ type='A' OU type='B' )
https://en.wikipedia.org/wiki/Superscripts_and_Subscripts_(U...
Anybody know the reason?
https://www.unicode.org/L2/L2009/09028-n3571-upa-additions.p...
Meanwhile, in a parallel universe where ASCII was invented in asia and latin characters were not available until unicode was established:
> For example, Ä was an error introduced while trying to record A below a dotted line. When reading the copy, the dotted line was added to the character by mistake. The original character (A) was not added to Unicode until much later and doesn't display on most sites for me.
These unused glyphs on the post are harmless, because unused.
Did Burmese typewriters contain an upside-down character, which subsequently became proper typewriter style?
https://skeptics.stackexchange.com/questions/49653/did-burme...
The function that swaps the arguments to a two-argument function: f ↦ ((x, y) ↦ f(y, x))
Alternatively, the f(x,y) = (y,x) notation would work (sort of, you tend to explicitly write out the base vector 𝑒ₖ for both notations: f(x,y) = f(x𝑒₁+y𝑒₂) = y𝑒₁+x𝑒₂ )
https://en.wikipedia.org/wiki/Function_(mathematics)#Arrow_n...
swap = f ↦ ((x, y) ↦ f(y, x))
The rule this satisfies is
swap(f)(x, y) = f(y, x)
and the type for swap is
swap : (X × Y → Z) → (Y × X → Z)
for some types/sets X, Y, and Z.
(Also, note that tuples might not be anything like a vector space. For example, for String × [Int], you wouldn't usually write ("hi", [2,3]) as "hi"𝑒₁ + [2,3]𝑒₂ since there's not really a good commutative addition operation for strings or lists.)
0: https://math.stackexchange.com/questions/64468/why-is-lambda...
I assume that the situation is even stranger in Japanese because these mistakenly created symbols do not have either an associated definition (even a mistaken one) or a pronunciation.
(Now we see how many people get the joke.)
When I posted, the post included this character:
U+262D
https://www.fileformat.info/info/unicode/char/262d/index.htm
Unicode Character 'HAMMER AND SICKLE' (U+262D)
Yep, it gets eaten by HN. It's in the BMP, dammit! It should be safe!
It's in Unicode because it's a hieroglyph. Now why is it a hieroglyph I think you would need to ask ancient Egyptians.
Probably the most interesting political symbol in Unicode is the Emblem of Iran: https://en.wikipedia.org/wiki/Emblem_of_Iran
It's referred to officially as "FARSI SYMBOL", which apparently was a euphemism chosen because during ISO standardization of Unicode the original name "SYMBOL OF IRAN" was deemed unacceptable. As a logo, it wouldn't make it into the standard today, but nobody knows how it originally got in there:
http://archives.miloush.net/michkap/archive/2005/01/29/36320...
The symbolism of the fasces suggested strength through unity: a single rod is easily broken, while the bundle is difficult to break."
So just look for a farming-related item with 2 or more stick-like objects bound together.