Unicode character “ꙮ” (U+A66E) is being updated
twitter.com
twitter.com
Today I learn that Centiclops effectively has a Unicode character. As Centiclops' representative in the world of the non-imaginary, we accept that a Unicode character with a hundred eyes is not practical and we accept the representation with just a few eyes, but generally agree that upgrading to 7 to 10 is a nice improvement, as 7 does not evenly divide into 100 but 10 does. This is important, because... reasons.
"It is true that I never leave my house, but it is also true that its doors (whose numbers are infinite) (footnote: The original says fourteen, but there is ample reason to infer that, as used by Asterion, this numeral stands for infinite.) are open day and night to men and to animals as well."
https://klasrum.weebly.com/uploads/9/0/9/1/9091667/the_house...
https://www.mythicalcreaturescatalogue.com/post/2016/06/10/t...
Fixed, thanks.
https://www.history.com/.amp/topics/christmas/santa-claus#si...
[1] Source: https://boston1775.blogspot.com/2016/12/st-claus-was-celebra...
Well, santa is a Spanish word meaning "holy" and saint is a cognate French word meaning the same thing. They descend from Latin sanctus; compare sanctify.
When the prayer goes "holy Mary, mother of god", "holy Mary" is an exact equivalent of "santa María".
[1] https://en.m.wikipedia.org/wiki/Hail_Mary
[2] https://glaemscrafu.jrrvf.com/english/avemaria.html
[3] https://hymnary.org/text/hail_mary_full_of_grace_the_lord_is...
[4] http://www.marysrosaries.com/Rosary_prayers_in_different_lan...
But it makes sense that Church Latin would be different.
Like oinops, referring to the “wine-like face” (ie dark read) of a sea passage in the evening.
May be wrong, though, I don’t have my tools handy to check now.
Though to be honest, that Unicode character looks more like a bunch of cells forming a tissue to me than eyes.
Not in any normal sense of "roots". Cent is a Latin root meaning 100. ops is a Greek form meaning eye. The -i- indicates that the word is being formed in Latin, and the -cl- is entirely spurious. The original Greek word divides as cycl-ops, not cy-clops.
https://en.wikipedia.org/wiki/Hybrid_word mentions automobile, chloroform, hexadecimal, micro-instruction, petroleum, television and a few others.
Truly the purity of English will never recover from the trauma that I have inflicted on it.
Sure, it does, in English, which stole prefixes, suffixes, and roots from Latin, Greek, and many other languages, and has no problem using them together, without special concern about where it got them from.
The real Greek for a hundred-eyed being would be something like "hekatonoptes", but Argus wasn't called that as far as I know.
ꙮ ꙮ
ꙮ ꙮ ꙮ
ꙮ ꙮ
ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮMake a new character. Updating the existing character ruins the meaning of all previous usages.
It's like trying to change an API. Don't disrespect your existing users. Make a new version.
(ꙮ ͜ʖꙮ)
Think of all the ASCII art this botches. That has to have some historical importance to the Unicode standards body.
(⌐ꙮ_ꙮ)
For scholarly digital (unprinted) documents where the correct character rendering matters, erroneous past usages can be trivially found with grep, a date search, and easily corrected. The domain experts will familiarize themselves with this issue and fix the problem. Don't take a shotgun to it!
This message wꙮn't have the ꙮriginally intended meaning if the characters are updated from underneath.
ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ ꙮ
You may disagree with me, and that's fine, but it doesn't change Unicode's mission. Besides, there's room for 1,112,064 codepoints[a], and only 149,146 are in use. It's predicted we'll never use it up, so what harm is there in one codepoint no one will ever need?
[a]: U+10'FFFF max; it used to be U+FFFF'FFFF, but UTF-16 and surrogates ruined that
It was introduced with other "ocular O"s which are seemingly more commonly used than this one.
It's not quite an illuminated initial.
"This document requests the addition of a number of Cyrillic characters to be added to the UCS. It also requests clarification in the Unicode Standard of four existing characters. This is a large proposal. While all of the characters are either Cyrillic characters (plus a couple which are used with the Cyrillic script), they are used by different communities. Some are used for non-Slavic minority languages and others are used for early Slavic philology and linguistics, while others are used in more recent ecclesiastical contexts. We considered the possibility of dividing the proposal into several proposals, but since this proposal involves changes to glyphs in the main Cyrillic block, adds a character to the main Cyrillic block, adds 16 characters to the Cyrillic Supplement block, adds 10 characters to the new Cyrillic Extended-A block currently under ballot, creates two entirely new Cyrillic blocks with 55 and 26 characters respectively, as well as adding two characters to the Supplementary Punctuation block, it seemed best for reviewers to keep everything together in one document.
(...)
MONOCULAR O Ꙩꙩ, BINOCULAR O Ꙫꙫ, DOUBLE MONOCULAR O Ꙭꙭ, and MULTIOCULAR O ꙮ are used in words which are based on the root for ‘eye’. The first is used when the wordform is singular, as ꙩкꙩ; the second and third are used in the root for ‘eye’ when the wordform is dual, as ꙫчи, ꙭчи; and the last in the epithet ‘many-eyed’ as in серафими многоꙮчитїй ‘many-eyed seraphim’. It has no upper-case form. See Figures 34, 41, 42, 55."
〜 ( ꒪ ꒳ ꒪ ) 〜
( ༎ຶ ෴ ༎ຶ )
I also wonder how these didn't become insanely popular overnight, like the famous eggplant.
I see that near it, there is an ef (Ф) with a very tall stem.
Why should that not be included as a standard unicode character? Surely it is used more often than the multiocular o.
You may say "it's a decorative flourish", which is of course true, but so is the multiocular o. Should we allow every conceivable decorative flourish into unicode? What is the standard for where flourishes become distinct characters?
Only for characters from existing coded character sets.
This doesn't contradict the stated goal exactly, but it seems against the spirit of it at least.
I mean, this ship has long sailed, but that was a mistake nevertheless. Not everything has to be a unicode character.
For example I could see there being an emoji for keyboard cat.
However, I could totally see some kind of open source GIF library of a few hundred meme videos and pictures, to standardize the "Reply with a GIF" thing in some P2P chat ecosystem, and maybe it could have a new URL scheme for referring to OpenMemes images.
So a peach emoji is not the same thing as the iOS peach-emoji-image. Similar to how changing my font doesn't change the actual characters.
I don't think including emojis was a great idea, but now that it's happened and people everywhere use them, emoji have become characters. I agree with your point, but it's already happened and so now there's not really any going back.
Don't tell me Presidents of the United States' song "Peaches" was about butts!?
https://en.wikipedia.org/wiki/Peaches_(The_Presidents_of_the...
The song was definitely a vaginal reference as far as I ever knew.
Personally I think there should be, actually. There's all these other body parts but these are left out. Emoji is almost becoming a language and the good thing is that everyone can understand them, regardless of language. For example I could imagine these could be very useful in an international medical setting. Or for sexting, obviously, we can pretend that's not a thing but that's a bit too Victorian for me.
Of course they're not appropriate in some settings but so are many words.
ꙮ seems to just be a fancy way of writing О. I haven't seen anything that says it has a different meaning. The arguments for excluding Klingon seem to apply even more so to ꙮ.
A weird solitary character from the 1400s isn’t subject to that, and even if it’s a mistake it’s probably not worth breaking compatibility at this point (I think the last such break with code points genuinely changing meanings was to repair a mistaken CJK unification some time in the 00s, and the Consortium may even have tied its own hands in that regard with the ever-more-strict stability policies).
Similarly, for example, old ISO keyboard symbols (the ⌫ for erase backwards, but also a ton of virtually unused ones) were thrown in indiscriminately at the beginning of the project when attempting to cover every existing encoding, but when the ISO decided to extend the repertoire they were told to kindly provide examples of running-text (not iconic) usage in a non-member-body-controlled publication. (Crickets. The ISO keyboard input model itself only vaguely corresponds to how input methods for QWERTY-adjacent keyboards work in existing systems—as an attempt at rationalization, it seems to mostly be a failed one.)
> I think the last such break with code points genuinely changing meanings was to repair a mistaken CJK unification some time in the 00s, and the Consortium may even have tied its own hands in that regard with the ever-more-strict stability policies[.]
Not exactly, the last break happened between Unicode 1.1 and 2.0 and the new CJK Unified Ideographs Extension A block still contains unified characters. The main reason for break was that both Hangul and CJK(V) ideographs required tons of additional code points and it became clear that 16-bit code space is dangerously insufficient; by 1.1 there was only a single big block of unassigned code points from U+A000 to U+E7FF (18,432 total), and there were 4,516 and 6,582 new Hangul and CJK(V) ideographs in 2.0 (11,098 total).
“A linguist has revealed he talked only in Klingon to his son for the first three years of his life to find out if he could learn to speak the 'language'.
[…]
Now 13, Speers' son does not speak Klingon at all.”
I can’t speak of other letters that were added in the same batch in 2007. Some of them seam meaningful, I donno, I don’t speak old church slavonic (although I am told it sounds like Croatian, which I understand a little)
Han unification only means that when you convert Japanese encodings (B) to Unicode (A), it is not distinguishable from non-Japanese encodings converted to Unicode. This means that the Unicode text doesn't always follow domestic conventions without out-of-band signaling or IVD or so. But if you know that the text was converted from a particular encoding, you can perfectly recover the original text encoded in that encoding.
Fonts bloat (do you want a font with 1 million characters in it ? I don’t. Do you want to have to install 1000 fonts having 1000 characters each to be sure to cover all the Unicode table ? I don't).
Lots of issues for everyday programmers (how do you handle weird unicode characters in your validation code ?) potentially leading to security issues (bypassing validation rules by close-but-different characters, phishing…)
There could have been a good case not to include it back in 2007, but once it has been included, excluding it would break stuff.
Speaking of which, do we have any similar hexagonal symbol ?
One known and surviving use. It is possible that it exists in other places, since the vast majority of the planet's written work has not been digitized. It may also have been used other places that have not survived.
Just because it's not important to you does not mean it is not important.
The fact that is survived for 600 years makes it interesting and worth saving. It is infinitely unlikely that anything you do, write, or say will last that long.
Ouch
Of course, the simple answer is that Unicode actually includes any character that someone cares enough to ask to be added, with rare exceptions.
If they’d treat the word characters the same way, it would only serve to confuse and do no favors to the remaining glyphs.
We do NOT keep discovering new planets, rather minor planets (I agree that the term is confusing), more than a million of them discovered in the Solar System now, like the 9007 James Bond.
When I think of a planet, I think of a world that has active geology that isn’t a moon (I know excluding moons is arbitrary, and perhaps I shouldn’t do that; but hey, that’s language for you). I honestly don’t care about the orbit, and I bet that when most people think about planets they aren’t thinking about the orbit either, let alone whether the planet has cleared the orbit or not. I doubt that will change.
Wouldn't that definition rule out gas giants?
I don’t think anybody considers geological activity as particularly useful for classifying things as ‘planet’ or ‘not planet’.
But there is still active geology on Mars. There is still moisture, winds and ice-caps that are shaping the environment. I consider that to be geologically active.
EDIT: And there are actual experts which consider active geology (or something similar) to be a planet, including Anton Petrov (https://www.youtube.com/watch?v=8-2HxrgqUnM)
The 'dwarf planet' distinction helps solve this! There are planets - distinctive in that they have clear orbits - and there are dwarf planets, which can be part of belt systems. This is a useful distinction.
I think it is fine that there are more objects planets then we can meaningfully count. Loads of things in our language act like that. E.g. a bug can be any number of things, and you know what a bug is by just talking about it. If some insect society then comes up with a meaningful definition of bugs which excludes spiders, that definition isn’t really doing the average user of that word any favor.
[1] http://www.asahi.com/special/kotoba/archive2015/moji/2011082...
I'm not sure we have space for another glyph in Unicode. Looks pretty packed in here...
Make a new character!
[1] https://lib-fond.ru/lib-rgb/304-i/f-304i-308/#image-251 [2] https://web.archive.org/web/20110927102700/https://www.stsl....
It seems to be preceded by other jocular glyphs among scribes. A regular o received a central dot when writing the word "eye": ꙩ
Words containing "two" or "both" had an o replaced by two conjoined os: ꚙ
It is only natural to carry on the in-joke for the dual and plural "eyes" like this: ꙭ
Our scribe simply got a little excited, following the pattern to it's logical end point in the term "many-eyed"
I'm glad I'm not responsible for unicode. Clearly I have the wrong mindset for it.
There's no way to know how many since-written documents will break if a whole codepoint is dropped.
It certainly made sense to include this package in Unicode, and the vast majority of those characters certainly should be in this proposal. You do have to draw the line somewhere, and obviously those close to the line will be debatable, no matter where you chose to draw it, like this particular symbol - but once you've decided that you will include the one-eyed O (small and capital) and the two-eyed O (small and capital), then putting in the many-eyed O as well to complete the set doesn't seem so far-fetched.
I just don't have the personal fortitude to attempt something so grandiose. Seems like a fool's errand.
Also, keep in mind there's not just one multiocular O. There's a bunch with varying numbers of eyes.
There are not quite 8000 spoken languages on Earth at the moment, and a lot of them are from cultures that never invented writing. SIL has sent a missionary to most of them to learn the language, invent a writing system for it, teach it to them, and translate the New Testament into it. Most of those are fairly standard alphabets using characters from the Latin scripts, plus perhaps a few new characters or new combinations of character and diacritical. The task is large, but finite.
But something just doesn’t feel right when you’ve got unicode with a character with one known use from forever ago.
Doesn’t this open up the flood gates to just a ridiculous amount of work or else biased gatekeeping?
How much work would it be to implement your own font of the entire unicode set? Or is that not actually a thing and fonts implement as-desired subsets?
You can't, and you are not expected to do so. You are limited by OpenType limit (65,535 glyphs), various shaping rules that possibly increase the number of required glyphs, and lack of local or historical typographic convention. Your best bet is either to recruit a large number of experts (e.g. Google Noto fonts) or to significantly sacrifice quality (e.g. GNU Unifont).
But yes, time constraints are the limiting factor. I don't think anyone is going to dedicate their entire life to making a single font.
Actually this character seems like a scribe's joke, no different from the illustrated characters at the beginning of medieval paragraphs (all of which are represented in Unicode as A, B or whatever). But the point still holds.
It even holds for modern languages -- consider the ghost characters needed for round trip compatibility: https://weekly-geekly.imtqy.com/articles/418717/index.html
(actually cuneiform is a poor example; perhaps Linear A would have been a better example)
Plus there's no shortage of space in the Unicode address space.
The glyph "ꙮ" was used to refer to an Angel with a whole buncha eyeballs, as one does. In terms of texts that survive today, this specific glyph has exactly one use in a single manuscript from the 1400's. It might have been used more, in texts which don't survive. But it is part of a larger trend, and I bet that its inclusion in Unicode depends strongly on that.
But yeah, in itself the ꙮ character exists solely so that modern computers are capable of a more-faithful rendition of the transcription of a single handwritten copy of the Book of Psalms.
"In the center, around the throne, were four living creatures, and they were covered with eyes, in front and in back. ... Each of the four living creatures had six wings and was covered with eyes all around, even under its wings."
I wonder if there is even a copy of the book transcribed to actual characters or if it only exists as scanned PDF copies? If anyone did transcribe it, would they have any knowledge that the ꙮ character even exists on computers?
In Unicode's case I think most of them are paid, at least.
However that is a bit harder with emojis, that have their own subcommittee, which seem to be more bureaucratic and also more popular than the rest of Unicode. Everyone wants to make a new emoji.
But the multiocular O really does seem like one monk got bored one time and did some doodling.
[1] https://twitter.com/BabelStone/status/1323440365429542919
[0]: https://www.unicode.org/wg2/docs/n5170-multiocular-o.pdf
https://news.ycombinator.com/item?id=32095502 ("A Spectre Is Haunting Unicode", 180 comments)
edit to add: The top thread in the 2020 repost was about ꙮ,
Hordes of Wizards of the Coast lawyers getting ready for the big fight
(When an apiarist is terrified.)
ꙮ>
===
The James Webb Space Telescope.I fear this will lead to a lot of "bug fixes and performance improvements" in Android. /s
> written in an extinct language, Old Church Slavonic
It’s absolutely not extinct and is used by the Eastern Orthodox Church in their religious texts almost exclusively. It’s taught to children alongside their Sunday school curriculum and, of course, in seminaries.
So then we all end up paying this massive complexity tax everywhere to pay for support for some Mongolian script that died out 200 years ago (or multi codepoint encodings of simple things like é - just why, it was so avoidable).
That said, what it's trying to do is enormously complex.
It would be nice if we could come up with some magical system that optimally encodes all the text that "matters" and ignores everything else, but history has shown that to be very hard. So we're left with Unicode, which takes the approach of giving us (effectively) infinite code points to represent characters, with (effectively) infinite ways to visually represent them. That does lead to a bunch of "unnecessary" baggage and headaches, but it also solves a bunch of real problems that you probably don't know exist.
Unicode is a pain in the ass, but it's a solution to a very hard problem. You can feel free to design your own solution, but you'll probably run head-first into all the problems Unicode was trying to solve from 40 years ago.
P.S.: Also, even for those, it would seem that one of the big reasons for things like combining characters was added to Unicode in order to be backwards compatible even with mutually incompatible encodings ?
This is not true. For a concrete example: the languages Hindi and Marathi, with ~500 million speakers, use the Devanagari script (also used by Nepali and Sanskrit), in which a grapheme cluster is (usually) a sequence of consonants followed by a vowel. For instance, something like "bhuktvā" (भुक्त्वा) would be two grapheme clusters, one (भु) for "bhu" and one (क्त्वा) for "ktvā". In Unicode each vowel and consonant (here, bh, u, k, t, v, ā) is separately encoded, which is the only reasonable thing to do, and inevitably means that grapheme clusters can have different lengths (number of code points). The alternative would have been to encode every possible (sequence of consonants + vowel) as a single codepoint, which gets ridiculous quickly: these sequences can be up to 5 consonants long, so you'd end up having to encode (33^5 * 13 ≈ 500M) codepoints for Devanagari alone (or completely prevent certain sequences of consonants from being expressed, which makes no sense either), not to mention that most of the scripts of the Indian subcontinent and south-east Asia follow the same principle and have similar issues (e.g. Bengali with 250M speakers, Telugu, Javanese, Punjabi, Kannada, Gujarati, Thai with over 50M speakers each, etc).
(See chapters 12–17 of the Unicode standard, currently version 15: https://www.unicode.org/versions/Unicode15.0.0/ch12.pdf)
[1] https://apple.stackexchange.com/questions/278937/is-there-a-... [2] https://forums.macrumors.com/threads/updating-maverickss-emo...
[1]: https://mufi.info
Hmmm. Well, on my screen, it’s lips/kiss. A Unicode fail.
How this stuff make it to Unicode?!
There was an accepted proposal to add many windings and webdings letters as unicode endpoints. Thus, levitating man in a suit.
There are things like the "ghost characters," which are codepoints in Japanese that map to characters that were basically transcription errors when the team was putting together a full set of Kanji. Some characters with an extra horizontal line snuck into the set; they were likely caused by a transcription error because the character got split onto two pieces of paper by lines of text being copy-pasted into a records book, and the shadow cast by the thin extra layer of paper was misinterpreted as another stroke.
I think I have a solution to decentralize Unicode:
1. Extend Unicode to 128-bits. We can still use UTF-8 variable length encoding which will limit the real size.
2. Use a blockchain to coordinate the characters. That way whoever wants to add a character can do it without gatekeeping.
These simple suggestions will go a long way in making Unicode less centralized.