Unicode 15.0 Slide Show
babelstone.co.uk
babelstone.co.uk
Too late now — it’s become an annually refreshed collection of fun fashionable clip art instead of an impartial repository of humankind’s symbols.
A tiny portion consists of "fun fashionable clip art." The vast majority of the changes are new scripts and symbols that allow people to correctly input existing texts.
/s
tbh my bigger worry is simply around implementation scope and complexity - kinda like browsers, Unicode and text rendering is perpetually getting harder to create new competitors / complete fonts / etc. Emoji are rather small in number and simple to implement though, if somewhat costly to include all those images / complex svgs.
Wikipedia's view of what's 'original research' has alway been a bit murky, on top of that. Are the primary documents that historians rely on acceptable as sources, or only if they've been synthesized by an 'accredited historian' (whatever that is) in book or research journal form? If publishing original research is wrong, why publish a journalist's original research which is incorporated into a newspaper article? What about original research published as a blog post, is that now an acceptable secondary source?
It isn’t as easy as you make it sound - to establish notability, they don’t accept self-published / vanity press sources - so you’d need to get your book published by a publisher with an established track record. That’s a lot harder - it is either something that money can’t buy, or at least you’d need a lot lot more money to buy it than self-publishing charges
Furthermore, one high quality reliable source is generally not considered enough for notability, they want multiple sources. If you get your biography published, and then get some journalists in established media outlets to publish articles about it, you’ll meet that hurdle too. But if you manage that, you are probably actually are notable, as opposed to just some random nobody trying to buy their way into Wikipedia
Refs: https://www.unicode.org/emoji/proposals.html http://www.unicode.org/pending/proposals.html
Even ZWJ sequences are forever: we're still stuck with the "eye in speech bubble" emoji despite the fact that it's a logo for a defunct anti-bullying campaign that doesn't even have a working website anymore.
Constraints are sometimes good. It's great to see how something like the Eggplant & Peach emoji got hijacked and used as a sexual reference.
Aside from that, I don't think the terms really apply to a discussion about emoji, because what makes emoji odd is that they are almost never in common usage as part of the language until Unicode adds them. That makes it very hard to claim that Unicode is describing anything in particular by adding emoji, and in fact makes their role much more prescriptive than not.
Klingon was explicitly rejected for a few reasons. Most prominently, there were serious concerns about Paramount asserting trademark rights over the characters. Especially at the time of its rejection, it wasn't clear that there was any organic use of the script, even among people speaking Klingon (note that in its earliest uses, it wasn't even used consistently as an alphabet for the language it is supposed to represent). Those reasons don't apply to Tengwar--and it's also been noted since then that the Unicode committee might look more favorably on a newer proposal for Klingon.
I think it does, doesn't it?
Here is some info on unicode rules for new characters, although there doesn't seem to be a single succinct statement, and there are slightly different standards for different categories ("letters" vs symbols vs emojis). But I am pretty sure you can't just make up a character and succesfully submit it to unicode on the argument that it would be useful, this thing you just made up; it has to have already established use to get into unicode, generally.
http://www.unicode.org/pending/proposals.html
Which characters are you annoyed by specifically?
It's true there are separate rules for emoji's, which are... weird. I seem to recall unicode initially tried to resist emoji's at least to some extent, before just giving in.
But emoji's are actually pretty clear, to me, as a success of unicode. People want them, they have achieved pretty widespread use globally, and if that's going to be so, we benefit from them being encoded in a standard vendor-neutral way. We're far better off with emoji's in unicode than, say, every vendor using their own non-standard use of private codepoints. (Unless you argue that if they weren't in unicode they wouldn't have caught on; I am sympathetic to the argument that we'd better off without emoji's, but I'm pretty sure if they weren't in unicode we'd still have them, just non-standardly).
Here are some of the guidelines for which emoji's will be encoded in unicode:
https://unicode.org/emoji/proposals.html#selection_factors
It's a bit fuzzy, but you generally can't simply get a new thing you just invented encoded in unicode on the argument that it's cool.
It feels like there’s this bias where people think they know what’s relevant and what’s just nonsense from the younger generation or what have you.
Mind you, Unicode stresses me out. It’s trying to catalog something that’s impossible to because it’s arbitrary and limitless. And it’s doing it within the arena of software engineering where we really crave rules and clear boundaries.
But if we didn’t try, we’d be worse off. Unicode is in some ways a horror, but it’s valuable, even if flawed.
I work mostly on Linux so I hacked a Character Viewer clone in Rust over a weekend recently[1].
It just does what I need but I'm planning to add features to it if I find them useful.
So I am curious: what functions does BabelMap offer that you can't live without, especially as a typographer?
font-family: Georgia, Serif;
I don't think those fonts support all of Unicode. Google created their Noto fonts [1] for this purpose; I wonder why those aren't being used.
Since 1998 MacOS has a “last resort” font that has glyphs (not necessarily unique) for every Unicode code point. They donated it to Unicode (https://en.wikipedia.org/wiki/Fallback_font#Unicode_Last_Res...), so I expect most OSes running full-blown modern browsers to have it or something similar (those running smaller browser engines may be too space constrained to have room for it)
http://www.chinaknowledge.de/Literature/Science/shuowenjiezi.html
I have the noto fonts and ctext dot org's hana fonts but still see tofu in the above page. Whatever font is used on the iPhone's Pleco app the correctness depends on context where you are in the app.These two examples often are confused: 日曰
I believe there are semi-exceptions to what I said, although you wouldn't want to use such a font for practical purposes. These exceptions are fonts that display the Unicode code point in a box, so they would display the ASCII 'A' as a box containing "0065" (I forget whether they use hex or decimal, or maybe there are both). I suspect these fonts create their glyphs on the fly.
I recognized the domain and tried to remember why, and now I remember.
I'm working on a game, and babelstone.co.uk has probably the world's most comprehensive (and high quality) set of runic fonts:
https://www.babelstone.co.uk/Fonts/
https://www.unicode.org/reports/tr9/tr9-46.html
because I speak a right-to-left language. Whoever wants to write an application involving text entry, and truly support localization or internationalization, should take the time to read at least section 3:
https://www.unicode.org/reports/tr9/tr9-46.html#Basic_Displa...
What should be _the_ value of `"I".lower()`? Or, "i".upper()?
And please don't bring up locales. The whole point of accepting the complexity of Unicode is to be able to take a document which stands on its own without external references.
> Early character encodings also conflicted with one another. That is, two encodings could use the same number for two different characters, or use different numbers for the same character.
> The Unicode Standard provides a unique number for every character, no matter what platform, device, application or language.[1]
Those statements are outright lies: Unicode does not provide a unique number fpr "upper case Turkish dotless i". Nor does it provide one for "lower case Turkish dotted i".
If it did, it would be possible to correctly map "i" to "I" or "İ" and "I" to "i" or "ı" without having to know anything other than the source codepoint.
The font does not even come into play here.
Even more complicated, there is a capital form of ß:
https://en.wikipedia.org/wiki/%C3%9F
> Until 2017, there was no official capital form of ⟨ß⟩; a capital form was nevertheless frequently used in advertising and government bureaucratic documents.[10]: 211 In June of that year, the Council for German Orthography officially adopted a rule that ⟨ẞ⟩ would be an option for capitalizing ⟨ß⟩ besides the previous capitalization as ⟨SS⟩ (i.e., variants STRASSE and STRAẞE would be accepted as equally valid).[11] [12] Prior to this time, it was recommended to render ⟨ß⟩ as ⟨SS⟩ in allcaps except when there was ambiguity, in which case it should be rendered as ⟨SZ⟩. The common example for such a case was IN MASZEN (in Maßen "in moderate amounts") vs. IN MASSEN (in Massen "in massive amounts"), where the difference between the spelling in ⟨ß⟩ vs. ⟨ss⟩ could actually reverse the conveyed meaning.[citation needed]
No character encoding standard can save you from that kind of complexity. You need special-case code on a language-by-language basis.
Also, Unicode didn't invent it. Germans invented it.
What makes you think this is the whole point? I don't believe that "whole point" is actually possible with global human languagues as actually used, nor do I think those behind unicode historically or presently have considered this the "whole point".
Do you have any reference suggesting this was meant to be the "whole point" of unicode?
I think it's actually a pretty amazing success that unicode does give us algorithms for manipulating text in various ways, that actually work pretty darn well... but yes, they sometimes require locale parameters.
What it does provide is a unique code point for "upper case I", "upper case dotted I", "lower case dotted i" and "lower case dotless i".
That means "ι".upper().lower() produces "ι" but "ı".upper().lower() produces "i" (unless you do it on a computer set to a Turkish locale:
>>> "ι".upper().lower()
'ι'
>>> "ı".upper().lower()
'i'
Note we have: LATIN CAPITAL LETTER IOTA
LATIN SMALL LETTER IOTA
GREEK CAPITAL LETTER IOTA
GREEK SMALL LETTER IOTA
MODIFIER LETTER SMALL IOTA
TURNED GREEK SMALL LETTER IOTA
CYRILLIC CAPITAL LETTER IOTA
CYRILLIC SMALL LETTER IOTA
APL FUNCTIONAL SYMBOL IOTA
MATHEMATICAL CAPITAL IOTA
MATHEMATICAL BOLD SMALL IOTA
MATHEMATICAL ITALIC CAPITAL IOTA
MATHEMATICAL ITALIC SMALL IOTA
MATHEMATICAL BOLD ITALIC CAPITAL IOTA
MATHEMATICAL BOLD ITALIC SMALL IOTA
MATHEMATICAL SANS-SERIF BOLD CAPITAL IOTA
MATHEMATICAL SANS-SERIF BOLD ITALIC CAPITAL IOTA
MATHEMATICAL SANS-SERIF BOLD ITALIC SMALL IOTA
That's lotta iotas[1]. There's more[2]: GREEK SMALL LETTER IOTA WITH DASIA
GREEK SMALL LETTER IOTA WITH PSILI AND VARIA
GREEK SMALL LETTER IOTA WITH DASIA AND VARIA
...
The point is that almost all of these variations on the theme with explicit duplicates get their upper and lower case distinct codepoints, yet "lower case Turkish dotted i" needs to for some reason map to 0x69 in ASCII and "upper case Turkish dotless i" needs to map to 0x49 ASCII.It is stupid and directly contradicts the statement:
> The Unicode Standard provides a unique number for every character, no matter what platform, device, application or language.[3]
The statement on the Unicode.org web site cannot be attributed to ignorance.
[1]: https://en.wikipedia.org/wiki/Iota
Wow, Unicode seems to be expanding way out of the original goal of representing written characters - how often would these symbols be even used? And how many font sets would bother including them?
Probably the most notable:
> Use figurine algebraic notation, which replaces the letter that stands for a piece by its symbol, e.g. ♘c6 instead of Nc6. This enables the moves to be read independent of language (the letter abbreviations of pieces in algebraic notation vary from language to language).