On the other hand we don't use pictures in the same way, the meaning of the message depends on the exact picture used as part of the "font". The whole thing about flags and skin tones are symptoms of the same problem.
On the other hand we don't use pictures in the same way, the meaning of the message depends on the exact picture used as part of the "font". The whole thing about flags and skin tones are symptoms of the same problem.
But anyway, images should have remained separate things, png or whatever. Fonts rendering is complex enough with complex scripts e.g. Arabic, but that is the domain of fonts/shaping/rendering.
Such a hugely expensive technical mess because people want their stupid color emojis.
With, of course, accessibility for those needing it.
For example, the " " emoji, described by Unicode as "health worker" is called "non-binary doctor" by Apple.
If you send that emoji in a text to somebody, and they're e.g. using accessibility software, wearing AirPods and have notifications read out to them by Siri, or have their phone integrated with CarPlay etc, the message they'll hear is very different from what you actually intended to say.
The point of any glyph at all being allocated a Unicode codepoint, is that the glyph already exists in documents somewhere (either real paper documents, or digital documents in some other legacy encoding); and if Unicode didn't include a codepoint to encode that glyph, then that text wouldn't be able to be faithfully represented in Unicode.
Which means that some ephemeral undocumented proprietary encoding might be used instead (as in “cellphone novels” in late-90s Japan, now permanently committed to the Internet in the form of SJIS encoding + some particular feature-phone’s SJIS pictograph codepage that no modern OS can decode.)
Or worse, that the character and its contribution to the meaning of the document might be lost entirely (as in MSN or ICQ messenger emoticons, where exporting chat logs as HTML would just strip the emoticons out.)
Putting whatever glyphs already exist into Unicode, no matter how rarely, is precisely what makes Unicode uni-code — i.e. an encoding that guarantees lossless archival "transmission into the future" of all texts, such that all texts should be better off transcoded into Unicode. (At least if the intent is to archive their semantic meaning as text, rather than their bytewise representation as data for software-archaeology purposes.)
---
The standard of "if it already has usage, then it must get a Unicode codepoint so that the document can be Unicode-encoded" is all well and good at face value, and worked well as a guideline for the first ~10 versions of the Unicode spec. Until ~2009, pretty much nobody was arguing against codepoint allocations being "warranted."
But then the modern mobile platforms realized that they essentially had a gun they could hold to Unicode’s head — in the sense that any glyph they wanted to invent, they could threaten to release into the wild as proprietary encoding if Unicode doesn’t include it. And because millions of people would be guaranteed to make tweets et al using that new symbol, Unicode is essentially forced to include it. Even though it doesn’t exist yet until that moment.
All the “dumb” emoji—and color emoji in general, and the explosion in variation selectors—came after that. In est, they’re Apple and Google’s fault, not the Unicode Consortium’s.
> Which means that some ephemeral undocumented proprietary encoding might be used instead
I'm not arguing against the need of some standard for representing those, I'm just objecting to the technical decision of incorporating them as code points, because they don't serve the same semantical purpose.
Instead you could have introduced -- for example -- a simple vector graphics format.
The standard says “pistol”, what a pistol looks like is left to the implementer.
You're missing the point. Unicode is descriptive like a dictionary is descriptive: it describes — and thereby locks into place and somewhat formalizes (really: ossifies) — existing usage.
If you prefer, rather than "prescriptive" vs "descriptive", you can pretend I said "proactive" vs "reactive." Unicode as a standard reacts to existing out-of-Unicode usage of glyphs, by defining into existence Unicode code-points to allow the encoding of those glyphs. (Or rather, to allow encoding of any of the equivalence-class-hypersphere of glyph-configuration-space described by the name of the codepoint, as that codepoint. Those particular codepoints have been defined through usage, whether the Unicode Consortium likes it or not.)
But even that's not quite accurate, because, especially with the modern forced assignments by Apple/Google, Unicode in fact is purely describing — documenting — codepoint assignments that the OEMs decided on and implemented unilaterally, ahead of standardization. (I.e., Apple/Google now just allocate new codepoints for emoji and assign them glyphs all on their own — and Unicode then must play catch-up, including those codepoints in the standard only after they're already shipping on real devices and thereby effectively already "locked in" in their meanings.)
> Because they don't serve the same semantical purpose. Instead you could have introduced -- for example -- a simple vector graphics format.
They quite clearly do serve a semantic purpose:
• Screen readers can describe an emoji — and smart ones can even use an emoji at the end of a sentence to add emotional color to their reading of the sentence. This would not be possible if emoji were just vector images.
• LLMs have particular trained associations on what a given arbitrary Unicode codepoint should relate to. It would be prohibitively difficult for them to form the same associations between regular word tokens, and the huge sequence of tokens that collectively represents a vector image.
• Search engines can index emoji just like any other text. No search engine that I know of could embed a vector image in its fulltext index in a useful way.
• If a font doesn't represent an emoji, people can still copy-and-paste the unrepresentable-codepoint-glyph rendering of the codepoint into their OS's character map to get a Unicode-standardized description of the codepoint. This wouldn't be true if emoji were just arbitrary vector images.
Also, if you wanted to oust some glyphs from Unicode in favor of embedded vector images, where would you stop? If U+1F60A "Smiling Face with Smiling Eyes" shouldn't be in Unicode, should U+2660 "Black Spade Suit"? How about U+21F6 "Three Rightwards Arrows", or U+2713 "Check Mark"? U+2766 "Floral Heart" / U+2042 "Asterism"?
And here the issue is not in font differences (or different pictures of the same thing getting different interpretations): it’s the thing represented that’s actually different.
Many typefaces are have problematic I/l/1 and O/0 characters, and that is material enough that applications avoid those characters entirely to avoid potential confusion.
The problem is deeper than unicode: the fact that any codepoints can be rendered differently from client to client or typeface-to-typeface was always problematic, even back it was mostly ASCII.