I think we've seen our fair share of derailed projects in IT. Unicode is certainly a very critical project, which makes its focus even more important.
For example, some are phonographic (the letters represent phonemes, i.e. speech sounds), others are logographic (the letters represent words, or morphemes).
Both of these can be featural: The Korean Hangul script is phonographic, but the letter shapes are not arbitrary. Some of them are visualizations of mouth shape or tongue position when producing the sound. For example the velar consonant ㄱ (somewhere inbetween the English g and k, roughly) is the tongue position as seen from the side.
A logogram can also be featural, i.e. it can look like the thing it represents. This would be the case for a glyph that means "thermometer" and looks to viewers like a thermometer. It could also be arbitrary - some shape that you just agree on to mean thermometer. Perhaps sharing some graphical aspects with topically related logograms.
They're all just spots on a spectrum.
(Edit: I briefly touched on pictograms vs. ideograms a bit more down below.)
I realize this message will not render correctly in Chrome because Chrome, at least Desktop Chrome, does not yet support emoji 😭
So, for the sake of interoperability and the ability to read back these documents in a future Unicode-only world, we have to encode them.
When I was a kid, there was a boy in my class who was mute. He carried a chalkboard and flashcards to facilitate communication. Eventually, he was chosen to pilot a device that was sort of like a keyboard with oversized keys. Each key had a pictograph on it[0].
Wouldn't it be nice for him if the flashcards and keyboard device had a consistent set of characters that were easily recognizable not only by him but by the people with whom he was communicating?
Now, would he ever need "MAN IN SUIT LEVITATING"? Probably not (you never know!) but I'm glad there is a standard set of otherwise useful pictographs for people like him.
[0] - I'm in my late 20s but this was when few families owned computers, let alone super-helpful handheld devices. Also worth mentioning that this boy knew ASL but attended a regular public school where few students and faculty could so the chalkboard and flashcards were infinitely more practical.
Btw. that device exists today, and it's called a tablet. I'm pretty sure i have seen stuff exactly like that and it serves a great purpose.
The consistent set of "characters" (pictures in this case) would still look very different on each device or application that uses another font. Because a banana in one font could look very different in another font. The only thing that actually would be consistent if the content he uses would use actual pictures! (A website displaying a banana on his touchpad can look different on his laptop, can't it?).
A fair bunch of Chinese characters are also actually pictographic (they're an image of something) or ideographic (they represent an abstract idea, not a morpheme or word) in origin, if not current usage.
This stuff - synergies and conflicts between writing systems and various media - fascinates me. Chinese writing for example now suffers from the problem that if your medium doesn't allow free-form painting/graphical compositing (like our computers largely don't right now) you can't easily coin new logograms, since you need a central registry that's slow to distribute to leaf nodes (cf. Unicode). So what's happening now is that new words get coined by recombining existing characters based on their sound value, adding a phonographic layer on top. And of course inputting Chinese has also become dependent on auxiliary systems like pinyin-based IMEs.
Meanwhile, the Korean Hangul alphabet has the interesting property that some of the letters graphically derive from each other with consistent patterns of "take this, add a stroke and you get this", making it easier to design keyboards with reduced numbers of buttons to enter combos. There's the theory that this contributed to the success of texting/mobile in Korea, and is a contributing reason for the healthy mobile industry there.
I'm getting far from the original topic with this rambling, but it still reads on it in one sense: There's a lot of variety to writing systems, and it's worth thinking about emoji as a spot on the spectrum instead of in isolation.
Except that emoji aren't a writing system. Writing systems encode language. Emoji encode emotions, perhaps, or suggest collections of ideas/emotions, but don't convey language per se because there's no conventionally defined correspondence between words in a language and a given emoji.
If I give you a string in English or Chinese or Arabic or Amharic, using the conventional writing systems for those languages, each string will be interoperable as a specific sequence of words in that language. A string of emoji doesn't work that way.
Edit: to circle back, you said:
How are emoji different from other logograms?
and everything thereafter was pointless pedantry on my part. Sorry for wasting valuable minutes of your life.Unicode should be used to represent language and not anti-piracy symbols or poo or snowmen. Just my opinion, though. Obviously a great many people like shit in their fonts.
But wait - what's a "normal language" and what's a normal graphical representation of it? Is it down to the development process? A lot of writing systems were originally designed/agreed upon by small groups of people, too, for example.
My questions (and other comments in this subthread) aim to get you to think more abstractly about the problem space and define your beef more accurately, because I'm interested in this question (where should Unicode draw the line) as well. Let's brainstorm, basically.
Fonts are not a clipart gallery and not a place for some committee to dump little pictures. If you want to show me a snowmen over the internet, use an image! Certainly people can understand an image of a snowmen.
Of course these are "official part of a language". But why? Under whose authority? You could cite "actual usage", but many of the emoji in Unicode certainly have seen actual usage for 10+ years in East Asia, so they qualify. And many of the letters in Unicode were in fact designed by committee prior to mass adoption by an actual population (for example the Hangul alphabet), so a lot of scripts already in Unicode originally didn't meet that barrier.
* = Letters vs. characters: Elements of alphabets are referred to as letters. Alphabets are graphical representations of phonemes, but many other writing systems encode syllables (syllabaries, like the Japanese kana), morphemes or words (logographies) instead. Unicode contains examples of all of these, and thus not just letters. So we want to talk about characters instead.
That's pretty much my point.
http://www.unicode.org/cgi-bin/GetUnihanData.pl?codepoint=7C...
You're just being petty or culturally arrogant at this point, TBH.
I say it again: If you want pictures and cliparts, use .png. If you want to write actual text, use fonts. It's as simple as that.
Why should my browser load thousands of pictures that are never used?!
Yes, it's ridiculous.
You're pulling numbers out of your ass. You also clearly don't know how font fallbacks work, despite having been told how they work in this very thread. Several times.
And you clearly don't understand that a) this thread contains unicode that my browser still doesn't show, despite your magic fallback and b) the fonts stil contain garbage that has nothing to do with fonts.
The nice part about having it Unicode is that if you copy and paste that character somewhere else, it'll still work. It never goes through what we used to have to deal where the same character value couldn't reliably be processed without knowing the intended encoding with certainty.
Also, adding political glyphs is a sure way to ruin the universality of Unicode (I'm looking at you no-piracy).
All that said, I love Unicode, it has flaws but I think it's brilliant (no sarcasm).
Also, can't wait to actually use the man-in-business-suit-levitating-glyph (sarcasm).
I'd use the Unicode sarcasm marks, but your font probably doesn't support them yet...
As someone else pointed out, it comes from Wingdings, for which there has been an effort to fully get into Unicode.
https://en.wikipedia.org/wiki/Webdings#mediaviewer/File:Webd...
You'd probably only have one or two installed that have the emoji characters.
It just looks like a big waste for no good reason.
Waste of what?
> Because if not every font implements it, what is the usefulness of it?
Most fonts only implements a fraction of unicode, is unicode a waste? Is cyrillic in unicode a waste? Is Cherokee? Braille symbols? Buginese and Meroitic cursive?
> Did you ever visit websites that use glyphs your font doesn't know?
Yes? The software will fall back on a font which does provide the glyphs, or display a block with the codepoint if it doesn't have any.
> How would you use those pictures on a website, when you can't even guarantee the style of the picture, because it may be implemented differently or not at all.
How would you use letters on a website, when you can't even guarantee the style of the letters, because it may be implemented differently or not at all?
Not only that, but unicode is not specific to websites for fuck's sake, not only will somebody find a use for them on websites even if you can't, somebody will definitely find a use for them outside of websites. Most (or all?) of the new emoji are inherited from windings[0] and webdings[1], they don't come out of nowhere.
> It just looks like a big waste for no good reason.
Again, a big waste of what?
No more so than right now.
> selecting a reasonable subset of bazyllion random, singular pictures to draw and test is a major headache.
Right, because there[0] were[1] no[2] "bazyllion[3] random[4], singular[5] pictures[6] to[7] draw[8] and[9] test[10]" dating back to Unicode 1.0[11] before and Unicode 7.0 changes everything.
[0] http://en.wikipedia.org/wiki/Mathematical_alphanumeric_symbo...
[1] http://en.wikipedia.org/wiki/Mahjong#Unicode
[2] http://en.wikipedia.org/wiki/Dominoes#Dominoes_in_Unicode
[3] http://en.wikipedia.org/wiki/Playing_Cards_(Unicode_block)
[4] http://en.wikipedia.org/wiki/Musical_Symbols_(Unicode_block)
[5] http://en.wikipedia.org/wiki/Arrow_(symbol)#Arrows_in_Unicod...
[6] http://en.wikipedia.org/wiki/Optical_Character_Recognition_(...
[7] http://en.wikipedia.org/wiki/Geometric_Shapes
[8] http://en.wikipedia.org/wiki/Unicode_Dingbats
[9] http://en.wikipedia.org/wiki/Phonetic_Extensions
[10] http://en.wikipedia.org/wiki/Miscellaneous_Technical_(Unicod...
[11] http://en.wikipedia.org/wiki/Box_Drawing_(Unicode_block)
Other comments have touched on this already, but I want to substantiate with additional information: In practice, fonts that cover many different writing systems at once are relatively rare. Your typical Western font will focus on Latin with varying coverage of descendants (like the Portuguese and Vietnamese alphabets) and perhaps expand to cover Cyrillic, but not more. Conversely, it's rare for a font by a Korean foundry to have equivalent Latin/Cyrillic coverage next to its native Hangul alphabet.
There's plenty of documents that mix these writing systems, though. I do in fact visit websites that aren't covered by any one font in my browser's default settings or any one family named in the site's markup/CSS all the time, and you probably do too.
Hence the font stacks on all relevant platforms perform glyph substitution: If font X doesn't contain a glyph, it'll go looking for a font Y that does and drop that glyph in. Many platforms allow you to manipulate substitution preference e.g. by defining aliases for font families.
Which means there's no sanely written font that would use a "t" for an icon. Wingdings is not a sanely written font.