Emoji by Year
taylor.town
taylor.town
Some of them don't display correctly in my browser, so to me it just looked like meaningless garbage. Because it's a personal website, my first assumption was that this was a list of the author's most used emoji every year (it never occurred to me to ask why anyone would track this; it's the internet).
The link at the bottom of the post makes it clearer.
I'm not even sure the exclusion of emojis is consistent with their own explanation [2] of what an identifier should consist of, given that IMO an emoji pretty clearly qualifies as a 'ideograph' (a graphic symbol that represents an idea or concept).
It is reasonable to exclude non-representational characters like punctuations marks, dashes, parentheses, brackets, etc that intuitively don't 'stand for' anything and instead solely serve to delineate syntactic structure. But the vast majority of emojis are entirely unlike punctuations marks and do intuitively 'stand for' some idea or object.
[1] https://unicode.org/reports/tr31/
[2] "The formal syntax provided here captures the general intent that an identifier consists of a string of characters beginning with a letter or an ideograph, and followed by any number of letters, ideographs, digits, or underscores" (https://unicode.org/reports/tr31/#Default_Identifier_Syntax)
Emojis are essentially colored symbols in terms of Unicode general categories and many languages do not accept symbols in an identifier, so it's no surprise that emojis are handled equally.
UTR #31 does have a profile for identifiers with emoji (see the section 7.2), but it also demonstrates why it can't be a default: there are cases where `AB` is not valid but `ABC` is valid for some characters A, B and C. So character classes can't fully describe identifiers with emoji after all! (And in fact, many Unicode character classes were defined only for describing standard algorithms.) You need a full emoji regular expression to define such identifiers.
> [2] "The formal syntax provided here captures the general intent that an identifier consists of a string of characters beginning with a letter or an ideograph, and followed by any number of letters, ideographs, digits, or underscores"
(Emphases original)
Unicode has an unfortunate but particular definition of "ideograph", defined in the third sense of [1]. As a rule of thumb you should always mentally replace it with "Han character", which is not an ideogram (like numerals) but a logogram. This misnomer was too old to fix so Unicode decided to go with it.
They missed an opportunity to just base it on citrus hybridisation; eg a key lime is micrantha X citron!
If I need to express myself without words, I stick to a few emoticons such as :-) ;-) and ¯\_(ツ)_/¯. That's all I really need.
(wit.)
I doubt that emoji are ever used well.
> they can help break up text and draw your attention to specific parts.
I'm not sure what this means?
> More importantly, however, they are much easier for screen readers to interpret, which has to help for accessibility.
Well, I use emoticons sparingly, mainly at the end of messages to specific people who I know don't use a screen reader. Nothing is easier to interpret than words.
Also the existing system is (mostly) backwards compatible. The main issue is that it's not forwards compatible, with the exception of said ZWJ sequences.
You're comparing demand with supply, which are not the same. In any case, the blog post is somewhat misleading, because the 2023 additions produce endless varieties of human "representation", and everyone wants to be represented in emoji. And everyone wants their food to be represented in emoji. Basically, everyone wants every object in the world to be represented in emoji.
It's also misleading to say that most of the new ones don't really require new unicode codepoints, because the new sequences depend on the new codepoints, otherwise they would be preexisting emoji, not 2023 additions.
That's fair, I did get them mixed up. The _supply_ is constrained since they introduced more strict rules for adding new emoji, the demand is still much larger than that.
> the 2023 additions produce endless varieties of human "representation", and everyone wants to be represented in emoji
Most of the 2023 additions are about horizontally flipping actions that have directionality. Only the new family pairings could be considered representation-related, but IMO they're mostly about completing the set of possible combinations.
> the new sequences depend on the new codepoints, otherwise they would be preexisting emoji, not 2023 additions
It's complicated, because Unicode publishes both the codepoints and also a list of recommended interoperable ZWJ sequences. Newer additions have been to this list.
Anyway, ultimately this is a matter of opinion, but arguing that unicode growing to encompass a lot of emoji as time goes by is not something scalable seems weird to me. One of the explicit goals of unicode was the ability to handle CJK ideograms. There are ~90,000 of those, and only ~3000 emoji. Even if they added 100 new emoji per year, which is far more than what they're doing, it would take about 900 years for the emoji set to match the size of the CJK ideograms set. This seems very scalable to me.
Now I just need to figure out how to get 2023 emojis working in Windows 11.