"Exponential explosion" is really putting it too strong; it's perfectly possible to just add ǿ and á̤ and a bunch of other things. The combinations aren't infinite here.
The problem with e.g. Latin script isn't necessarily that combining characters exist, but that there's two ways to represent many things. That really is just a "mess": use either one system or the other, but not both. Hangul has similar problems.
Devanagari doesn't have any pre-compose characters AFAIK, so that's fine.
That's really the "mess": it's a hodgepodge of different systems, and you can't even know which system to use a lot of the time because it's not organised ("look it up in a large database"), and even taking in to account historical legacy I don't think it really needed to be like this (or is even an unfixable problem today, strictly speaking).
At least they deprecated ligatures like st and fl, although recently I did see ij being used in the wild.
They certainly are. Languages are a creative space driven by the human imagination. Give people enough time and they'll build new combinations for fun or for profit or for research or for trying to capture a spoken word/tone poem in just the right sort of exciting way. You may frown on "Zalgo text" [1] (and it is terrible for accessibility), but it speaks to a creative mood or three.
The growing combinatorial explosion in Unicode's emoji space isn't an accident or something unique to emoji, but a characteristic that emoji are just as much a creative language as everything else Unicode encodes. The biggest difference is that it is a living language with a lot of visible creative work happening in contemporary writing as opposed to a language some monks centuries ago decided was "good enough" and school teachers long ago locked some of the creative tools in the figurative closets to keep their curriculum simpler and their days with fewer headaches.
We've got 150K assigned codepoints assigned, leaving us with 950K unassigned codepoints. There's truly massive amounts of headroom.
To be honest I think this argument is rather too abstract to be of any real use: if it's a theoretical problem that will never occur in reality then all I can say is: <shrug-emoji>.
But like I said: I'm not "against" combining marks, purely in principle it's probably better, I'm mostly against two systems co-existing. In reality it's too late to change the world to decomposed (for Latin, Cyrillic, some others) because most text already is pre-composed, so we should go full-in on pre-composed for those. With our 950k unassigned codepoints we've got space for literally thousands of years to come.
Also this is a problem that's inherent in computers: on paper you can write anything, but computers necessarily restrict that creativity. If I want to propose something like a "%" mark on top of the "e" to indicate, I don't know, something, then I can't do that regardless of whether combining characters are used, never mind entirely new characters or marks. Unicode won't add it until it sees usage, so this gives us a bit of a catch-22 with the only option being mucking about with special fonts that use private-use (hoping it won't conflict with something else).
Unicode can't get rid of the many precombined characters for a huge number of backward compatibility reasons (including compatibility with ancient Mainframe encodings such as EBCDIC which existed before computer fonts had ligature support), but they've certainly done what they can to suggest the "normal" forms in this decade should "prefer" the decomposed combinations.
> If I want to propose something like a "%" mark on top of the "e" to indicate, I don't know, something, then I can't do that regardless of whether combining characters are used
This is where emoji as a living language actually shines a living example: It's certainly possible to encode your mark today as a ZWJ sequence, say «e ZWJ %», though you might want to consider for further disambiguation/intent-marking adding a non-emoji variation selector such as Variation Selector 1 (U+FE00) to mark it as "Basic Latin"-like or "Mathematical Symbol"-like. You can probably get away with prototyping that in a font stack of your choosing using simple ligature tools (no need for private-use encodings). A ZWJ sequence like that in theory doesn't even "need" to ever be standardized in Unicode if you are okay with the visual fallback to something like "e%" in fonts following Unicode standard fallback (and maybe a lot of applications confused by the non-recommended grapheme cluster). That said, because of emoji the process for filing new proposals for "Recommended ZWJ Sequences" is among the simplest Unicode proposals you can make. It's not entirely as Catch-22 on "needs to have seen enough usage in written documents" as some of the other encoding proposals.
Of course, all of that is theory and practice is always weirder and harder than theory. Unicode encoding truly living languages like emoji is a blessing and it does enable language "creativity" that was missing for a couple of decades in Unicode processes and thinking.
Yes, and that only makes things worse since the overwhelming majority of documents (99.something% last time I checked) uses pre-composed. Also AFAIK just about everyone just ignores that recommendation.
This is a classic "reality should adjust to the standard" type of thinking. Previous comments about that: https://news.ycombinator.com/item?id=36984331
I suppose "e ZWJ %" is a bit better than Private Use as it will appear as "e%" if you don't have font support, but the fundamental problem of "won't work unless you spend effort" remains. For a specific niche (math, language study, something else) that's okay, but for "casual" usage: not so much. "Ship font with the document" like PDF and webfonts do is an option, but also has downsides and won't work in a lot of contexts, and still requires extra effort from the author.
I'm not saying it's completely impossible, but certainly harder than it used to be, arguably much harder. I could coin a new word right here and now (although my imagination is failing me to provide a humorous example at this moment) and if people like it, it will see usage. In 1960s HN when we would have exchanged these things over written letters, and it would have been trivial to propose a "e with % on top" too, but now we need to resort to clunky phrases like this (even for typewriters you can manually amend things, if you really wanted to).
Or let me put it this way: something like ‽ would see very little chance of being added to Unicode if it was coined today. Granted, it doesn't see that much use, but I do encounter it in the wild on occasion and some people like it (I personally don't actually, but I don't want to prevent other people from using it).
None of this is Unicode's fault by the way, or at least not directly – this is a generic limitation of computers.
It shouldn't matter what's in the wild in documents. That's why we have normalization algorithms and normalization forms. Unicode was built for the ugly reality of backwards compatibility and that you can't control how people in the past wrote. These precomposed characters largely predate Unicode and were a problem before Unicode. Unicode won in part because it met other encodings where they were rather than where they wished they would be. It made sure that mappings from older encodings could be (mostly) one-to-one with respect to code points in the original. It didn't quite achieve that in some cases, but it did for, say, all of EBCDIC.
Unicode was never in the position to fix the past, they had to live with that.
> This is a classic "reality should adjust to the standard" type of thinking.
Not really. The Unicode standard suggests the normal/canonical forms and very well documented algorithms (including directly in source code in the Unicode committee-maintained/approved ICU libraries) to take everything seen in the wilds of reality and convert them to a normal form. It's not asking reality to adjust to the standard, it is asking developers to adjust to the algorithms for cleanly dealing with the ugly reality.
> Or let me put it this way: something like ‽ would see very little chance of being added to Unicode if it was coined today.
Posted to HN several times has been the well documented proposal process from start to finish (it succeeded) of getting common and somewhat less common power symbols encoded in Unicode. It's a committee process. It certainly takes committee time. But it isn't "impossible" to navigate and is certainly higher than "little chance" if you've got the gumption to document what you want to see encoded and push the proposal through the committee process.
Certainly the Unicode committee picked up a reputation for being hard to work with in the early oughts when the consortium was still fighting the internal battles over UCS-2 being "good enough" and had concerns about opening the "Astral Plane". Now that the astral plane is open and UTF-16 exists, the committee's attitude is considered to be much better, even if its reputation hasn't yet shifted from those bad old days.
> None of this is Unicode's fault by the way, or at least not directly – this is a generic limitation of computers.
Computers do anything we program them to do and in general people find a way regardless of the restrictions and creative limitations that get programmed. I've seen MS Paint drawn symbols embedded in Word documents because the author couldn't find the symbol they needed or it didn't quite exist. It's hard to use such creative problem solving in HN's text boxes, but that from some viewpoints is just as much a creative deficiency in HN's design. It's not an "inherent" problem to computers. When it is a problem they pay us software developers to fix it. (If we need to fix it by writing a proposal to a standards committee such as the Unicode Consortium, that is in our power and one of our rights as developers. Standards don't just bind in one-direction, they also form an agreement of cooperation in the other.)
This comes up in specifications that have a broad range of use cases; when I was involved in this my idea was to just spec things so that there's only one allowed form; you'll still need a small-ish table for this, but that's fine. But that's currently hard because for a few newer Latin-adjacent alphabets some letters cannot be represented without a combining character.
So then you have either the "accept that two things which seem visually similar are not identical" (meh) or "exclude embedded use cases" (meh).
I never really found a good way to unify these use cases. I've seen this come up a few times in various contexts over the years.
> Posted to HN several times has been the well documented proposal process from start to finish (it succeeded) of getting common and somewhat less common power symbols encoded in Unicode.
Would this work for an entirely new symbol I invent today? It's not really the Unicode people that are "difficult" here as such, they just ask for demonstrated usage, which is entirely reasonable, and that's hard to get (or: harder than it was before computers) especially for casual usage. I'm sure that if some country adopts/invents a new script today, as seems to be happening in West-Africa at in recent years, the Unicode people are more than amendable to work with that, but "I just like ‽" is a rather different type of thing.
Sure, they want demonstrated usage as inline in the flow of text as textual elements as opposed to purely iconography or design elements (because such things are outside of Unicode's remit, modulo some old Wingdings encoded for compatibility reasons and the fine line between emoji are expressive text and also emoji are useful for iconography in many cases). But at this point (again in contrast to the UCS-2/no-Astral-plane days) the committees don't seem to care how it was mocked up (do it on a chalkboard, do it in paint, do it in LaTeX drawing commands, whatever gets the point across) or how "casual" or infrequent the usage is, so long as you can state the case for "this is a text element" (not an icon!) used in living creative language expression. There's more "provenance" requirements for dead languages and they'll want some number of academic citations, but for living languages they've come to be flexible (no hard requirements) on the number of examples they need from the wild and where those are sourced from. Showing it in old classic documents/manuals/books, for instance, helps the case greatly, but the committees today no longer seem as limited to just what can be used to demonstrate usage. "I just like it" is obviously not a rock solid proposal/defense to bring to a committee (any committee, really), but that doesn't mean that is impossible for the committee to be swayed by someone making a strong enough "I just like it" case if they demonstrate well enough why they like it and how they use it and how they think other people will use it (and how those uses aren't just iconography/decorative elements but useful in the inline context of textual language).
Even "put the code points forming the composed character in descending numerical order" would be better than nothing. If it was there from the start.
However, the Unicode commitee is too busy adding new emojis to make their standard sane.
But, of course, unicode can't define that the standard will cover only the canonical forms, and couldn't do that since the start, as it needed backwards compatibility with various pre-unicode encodings which had mutually incompatible principles, so it needed support for both composed and decomposed versions of the same characters.
There' your problem right there. Plural formS. It's not canonical if there are more of it.
Not just emojis, in general I believe Unicode has just said they're not going to add new pre-composed characters and that using combining characters is the Right Way™ to do things (well, the only way for newer scripts).
One of the downsides of writing down specifications is that they tend to attract people with Very Strong Opinions on the One And Only Right Way and will argue it to no end, and essentially "win" the argument just by sheer verbosity and persistence.
That's certainly what I've seen happen in a few cases, and is what happens on e.g. Wikipedia as well at times.
But yeah, emojis is even worse. Something things can look rather different depending on which invisible variation selector is present. We've got tons and tons of unassigned codepoints and we need to resort to these tricks to save a few of them?
Firefighter is "(man|woman|person) + ZWJ + firetruck". Clever, I guess. Construction worker is "Construction worker (+ ZWJ + (male sign|female sign))?" (absence is gender-neutral). Why are there 2 systems to encode this? Sigh...
All of this is too clever by a mile.
[1]: HN will strip stuff, but try something like:
echo $'↔\ufe0f ↔\ufe0e'
May not display correctly in terminal, but can xclip it to a browser – screenshot: https://imgur.com/a/iFmBDQkBut the technical implementation? Yeah, that could have gone a lot better IMHO.
One must also wonder if some things really had to be added in the first place, e.g. for people kissing it's:
(person|man|woman)(skin-tone)? ZWJ <heart> ZWJ <kissing lips> ZWJ (person|man|woman)(skin-tone)?
This is NOT a complaint about that they added diversity as such, in principle I'm all for that, it's just that few seem to actually use these emojis, and both in terms of code and UI it all gets pretty complex; there's 98 combinations to choose from here.I don't really get why <heart> or <killing lips> or <kissing face> isn't enough. That's actually what most people seem to use anyway, because who finds it convenient to pick all the correct genders and skin tones from the UI for both people?
Oh. So that's why HN discussion always looks sane. They strip the pollution.
Less than that since a default skin color can be set in most apps. I'm sure setting a gender will come soon so the entire first part of that emoji can be auto-guessed. Then its just showing the other options in the UI. Really all of this is UI design as even with the 98 combinations you can still display it as 4/5 options you drill down.
> who finds it convenient to pick all the correct genders and skin tones from the UI for both people?
I just checked and searching "kissing" in my iOS emoji keyboard inside Messenger showed just 4 of the emoji's your describing - defaulting both skin tones to my settings and then the four M/F pair ups. Plus some non-related kissing emojis like the cat kissing.
But that's kind of wrong, no? The entire point is that you can choose both sides individually. What if you set it to black and want to kiss some white bloke?
If anything that only underscores my point that it's too complex and that no one is using them (certainly not as intended anyway).
In the Windows 11 emoji picker it works like this:
1. Search "kissing". See two generic yellow people kissing. Notice a blue dot in the bottom right corner.
2. Clicking the emoji brings up previously used versions of the kissing emoji, with a + button.
3. Clicking + brings up a dialog like I described previously. Two generic figures at the top, then a row of skin tones.
4. You can click on each generic person and choose a gender, then select a skin tone. You can do this for each person in the group.
5. Click done. This emoji is now in your default emoji list and you won't need to recreate it again.
The problems with Unicode are mostly to do with internal inconsistencies and churn, problems that usually only affect programmers.
1. Different ways to encode the same visually indistinguishable set of characters as code points leading to normal forms, text that compares unequal even when it appears to be identical, the disastrous "grapheme clusters" concept and so on.
2. Many different ways to encode the same sequence of code points as bytes. Not only UTF-32/16/8 but also curiousities like "modified UTF-8".
3. Emoji. A fractal of disasters:
3.a. Updates frequently. Neither Unicode nor software in general was built on the assumption that something as basic as the alphabet changes every year. If you send someone an emoji, can their device draw it? Who knows! In practice this means messaging apps can't rely on the OS system fonts or text handling libraries anymore which is a drastic regression in basic functionality.
3.b. (Ab)uses composition so much it's practically a small programming language, e.g. flags are composed of the two letter country code spelled using special characters. People are represented as as generic person plus skin color patch, families are represented using composed individual people etc.
3.c. Meaning of a character is theoretically specified but can subtly depend on the font used, e.g. people use a fruit emoji in visual puns because of how it looks specifically on Apple devices, so a "sentence" can make no sense if it's rendered with a different font.
3.d. Unbounded in scope. There's no reason the Unicode committee won't just keep adding new pictograms forever.
3.e. Encoded beyond the BMP which in theory every correct program should handle but in practice some don't because nobody except a few academics used characters beyond it much until emoji came along.
3.f. Disagreement over single vs double width chars, can only know this via hard-coded tables, matters for terminals and code editors.
Some of these can potentially be cleaned up outside of the Unicode consortium in backwards compatible ways. You could have a programming language that automatically normalized strings to fully composed form when deserializing from bytes, and then automatically folded semantically identical code points together (this would be a small efficiency win for some languages too). You could campaign to build a consensus around a specific normal form, like how UTF-8 gained consensus as a transfer encoding. You could also define a fork of Unicode (using private use areas?) that allocates a single code point to the characters that are unnecessarily using composition today but don't yet have one and then just subset out the concept of composition entirely.
Emoji are a big problem. It's tempting to say that these should not be encoded as characters at all. Instead there could be a set of code points that define bounds that contain a tiny binary subset of SVG, enough to recreate the Apple pixel art somewhat closely. Emoji would always be transmitted as inlined vector art. Text rendering libraries would call out to a little renderer for each encoded glyph, using a fast fingerprinting algorithm to deduplicate the bytes to an internal notion of a character. To avoid wire bloat, text can simply be compressed with a pre-agreed zstd or Brotli dictionary that contains whatever images happen to be popular in the wild. At a stroke this would avoid backwards compat problems with new emoji, enabling programs working with text to be upgraded once and then never again, eliminate all the ridiculous political committee bike-shedding over what gets added, let apps go back to using system text support and get rid of the bajillion edge cases that emoji have spewed all over the infrastructure.
I've written unicode-aware software for over a decade, doing a wide variety of programs, and I've never had to bother with all that mess.
If I'm parsing strings I'm looking for stuff in the 7-bit ASCII range which maps neatly onto the Unicode representations, and so I just need to take care to preserve the rest.
The only trouble I've had is that a lot of programmers haven't learned, or don't get, that text encoding is a thing and that it needs to be handled.
So they'll hand me an XML they claim is UTF-8 encoded, except that XML header was just copypasta and the actual XML document is encoded in some other system encoding like Windows-1252. Or worse, a mix of both.
For site local, fec0::3. Yeah site-local is discouraged but you can still do it. Or you can slightly misuse fd00::3.
You only get those latter 16 hex characters if you explicitly don't want to choose addresses.
What on earth are you doing that it's leading to crashes? Are you not validating the result?
https://www.reddit.com/r/apple/comments/37e8c1/malicious_tex...
In fact, I don't know that there's any reason to believe normalization happens at all in the process of executing this.
Also, the % should measure people, not languages, that would greatly decrease the imaginary 99%