Emojis paved the way for UTF-8 everywhere
developers.ibexa.co
developers.ibexa.co
There are 16 of these planes. This first block of 65,536 characters is what you can encode with only two bytes (e.g. UTF-16), and it includes most of what anyone alive needs to encode their languages adequately enough. For a long while anything encoded beyond this block had only limited support, and plenty of bugs and limitations meant that using it was tricky (well, it worked fine in LaTeX of course; via xelatex for example). This was back in 2008/2009.
Characters encoded beyond the BMP in plane 1 and 2 included things like esoteric CJKV additions (East Asian ideographs) not usually in daily use, but part of historic documents.
Then came the emoji additions (a core set is part of the BMP and came from Japanese telecom standards), and support is now ubiquitous. Using UTF-8 is a no-brainer for most applications, and a good things that is too!
UTF-16 allows representing characters outside of the BMP by using a reserved area to split a single codepoint into two surrogates that form a pair.
This makes UTF-16 complicated and in some ways worse than UTF-8: the encoding is longer for many typical texts, but is still not fixed-width. The bug you typically see is that codepoints outside of the BMP are munged when clipping the text to a certain length (or reversing it, but that doesn't happen in real systems generally.)
Yes, that was rather my point: if you're using a Unicode-based character encoding, you're going to have variable-width characters regardless, so you might as well use UTF-8.
> UCS-2 uses a fixed number of (16-bit) code units to represent a Unicode scalar value (code point).
Sure, but that's a implementaion detail of the mapping from characters (at the application level) to bytes (at the physical(-ish) representation level).
In other words, to what extent are surrogate pairs a UTF-16 thing, rather than a Unicode thing that exists to accommodate for UCS-2 -> UTF-16?
> The definition of UTF-8 prohibits encoding character numbers between U+D800 and U+DFFF, which are reserved for use with the UTF-16 encoding form (as surrogate pairs) and do not directly represent characters.
They could be replaced by the replacement character to produce a valid string.
https://simonsapin.github.io/wtf-8/
To be clear, this is not an official Unicode spec. It's a hack (albeit a pretty natural and obvious one) to deal with systems that don't do Unicode quite right.
I recently came across some old code that narrows wchar_t to UCS-2 by zeroing out the high-order bytes. Even though my test was careful not to generate any surrogates in the input, they showed up in the output when a randomly generated code point like U+1DF7C was mangled into U+DF7C.
A corrupted value like that is not necessarily a great example of something you want to preserve, but it's the sort of thing that late 90s code assumed about Unicode.
Now, the issue of clipping or reversing strings is a problem not just because of encoding. It simply doesn't work even with UTF-32. You're going to end up cutting off combining characters for example. Manipulating strings is very difficult, and software should never really try to do it unless they know what they are doing, and even then you need to use a library to help you do it.
Not sure what you mean by 'being smart' when all of those were released before Unicode 2.0.
That said, UTF-8 was already 4 years old by the time Java came out. Surrogate pairs was added to Unicode in 1996, one year prior to the release of Java.
I joined Sun Microsystems around that time, and Unicode really wasn't a thing in the Solaris world for a few more years, so the fact that people wasn't aggressively pushing good Unicode support at the time is understandable. People just didn't have much experience with it.
The problem with UTF-8 is that the density is really good for North America and Western Europe but drops off quite a bit for other languages, and you have to trade CPU for bandwidth (eg, gzip) to do much about it.
Japan has several encodings (though shiftJIS is the only one that I can recall) that use escape characters to switch code pages. As long as you don't switch too rapidly between kanji and borrow words, it's more compact, but more complex to implement (I would say less so than implementing gzip but if you aren't using zlib, one of the most portable libraries in existence, you have much bigger issues than character encoding).
UTF-8 takes 3 bytes for all of the first block. Only the first 2048 characters fit into 2 bytes, which is mostly European languages.
Taking a random Wikipedia page as sample I get 46kB (UTF-8) versus 35kB (Shift-JIS). A random Japanese text from Project Gutenberg is roughly ⅔ of the size of the UTF-8 text in Shift-JIS.
Those are impressive enough numbers, but add just a single photograph to the Wikipedia page and it doesn't matter at all. Text is just pretty efficient, even if you use an encoding that supports every language in the world.
Second, text is so comparatively tiny relative to photos, video, code, etc. that it really doesn't matter at all anyways.
Third, text is often zipped as well. It's often zipped over HTTP. It's zipped when it sits inside of an EPUB. It's zipped when it sits inside a Word document. You can even configure MySQL to zip text fields in a database. Basically, whenever space is an issue, you can fix it.
So it's hard to see how this is any problem in practice at all, when phones and computers mostly ship with 32 GB of SSD minimum.
Shift JIS requires you to track extra context when you're actively using the text. That's going to take extra space. I bet that in most of the situations where shift JIS meaningfully wins out, you could get more benefit by using a combination of UTF-8 and Zstd.
But whoever came up with this cute analogy got the analogy wrong — the higher Unicode planes are analogous to the "outer planes" themselves; while the "astral plane" would be some sort of glue allowing you to access these outer planes from within the BMP. Like... surrogate-pair characters! One could nickname the reserved surrogate-pair range in the BMP, the "astral projection" range ;)
Early discussion of "astral character" or "astral plane" for the Unicode supplementary planes at: https://unicode.org/mail-arch/unicode-ml/Archives-Old/UML024... Even earlier 1998 use: https://www.unicode.org/L2/L1998/98354.pdf
UTF-16 is variable width, not two bytes, and it can encode any Unicode character.
This is also a little confusing, since "block" already has a specific meaning in relation to Unicode.
Unfortunately, this has been false for a long time. BMP turned out to be not even barely enough even by Unicode 3.0 [1] where the initial set of Unicode emoji (722 characters) would barely fit in the undesignated area of BMP. Many important characters, starting from a larger set like CJKV and eventually to almost everything by Unicode 6.0 [2], got allocated in SMP and SIP instead as a result. HKSCS additions in the CJK Unified Ideograph Extension B block (U+20000..2A6FF) are notable examples.
I was using some of these (from B and probably C) for very specific purposes at that time, and general support was a long way off in 2009 (although already good on GNU/Linux distributions).
https://en.wikipedia.org/wiki/UTF-8#/media/File:Utf8webgrowt...
I was expecting some actual insightful anecdote about a pivotal choice by a major tech company that made all the difference. But nothing at all.
UTF-8 was the sane refactoring after the initial incompatible parallel standards were established.
I work with WordPress a lot, and up to 3 years ago it was quite common for MySQL setups at shared hosting providers to only support utf8mb3. And Emoji support really did help here to move it forward.
The more interesting thing is why basically no one uses Eastern ideograms in the West, except maybe the Korean ideogram for crying (ㅠㅠ) and rarely, other kaomoji-like stuff. Some kanji also tell visual stories, and most children learn them just fine, so it’s not as simple as accessibility. Borrowing kanji was also anticipated by many sci fi writers and yet is not to be.
The character I can think of that kind of matches the description is 囧, which is a Chinese character.
Yu. Having two of them looks like a crying face. Although tha (ಥ) is also a common component of crying face (ಥ_ಥ). They're talking about kaomoji which use various non-latin or fullwidth symbols (though you're right that they're largely not ideograms) to compose pretty extensive "smileys" e.g. the look of disapproval uses kannada, denko uses greek and katakana, ...
The Gboard keyboard on android has a tab for many of these common “emoticon” faces / character sequences. If you open the emoji picker on the keyboard and then tap the far right bottom tab icon “:-)”
They can get very elaborate though, these are just very basic common faces.
iOS also has that on the standard Japanese "Kana" keyboard (and possibly others), under the "^_^" key.
Some programming languages also start adopting them, too. Raku is the one I know (it allows French and German quotes, too). Maybe Julia, too? I think some language communities tend to be more open to widespread Unicode usage in source code than others.
Once upon a time, I wanted to rely on these distinctions in a TTS frontend to distinguish between 5" floppy disk and "Mambo No. 5"
I soon realized that people use quotation marks and dashes in such a random manner that insisting on treating the semantics literally would create more confusion than it would resolve.
Edit: Now that my comment has posted, I realized you used the correct characters, but the font on my browser rendered them in a seemingly incorrect way, so they looked like double primes.
https://jakubmarian.com/map-of-quotation-marks-in-european-l...
I like to call «this» the Swiss system, because in Switzerland they use it for four different official languages.
That would really be handy on the CLI instead of doing a bunch of escaping with backslashes.
I've had to run double-depth for loops which then execute an awk script remotely. Three would have been fine.
As far as I know this only works for characters in the BMP, but for such characters it may be faster than picking them from that cmd-ctrl-space popup.
If you don't have that input source available, you can add it via the "Input Sources" tab of Keyboard preferences. There you can also enable showing the input menu in the menu bar giving you an easy way to switch between "Unicode Hex Input" and your normal input source.
I will bite, wtf has Chrome to do with UTF-8? As far as I can find the last browser to struggle with it was IE5, IE6 was released almost a decade before Chrome was a thing.
There is no requirement whatsoever that the browser actually use 16-bit code units to represent strings. This is what Simon Sapin achieved for Servo with the WTF-8 encoding, which extends UTF-8 to allow surrogates, including unmatched surrogates: it makes it able to represent these strings with 8-bit code units, commonly halving memory usage and improving the speed of various sorts of linear processing (though at the cost of random access speed which becomes O(n)).
¯\_(ツ)_/¯
I don't know!
(Yeah - I know - it's not an ideogram)
That’s false. There is a lot of students of CJK languages in European universities as Japanese and more recently Korean as popular language. Chinese and Japanese can be learn in some high-school too. There are publishers publishing books containing such characters (e.g. You Feng in Paris). And of course the diaspora and heritage learners... That a lot of people, even more in countries with significant East-Asian population like Australia or the US.
"update now to get access to :burrito: and :taco:! (also fix the following 12 CVEs that 90% of our userbase doesn't know about or read)"
"add a glyph to the emoji keyboard" is more precise.
Burrito: Taco:
edit No, it stripped those :'(.
https://forums.macrumors.com/threads/updating-maverickss-emo...
In any case, this must be the most ridiculous carrot yet. :D
I’ve seen sentences sometimes where words are actually replaced with emojis... is this how some subset of people actually communicate online or that’s just for some effect of irony?
My use of emoji in messaging applications is primarily limited to quick rebus-like reaction replies to meme images, a quick and dirty reaction to a message in slack expressing some vague emotional response, or making complicated fart-jokes with my partner that rely on a lot of out of band information.
(i also use IRC on the daily, and have been known to use emoji there too, so these communication forms are not disjoint)
Text is a flattening of speech, and emojis can add some of those missing dimensions back -- and, like our IRL verbal cues, tics, and gestures, they can be hard to decode if you're not "in" on the game.
This article does that. Here it might be more often than the author would normally do, but I've seen things like that non-ironically.
> To stay relevant in the age of social media you had to support emoji or you were in the .
EDIT: HN stripped the emoji. The end of that sentence was "or you were :skull: in the :droplet:", read as "or you were dead in the water".
Maybe not ironically, but doing that definitely adds informality - exactly the kind of out-of-band context that GP is talking about.
It’s an ingroup signal, indicating (by “correctly” using emoji) that you’re part of a specific subculture to the recipient. “Praying_emojii flame_emojii YASS” indicates to me that the writer is young and hip. “Do you want to have a BBQ bbq_emojii ?” says to me that the writer is older and less hip.
It can be hard to communicate tone through writing. Emojis allow one to instantly mark a piece of writing as informal/non-serious with minimal effort. This includes irony - “eggplant_emojii” is often a non-serious reply indicating that “I am jokingly acting like this this is sexual or attractive”
It’s a proxy for longer writing. “Thumbsup_emoji” is a substitute for some marginally harder to articulate feelings of “looks good / I like it”
Of course, there are many subcultures that use emoji in different ways and as proxies for other things as well. At a previous employer we’d often just send “taco_emoji?” to ask who was buying lunch. It’s the sort of thing that can be used/abused in many different ways.
Emoji are just fonts, and with some search engine sleuthing, you can find an OG pistol emoji out there, pop open your system emoji font in an editor, and replace the water pistol with a pistol pistol.
Everyone else will still see the Nerfed version, of course.
†For some value of easy
It has a lot of follow on implications too.
It is known that developers from English-speaking countries are generally oblivious to encoding problems. Probably they could get by on ASCII far longer than the rest of the world, so no wonder that they might confuse cause and effect in this case.
Or treated UTF8 as the nonsense that MySQL's utf8 is (it's 3-bytes utf8 aka only the BMP, and silently drops anything from the first non-BMP codepoint).
Edit: I didn't realize HN doesn't support emojis
They may not care about i18n, but they do care about cute emoticons.
So we get not just unicode support everywhere, but also character pickers inside the keyboard.
And now we also get to benefit from unicode for things we find pretty - for example, the famous powerline https://github.com/powerline/powerline
Information density matters, and I can't wait for someone to replace "old" color coding of files (.dircolors in batch) by 1 emoticon : a music note for music files, etc.
I'm not really too sure why. I don't mind them in personal messages or texts and stuff, but seeing them on public pages just kind of annoys me for some reason.
I see it on public repositories sometimes, but it never really seems to add anything useful.
I'm used to them at this point, and it's kind of nice when scanning commits to be able to see what type they are.
Don't get me wrong, i'm not going to get mad or lose my mind or anything when I see an emoji somewhere, it just I dunno it looks wrong or something.
I'm not because it's completely arbitrary about it e.g. you can include 🁓, 🀕, ⏱, 🃅, box drawing, or Z̸̠̽a̷͍̟̱͔͛͘̚ĺ̸͎̌̄̌g̷͓͈̗̓͌̏̉o̴̢̺̹̕ but not trigrams, die faces, box elements, musical notes or flags. They just whitelisted/blacklisted entire blocks and called it a day.
Which obviously is par for the course when it comes to HN's comment box, the markup system is even more half-assed.
(The SI standard for separating the integer part from the fractional part is to use "." or ",", whichever is customary in your location. Using thin space for grouping removes the ambiguity that you get in places that use one of "."/"," for grouping and the other for a decimal point).
Let's see what it does with `U+00BC Vulgar Fraction One Quarter Unicode Character`... ¼
Nope it allows it instead of turning into `1/4`, so that's not canonical normalization. I guess it's custom rules? Or some other unicode transformation we're not thinking of, or other third-party re-usable transformation.
They just blacklisted (or whitelisted) blocks or categories.
Country flags seem like they could be used for political trolling.
Die faces could lead to weird rolling threads or other things.
Musical notes, you got me, can't really think of anything too bad for those.
The markup's not great, but too much formatting is distracting. I personally prefer the limited options. You focus more on the content of your comment than making it look pretty.
The only thing i really despise about hn's formatting is the code blocks or whatever they are, the one on mobile that vanishes off the side and you have to scroll horizontally to read everything. I really can't stand when people use those for quotes.
Other than that though, hn's formatting makes everything uniform and fairly easy to read through. There's no fancy nonsense getting in the way of things.
Actually, that's part of why those code block things piss me off, they're probably the fanciest piece of formatting you can do and all it does is obstruct information and make me waste time while reading.
That’s not really believable given how arbitrary it is.
> Die faces could lead to weird rolling threads or other things.
As if tiles or playing cards could not be used that way.
> The markup's not great, but too much formatting is distracting.
The problem is that despite having only two directives half of HN’s markup is actively detrimental: because there is no escaping, no inline literals, and the parsing is sub-par, in my experience the “emphasis” directive causes issues more often than it helps. HN’s markup would be significantly improved by removing it entirely.
> I really can't stand when people use those for quotes
Which would be way less likely if HN actually supported quotes.
>Which would be way less likely if HN actually supported quotes.
But look how well this works ;p.
Sorry...couldn't resist.
I dunno, I like the 'hackish' nature of it.
You're right i'm sure the tiles or playing cards could be used like that too, it may be arbitrary, I don't know. But, those were just some reasons off the top of my head, i'm sure when HN was being programmed a bit more thought went into it, or maybe not, who knows?
My main point is, I like the simplicity of it all, sure it could be better, but better doesn't necessarily lead to better quality content.
There's a minimum amount of distractions, most users find reasonable ways to communicate the context of the content of their posts and scrolling through most threads tends to be a mostly uniform experience where if users are following a few established conventions, you can follow the flow of things pretty well.
It's not perfect, it's not the best, but I feel like it fits the general vibe and nature of the site. It gives HN an identity among all the other news aggregators and forums.
No, that's exactly what they did. I asked.
Let's see if they work here:
Edit: No, they were filtered out.
I can't be the only one who thinks emoji are a terrible idea. Granted, I also don't think logographic characters are a good idea but at least they have thousands of years of use and agreed upon semantic meaning behind them.
Language is just an encoding for ideas, and emojis are a new compression algorithm. Using a single character, you can now convey certain thoughts and sentiments that you previously needed many more characters to reference or explain.
"So then just explain it!" some might respond. "Why can't people be bothered to spend even a little time to write down what they think?" It's an accessibility issue. People have limited time every day to get their ideas across, and they deserve ways of conveying their ideas concisely. There is precedent for this too -- this is why we have acronyms and new words. "lol" and "minivan" don't have thousands of years of agreed-upon semantic meaning behind them.
A final thought -- whether you think emojis are a terrible idea might not be relevant to whether they should exist. Letting people live their own lives to the fullest is much more important than making sure you, I, a future historian, or any other third party understands what they are saying. But you don't have to worry about not being to understand conversations. Given what you prefer, if someone wants to address you as a target audience, then they probably won't use emojis.
This makes me curious. What things can people talk about with emojis that they can't talk about in a proper language?
But you just did. So what does the emoji add to communication? And how does one get the secret decoder ring for it?
There are plenty of emojis that mean different things in different cultures. Some quite amusing.
— Every human being in recorded history when they hit 35 years old.
I highly doubt this. Common shortcurts (such as "it's" for "it is") and dialect is still absent from serious journals (unless it's the research target, of course) and I'm rather sure emojis will similarly be disregarded.
OMG, that makes so much sense. I was the opposite, grumbling about not caring about that silliness, not realizing the psychology of things.
Having suffered through the dark ages with Microsoft increasingly ruining the world with an endless stream of proprietary crap (I still hate them for making people think tab width is configurable), it's amazing to step back and witness how much things have improved (on this narrow slice).
0: https://kitemaker.co, an awesome new issue tracker with tons of hotkeys (and emojis)
(And by the way, HN stripped out the emoji I put in this comment; perhaps that is for the better, but it's kind of funny that the software we use here is quite opposite to the software we make in our day jobs)
If anything, I would say the UTF-8 paved the way for emojis, not the other way around as the ubiquity of a unicode encoding allowed for the existence of emojis. Can't encode emojies with ASCII. You have to have unicode and its encoding first before you can have emojis.