How UTF-8 Works
sethmlarson.dev
sethmlarson.dev
A sketch on a diner placemat has lead to every person in the world being able to communicate written language digitally using a common software stack. Thanks to Ken Thompson and Rob Pike we have avoided the deeply siloed and incompatible world that code pages, wide chars and other insufficient encoding schemes were guiding us towards.
It's one of those rabbit holes where you can see people whose entire career is wrapped up in incredibly tiny details like what number maps to what symbol - and it can get real political!
Long live UTF-8. Finally I can write any Central European name without mutilating it.
What do you mean by that? Unicode does have double-wide characters and, I discovered recently, some other characters called 'wide' that at least are wider than one column but smaller than two in at least some monospaced fonts. Try:
* Small Hyphen-minus (U+FE63) "﹣": Seems to be >1 and <2 columns in at least some monospaced fonts.
* Fullwidth Hyphen-minus (U+FF0D) "-": ditto
Therefore, by going to utf8 instead of ucs2, we did not have to go to wide chars.
But the whole emoji modifier (e.g. guy + heart + lips + girl = one kissing couple character) thing is a disaster. Too many rules made up on the fly that make building an accurate parser a nightmare. It should have either specified this strictly and consistently as part of the standard, or just left it out for a future standard to implement, and just just used separate codepoints for the combinations that were really necessary.
This complexity is also something that has led to multiple vulnerabilities especially on mobiles.
See here all the combos: https://unicode.org/emoji/charts/full-emoji-modifiers.html
And yeah this problem only came later after UTF-8 was already invented. It's a smart solution also because at the time a lot of detractors of Unicode were opposed to the extra data requirements of the long form representations.
Windows machines in 1990s had several megabytes of main memory, and people could barely get it to support one East Asian language at a time, never mind multiple of them. No sane person would propose using three bytes per a Korean character when two would do - that would mean your word processor will die after adding 50 pages of document, while your competitor can do 75.
And even if you did have UTF-8, you wouldn't see those Thai characters anyway, because who would even have these fonts when your OS must fit in a handful of stacked floppies.
It took years before UTF-8 made technical sense for most users.
Consider also that at this very time we’re talking of (early 1990s) the industry was shifting away from largely 8-bit code pages to 16-bit UCS-2, which is an even more extreme cost when compared to UTF-8, doubling space requirements for most people, rather than merely the 50% increase yongjik speaks of for certain languages. Yet this change was being done (more’s the pity).
Concerning the scarcity of bytes, yongjik’s point would certainly be valid if it referred to the 1970s, was probably valid of the 1980s, but is not valid of the 1990s. (But the point about keeping the full document in RAM is an unrealistic strawman.)
How is that defined and enforced? Very narrowly, it seems to me:
* ASCII Hyphen-minus (U+002D) has similar functions and appearance to Small Hyphen-minus (U+FE63), Fullwidth Hyphen-minus (U+FF0D), Hyphen (U+2010), Minus Sign (U+2212), Heavy Minus Sign (U+2796), En dash (U+2013), Em Dash (U+2014), Small Em Dash (U+FE58), Horizontal Bar (U+2015), Figure dash (U+2012). (I'm probably missing a few!)
* There are separate delta symbols for Greek and for mathematics (sorry, no more time for looking up code points).
* Very many other characters have appearances so similar that nobody could tell them apart.
So characters have apparently, to users and almost anyone not looking at the actual codes, identical functions and appearances.
ASCII is itself valid utf-8, because ASCII is a subset of utf-8. But a multi byte encoded codepoint in UTF-8 cannot be confused with ASCII, because the highest bit is set in all the octets.
There’s 3 valid (and useful) ways to measure a string depending on context:
- Number of Unicode characters (useful in collaborative editing)
- Byte length when encoded (these days usually in utf8)
- and the number of rendered grapheme clusters
All of these measures are identical in ASCII text - which is an endless source of bugs.
Sadly these languages give you a deceptively useless .length property and make you go fishing when you want to make your code correct.
This is also rarely useful unless you are working with a monospace font where all grapheme clusters have the same width, which is probably none if you support double-width characters. More likely what you are interested in is the display length with a particular font or column count with a monospace font.
I would give it to Java outright if not for the fact that C's char type doesn't define how big it is at all, nor whether it is signed. In practice it's probably a byte, but you aren't actually promised that, and even if it is a byte you aren't promised whether this byte is treated as signed or unsigned, that's implementation dependant. Completely useless.
For years I thought char was just pointless, and even today I would still say that a high level language like Java (or Javascript) should not offer a "char" type because the problems you're solving with these languages are so unlikely to make effective use of such a type as to make it far from essential. Just have a string type, and provide methods acting on strings, forget "char". But Rust did show me that a strongly typed systems language might actually have some use for a distinct type here (Rust's char really does only hold the 21-bit Unicode Scalar Values, you can't put arbitrary 32-bit values in it, nor UTF-16's surrogate code points) so I'll give it that.
C’s character type, FWIW, has a use: it more or less indicates the granularity that is efficiently addressable by the host architecture. Trying to use it for more than that is generally not that fruitful, but it definitely has a purpose and it’s pretty good at that.
Finally, speaking of unfortunate decisions, Rust happens to make one that I don’t particularly like: it lets you misalign characters (and panics), which is…not great. It would be much nicer if the view just don’t let you do this unless you specifically asked for bytes or something.
And again, I don't think Java probably needed 'char' at all, it's the sort of low-level implementation detail Java has been trying to escape from so this is a needless self-inflicted wound. I think there's a char in Java for the reason it has both increment operators - C does it and Java wants to look like C so as not to scare the C programmers.
C's unsigned char could just be named "byte" and if signed char must exist, call that "signed byte". The old C standard actually pretends these are characters, which of course today they clearly aren't which is why this is a thread about UTF-8. I don't have any objection to a byte type especially in a low-level language.
Presumably your Rust annoyance is related to things like String::insert? But I don't understand how this problem arises, if you are inserting characters at random positions in a String, that's just going to be nonsense. I can't conceive of a situation where I want to insert characters (or sub-strings) unless I know where they're supposed to go exactly relative to what is in the string already, whereupon it won't panic.
Compare for example decryption, where we learned not to provide decrypt(someBytes) and checkIntegrity(someBytes) even though that's what People often want, it's a bad idea. Instead we provide decrypt(wholeBlock) and you can't call it until you've got a whole block we can do integrity checks on, it fails without releasing bogus plaintext if the block was tampered with. An entire class of stupid bugs becomes impossible.
Java should have provided APIs that work on Strings, and said if you think you care about the things Strings are made up of, either you need a suitable third party API (e.g. text rendering, spelling) or you want bytes because that's how Strings are encoded for transmission over the network or storage on disk. You don't want to treat the string as a series of "characters" because they aren't.
The idea that a String is just a vector of characters is wrong, that's not what it is at all. A very low level language like C, C++ or Rust can be excused for exposing something like that, because it's necessary to the low-level machinery, but almost nobody should be programming that layer.
Imagine if Java insisted on acting as though your Java references were numbers and that it could make sense to add them together. Sure in fact they are pointers, and the pointer is an integral type and so you could mechanically add them together, but that's nonsense, you would never write code that needs to do this in Java.
K&R C claimed that char isn't just for representing "ASCII" (which wasn't at that time set in stone as the encoding you'll be using) but for representing the characters on the system you're programming regardless of whether they're ASCII. 'A' wasn't defined as 65 but as whatever the code happens to be for A on your computer. Presumably the current ISO C doesn't make the same foolish claim.
These requirements are perfectly good for the needs of a CHARacter type. If you need to control signed / unsigned because you want to use the char as a small integer, you can specify yourself whether it is signed or not.
In reality, where chars are used to store ASCII, the signdness of the datatype is meaningless because the highest bit is never set.
The output of files on disk can be UTF-8. The continued use of UCS-2 (later revised to UTF16) is happening in the runtime because things like the Win32 API which C# uses is UCS-2. The internal raw memory of layout of strings in Win32 is UCS-2.
*EDIT to add correction
The biggest problem is when you're working in an ecosystem that uses a different encoding and you're forced to convert back and forth constantly.
I like the way Python 3 does it - every string is Unicode, and you don't know or care what encoding it is using internally in memory. It's only when you read or write to a file that you need to care about encoding, and the default has slowly been converging on UTF-8.
Edit: The proposal for anyone interested https://github.com/dotnet/runtime/issues/933
I remember that I got a home assignment in an interview for a PHP job. The person evaluating my code said I should not have used UTF8, which causes "compatibility problems". At the time, I didn't know better, and I answered that no, it was explicitly created to solve compatibility problems, and that they just didn't understand how to deal with encoding properly.
Needing-less to say, I didn't get the job :)
Same with Python 2 code. So many people, when migrating to Python 3, suddenly though python 3 encoding management was broken, since it was raising so many UnicodeDecodingError.
Only much later people realize the huge number of programs that couldn't deal with non ASCII characters in file paths, html attributes or user names, because they just implicitly assume ASCII. "My code used to work fine", they said. But it worked fine on their machine, set to an english locale, tested only using ascii plain text files on their ascii named directories with their ascii last name.
It's not the type system, but many dynamic languages are interpreted instead of compiled.
(It doesn't even need to be a very good or strongly-enforced type system - Go makes it dangerously easy to convert between `[]byte` and `string` by other-type-system standards, and yet everything works pretty well. It's enough to hitch your thinking and make you realize you need another step.)
It's about the runtime error handling making it mandatory to deal with the error and won't let you panic at run time for this specific error.
That is so far from the truth.
This is quite literally the job of a type system: to impose a semantic interpretation on sequences of "raw bits" and let you specify legal (and only legal) operations in terms of the semantic interpretation rather than the bits.
(And it had nothing to do with UTF-8. Actually, it's at least partially caused by the CPython developers avoiding UTF-8 for poor reasons.)
But even today, codepoint indexing is too easy and doesn't crash so lots of code is still subtly wrong. Memory usage of a string grows 3x if you add a single emoji. Most libraries are still agnostic about whether they take bytes or str so you still get exceptions thrown with no easy solution. (The growing popularity of type hints is fixing this, but that's not really related to Python 3.)
I was delighted recently to stumble on a history of modern women course HIST1158 named Liberté Egalité Beyoncé and I immediately thought two things: 1. Why are our Computer Science courses given unimaginative names? and 2. What a useful test input, I bet some of our systems don't work correctly for this input even though an acute accent is hardly a bleeding edge feature.
I haven't been able to interest any Computer Science professors in fun names for their courses, but I was able in my test environment to name a COMP series course "Untitled Course Name" with a description explaining that "It is a lovely day in the village and there are only two hard problems in Computer Science".
You mean cache invalidation and naming things? :)
For example, constant time subscripting, or improved length calculations, are made possible by encodings other than utf-8.
But when performance isn't critical, utf-8 should be the default. I don't see a reason for any other encoding.
Assuming you mean different encoding forms of Unicode (rather than entirely different and far less comprehensive character sets, such as ASCII or Latin-1), there are very few use cases where "subscripting" or "length calculations" would benefit significantly from using a different encoding form, because it is rare that individual Unicode code points are the most appropriate units to work with.
(If you're happy to sacrifice support for most of the world's writing systems in favour of raw performance for a limited subset of scripts and text operations, that's different.)
If you're hoping that a fixed offset gives you a user-percieved character boundary, then you're not handling composed characters or zero-width-joiners or any number of other things that may cause a grapheme cluster to be composed of multiple UTF code points.
The "fixed" size of code points in encodings like UTF-32 are just that: code points. Whether a code point corresponds with anything useful, like the boundary of a visible character, will always require linear-time indexing of the string, in any encoding.
(*) Approximately nothing. If you're in a position where you've somehow already vetted that the text is of a subset of human languages where you're guaranteed to never have grapheme clusters that occupy more than a single code point, then you maybe have a use case for this, but I'd argue you really just have a bunch of bugs waiting to happen.
What about UTF-256? Maybe not today, maybe not tomorrow, but someday...
E.g., Ruby runtimes scan the bytes in a string and then cache data about them in value called a code range. Knowing the code range, you can optimize many operations to not require additional linear scans of the string. Knowing a UTF-8 string consists only of ASCII characters can allow operations to be just as fast as if the string truly were ASCII-only (Ruby supports 100+ string encodings). And that fact is used throughout the core library to provide fast implementations of many operations (upcase, downcase, capitalize, gsub, substring, and so on). Moreover, a JIT can generate extremely tight code in those situations. Having to take a linear pass through the string to discover codepoint boundaries incurs a huge performance cost. While all strings could be treated uniformly and use Unicode tables for case mapping and such, the extra overhead is brutal. It has a measurable impact on string-heavy applications, such as template rendering and text processing.
In the most general case, yes, you know nothing about the string and can't make any assumptions. You can't even be sure the byte sequence is valid UTF-8. But, very often you do know properties of those strings. And you can manage boundaries where strings with known properties are joined with strings with unknown properties (e.g., variable interpolation in a template file).
With the exception of conversion from numbers (which has its own optimizations that are likely equally applicable in UTF8 since Arabic numbers are just ASCII anyway), I’d say all of your examples sound like bugs waiting to happen.
Why shouldn’t string literals be allowed to contain complex emoji? Why should identifiers disallow them? Why should there be a company policy around putting complex emoji places?
Just saying “let’s just declare things such that strings aren’t allowed to have multi-code point grapheme clusters” sounds great until you accidentally let that assumption leak into a place where a user wants to use an emoji and can’t make it match their skin tone.
I’d also say that such restrictions are putting the cart before the horse; the typical reasons for restricting the allowed character set, are precisely because you want to make lazy assumptions about things like string offsets. Saying that such assumptions are a good thing because you have these restrictions in place, seems like circular logic to me.
I didn't say they shouldn't, just that many do not and you know that at parse time.
> Why should identifiers disallow them?
I don't write the language specs. Many languages don't allow classes, methods, variables, etc. to have complex grapheme clusters in them.
> Why should there be a company policy around putting complex emoji places?
Performance. Code sanity. Indexing. Ease of typing. Again, I'm not the one writing the policies. But, they exist.
> Just saying “let’s just declare things such that strings aren’t allowed to have multi-code point grapheme clusters” sounds great until you accidentally let that assumption leak into a place where a user wants to use an emoji and can’t make it match their skin tone.
I'm making a clear distinction between situations where you have user-supplied data and data under control of the language runtime, developer-created files, or those just adhering well-defined file formats. These are all strings and commonly consist of simple codepoints; indeed, many times they're just ASCII characters.
I addressed user-supplied values when I wrote "And you can manage boundaries where strings with known properties are joined with strings with unknown properties (e.g., variable interpolation in a template file)." TruffleRuby, for example, uses ropes as its underlying structure, so if you have a template written using all ASCII characters (rather common) and interpolate a user-supplied value, you can put the user string in one rope, the template in others and link them all together into a tree with ConcatRopes. The template ropes still know they only have simple codepoints and operations on those parts can be fast. The user variable only knows it's a generic UTF-8 string and operations on that string, if any, can go down the slower path. Oftentimes, there are no operations to perform on that user string other than to display it. Its mere presence doesn't need to adversely affect the rest of the template.
> I’d also say that such restrictions are putting the cart before the horse; the typical reasons for restricting the allowed character set, are precisely because you want to make lazy assumptions about things like string offsets. Saying that such assumptions are a good thing because you have these restrictions in place, seems like circular logic to me.
I'm not making any assumptions. I've spent an awful lot of time optimizing string performance in the context of a Ruby runtime and your initial claim of constant-time subscripting being a myth doesn't match my experience. Ruby allows as complex of a string as you want, but the reality is there are many situations where strings, by either by restrictions or de facto, will not have multi-codepoint grapheme clusters. In many situations you'll have strings with all the codepoints in the ASCII range. If the only information the runtime records when parsing a string is that "this is a UTF-8" string and then operates on all UTF-8 strings uniformly, you leave a lot of performance on the table. The best performing situation is when you don't have to deal with variable-width codepoints in a UTF-8 string. UTF-16 and UTF-32 aren't terribly common in Ruby, but they exist as valid encodings (well, UTF-16BE/UTF-16LE and UTF-32BE/UTF-32LE) and have simpler execution paths than UTF-8 for many use cases.
Source code manipulation is frequently Unicode aware but doesn't care about combinations or things outside of a strict subset of Unicode to modify lexing control flow.
Being able to store (and later refer to) character offsets in the source code is a plus because they'll only ever occur in places where the strict subset is enforced.
This is especially true of languages with line-only comments, etc, where different writing systems being used won't affect the error message information.
Like I said, there are a few useful cases where having a fixed width encoding is beneficial. It's less helpful to the discussion to assert you know better for every case, ever.
It doesn't add anything over storing byte offsets into an UTF-8 string.
It also wastes valuable code space that could be used to have more characters encode as 3 bytes instead of 4 in UTF-8.
> Thompson's design was outlined on September 2, 1992, on a placemat in a New Jersey diner with Rob Pike.
If that isn't a classic story of an international standard's creation/impactful update, then I don't know what is.
Parallel with my annoyance with Microsoft, I realized how long it’s been since I encountered any kind of text encoding drama. As a regular typer of åäö, many hours of my youth was spent on configuring shells, terminal emulators, and IRC clients to use compatible encodings.
The wide adoption of UTF-8 has been truly awesome. Let’s just hope it’s another 15-20 years until I have to deal with UTF-16 again…
However, Powershell (or more often the host console) has a lot of issues with handling Unicode. This has been improving in recent years but it's still a work in progress.
UTF-32 / UCS-4 is fine but feels very bloated especially if a lot of your text data is more or less ASCII, which if it's not literally human text it usually will be, and feels a bit bloated even on a good day (it's always wasting 11-bits per character!)
UTF-8 is a little more complicated to handle than UTF-16 and certainly than UTF-32 but it's nice and compact, it's pretty ASCII compatible (lots of tools that work with ASCII also work fine with UTF-8 unless you insist on adding a spurious UTF-8 "byte order mark" to the front of text) and so it was a huge success once it was designed.
For example, this Chinese Wikipedia page is almost twice as big in UTF-16 as in UTF-8, because the vast majority of the content seems to be HTML tags:
$ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | wc -c
1877366
$ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | iconv -f utf-8 -t utf-16 | wc -c
3345724
If I strip out all the printable ASCII characters, it does indeed become about 33% more compact in UTF-16: $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | wc -c
311226
$ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | iconv -f utf-8 -t utf-16 | wc -c
213444
With gzip compression it is only about 10% more compact: $ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | gzip -9c | wc -c
98373
$ curl -s 'https://zh.wikipedia.org/wiki/中华人民共和国' | tr -d '[:print:]' | iconv -f utf-8 -t utf-16 | gzip -9c | wc -c
89031
So I guess if you're archiving pure CJK text, maybe you could get a 10% benefit, though I suspect non-Unicode encodings of that text would be more compact anyway.I can't find the more detailed article, but this summarizes it: https://techcommunity.microsoft.com/t5/sql-server-blog/intro...
From what I can remember, UTF-8 consumes more CPU as it's more complex to process, has space savings for mostly ascii & European codepages, but can significantly bloat storage sizes for character sets that consistently require 3 or 4 bytes per character.
UTF16 is really not noticeably simpler. Decoding UTF8 is really rather straightforward in any language which has even minimal bit-twiddling abilities.
And that’s assuming you need to write your own encoder or decoder, which seems unlikely.
Big endian or little endian?
My team managed a system that did a read from user data, doing input validation. One day we got a smart quote character that happened to be > U+10000. But because the data validation happened in chunks, we only got half of it. Which was an invalid character, so input validation failed.
In UTF-8, partial characters happen so often, they're likely to get tested. In UTF-16, they are more rarely seen, so things work until someone pastes in emoji and then it falls apart.
- 0xxxxxxx -> 7 bits, ASCII compatible (same as UTF-8)
- 10xxxxxx -> 6 bits, more bits to come
- 11xxxxxx -> final 6 bits.
It has multiple benefits: - It encodes more bits per octet: 7, 12, 18, 24 vs 7, 11, 16, 21 for UTF-8
- It is easily extensible for more bits.
- Such extra bits extension is backward compatible for reasonable implementations.
The last point is key: UTF-8 would need to invent a new prefix to go beyond 21 bits. Old software would not know the new prefix and what to do with it. With the simpler scheme, they could potentially work out of the box up to at least 30 bits (that's a billion code points, much more than the mere million of 21 bits).The
Even for pure text data, if a previous field was over-read (the only plausible way to have start-truncation), then you probably are decoding incorrect data from then on.
IOW, this upside is both ludicrously improbable and much more damning to the decoding than simply be able to skip a character.
10000011 11101001 U+E9 "é"
10000001 10000011 11101001 U+10E9 "ჩ"
This would make it incompatible with many existing processes that already handle text, which was one of the goals of UTF-8.With a code like this, your dumb terminal can be built so it just automatically works. Without it, the terminal may be off by some number of bytes, and you have to have some sort of other (probably manual) synchronization procedure.
Similarly, imagine you've dialed in over modem to a remote system. You get some errors or lost text due to noise on the phone line. Or some bytes are lost due to bad RS232 flow control or no flow control. Your terminal emulator would be able to recover.
There might also be uses where text is broadcast, such as in TV closed captions, although maybe those have their own framing so you know where strings begin.
EDIT: another huge advantage is that lexicographical comparison/sorting is trivial (usually the ascii version of the code can be reused without modification).
I have printed this out and inserted it into my safe deposit box, so my children's children's children can take it out and have a laugh.
Prefix code: The first byte indicates the number of bytes in the sequence. Reading from a stream can instantaneously decode each individual fully received sequence, without first having to wait for either the first byte of a next sequence or an end-of-stream indication. The length of multi-byte sequences is easily determined by humans as it is simply the number of high-order 1s in the leading byte. An incorrect character will not be decoded if a stream ends mid-sequence.
https://en.wikipedia.org/wiki/UTF-8#Comparison_with_other_en...
Also, the human readability sounds fishy. Humans are really bad at decoding high-order bits. For example can you tell the length of a UTF-8 sequence that would begin with 0xEC at a glance? With my scheme, either the high bit is not set (0x7F or less), which is easy to see you only need to compare the first digit to 7. Or the high bit is set and the high nibble is less than 0xC, meaning there is another byte, also easy to see, you compare the first digit to C.
The quote also implicitly mis-characterized the fact that in my scheme an incorrect character would also not be decoded if interrupted since it would lack the terminating flag (No byte > 0xC0).
> - It is easily extensible for more bits.
UTF8 already is easily extensible to more bits, either 7 continuation bytes (and 42 bits), or infinite. Neither of which is actually useful to its purposes.
> The last point is key: UTF-8 would need to invent a new prefix to go beyond 21 bits
UTF8 was defined as encoding 31 bits over 6 bytes. It was restricted to 21 bits (over 4 bytes) when unicode itself was restricted to 21 bits.
Extending UTF-8 to 7 continuation bytes (or more) loses the useful property that the all-ones byte (0xFF) never happens in a valid UTF-8 string. Limiting it to 36 bits (6 continuation bytes) would be better.
Maybe not in your bubble, but in my bubble this is the common case and highly optimized components using SIMD parsing is the exception.
By byte-by-byte, I presume you mean logically, not invoking a syscall for every byte. A parsing library is often handed a buffer to parse. If every layer had it's own buffering, you'd end up with too much data copying, which compounds and can spill your CPU caches in ways that microbenchmarks won't reflect, particularly in a streaming pipeline. So you stick to simpler, straight-forward code and data paths, unless and until you know a particular component is a bottleneck. I've had much better results optimizing globally first before optimizing locally. You can usually go back and optimize locally whenever you want, but optimizing globally is more often a one-shot deal as it typically requires non-local (i.e. cross-component) analysis and refactors, which nobody has time for.
The history of UTF-8 as told by Rob Pike (2003): http://doc.cat-v.org/bell_labs/utf-8_history
Recent HN discussion: https://news.ycombinator.com/item?id=26735958
Fun fact, UTF-8's prefix scheme can cover up to 31 payload bits. See https://en.wikipedia.org/wiki/UTF-8#FSS-UTF , section "FSS-UTF (1992) / UTF-8 (1993)".
A manifesto that was much more important ~15 years ago when UTF-8 hadn't completely won yet: https://utf8everywhere.org/
It’d probably be more correct to say that it was originally defined to cover 31 payload bits: you can easily complete the first byte to get 7 and 8 byte sequences (35 and 41 bits payloads).
Alternatively, you could save the 11111111 leading byte to flag the following bytes as counts (5 bits each since you’d need a flag bit to indicate whether this was the last), then add the actual payload afterwards, this would give you an infinite-size payload, though it would make the payload size dynamic and streamed (where currently you can get the entire USV in two fetches, as the first byte tells you exactly how many continuation bytes you need).
The awful truth is that there is such a beast. UTF-8 wrapper with UTF-16 surrogate pairs.
This is incorrect. You can only find boundaries between code points this way.
Until your you learn that not all "user perceived characters" (grapheme clusters) can be expressed as single code point Unicode seems cool. These UTF-8 explanations explain the encoding but leave out this unfortunate detail. Author might not even know this because they deal with subset of Unicode in their life.
If you want to split text between two user perceived characters, not between them, this tutorial does not help.
Unicode encodings are is great if you want to handle subset of languages and characters, if you want to be complete, it's a mess.
I do briefly mention grapheme clusters near the end, didn't want to introduce them as this article was more about the encoding mechanism itself. Maybe a future article after more research :)
Usually people write just the UTF-8 encoding part, then don't mention the rest of the Unicode, because it's clearly not as good and simple.
Displays as I assume intended in Firefox on the same machine: American flag emoji then when broken down in the next step U-in-a-box & S-in-a-box. The other examples seem fine in Chrome.
Take care when using relatively new additions to the Unicode emoji-set, test to make sure your intentions are correctly displayed in all the brower's you might expect your audience to be using.
Here's how it's rendering for me on Firefox: https://pasteboard.co/rjLtqANVQUIJ.png
Really nice diagrams nevertheless
And Chrome on Android, though with a rendering difference on the not-plain-ol'-U and not-plain-ol'-S (in both cases, blue letters rather than white with a blue background).
>From the previous diagram the value 0x1F602 falls in the range for a 4 octets header (between 0x10000 and 0x10FFFF)
Using the diagram in the post would be a crutch to rely on. It seems easier to remember the maximum number of "data" bits that each octet layout can support (7, 11, 16, 21). Then by knowing that 0x1F602 maps to 11111011000000010, which is 17 bits, you know it must fit into the 4-octet layout, which can hold 21 bits.
I'd like to add a correction. The binary/ascii/utf-8 value of 'a' (hex 0x61) is not 01010111, but instead 01100001.
This is used incorrectly in both the Giant reference card, and in the "ascii encoding" diagram above it.
The bytes come out as:
0xF0 0x9F 0x87 0xBA 0xF0 0x9F 0x87 0xBA
but the bits directly above them all of the bit pattern: 010 10111
https://speakerdeck.com/alblue/a-brief-history-of-unicode-45...
I'm also waiting for new emojis, they recently added more and more that can be used as icons, which is simpler than integrating PNG or SVG icons.
There are some attempts at font families to cover the majority of characters. Like Noto ( https://fonts.google.com/noto/fonts ), broken out into different fonts for different regions.
Or, Unifont's ( http://www.unifoundry.com/ ) goal of gathering the first 65536 code points in one font, though it leaves a lot to be desired if you actually use it as a font.
I'm guessing you could extract a subset of UTF-8 - but has anyone done anything like that?
Even if you don’t appreciate that convenience (and the highly standardized, easily abstracted rules of combined bytes), maybe you’ll appreciate that the complexity of supporting multilingual users would mean multipart data, arbitrarily fractured. As in a single text message using multiple character sets could be dozens of payload boundaries. Does that really sound easier or more painless than UTF-8’s variable-length code points?
Couldn't you directly provide an index in the grapheme "cache" ?
> I'm sure such a system wouldn't support some languages
> I'm guessing you could extract a subset of UTF-8
I can't tell if you're joking or not. You're perfectly describing ASCII.
“pretty well, all things considered”