Adopting Erlang: Hard things to get right (2019)
adoptingerlang.org
adoptingerlang.org
Having a list of codepoints is kind of useful, but unicode gets pretty complex, some things are multiple codepoints but essentially indivisible (unicode flags, for example don't make much sense with only one code point), so codepoints aren't the useful unit that byte/characters are in 8-bit encodings.
Complete tangent, but:
To me, the Unicode flag codepoints aren't really "combining codepoints with weird behavior" in the sense that e.g. a zero-width joiner is a combining codepoint with weird behavior; so much as they're a special alphabet that colors the text written in it as an ISO country code; where valid country-codes spelled out in that alphabet just happen to be ligature-rendered in a special way by font authors.
It makes just as much sense to have an isolated Regional Indicator Symbol Letter C (U+1F1E8) in text as it does to have an isolated regular C in text. (HN strips out the former, but it does have a visual representation!) In both cases, they're one letter of a "word", and don't make sense to use, but do make sense to mention. You can talk about "countries whose ISO country codes start with U+1F1E8" just like you can talk about "countries whose names start with C." This is unlike e.g. the individual surrogate-pair codepoints in UTF-16, which don't make sense to use nor mention on their own.
And it can also make sense to use the Regional Indicator Symbol Letters to spell out country-codes that don't exist, without the expectation that they'll be rendered as a flag at present; simply because you're adding a layer of semantic meaning to your text for machines (and potentially getting a future flag-symbol for humans as a bonus.) The meaning — "this text is colored to be an ISO country-code" — is still there, whether the text renders a flag or not. It's a forward-compatible encoding, in a way that "individual flag emojis for each country" wouldn't be.
Mind you, I'm on the fence about whether I like the Unicode Committee's choice to implement these codepoints as a special alphabet. On the one hand, I appreciate that these characters need to be their own set of semantically-distinct codepoints in some sense, so that there's no semantic ambiguity of these being interpreted by machine-parsers as "letters in a word", like there would be if these were letters spelled out in the Wingdings font. But on the other hand, would it have really hurt the design of Unicode so much to just come up with a semantic-meaning equivalent to the Variation Selectors, so that we could just use regular letters A+Z, plus a prefix codepoint on each that means "interpreted as an ISO country-code"? Semiotics Selectors, per se?
I had a junior that came from python that did not want to work in elixir because he refused to write unit tests and the fact that double quotes were different from single quotes (and double quotes were default) he couldn't deal with -- joke was on him because python3 made strings a whole lot more complicated!
BEAM's JIT is different. It does some optimization, but it applies to all code loaded, and there's no fallback to an interpreter: the interpreter still exists, but you either run BEAM with the interpretter or BEAM with the JIT; there's no mixed mode, previous attempts at a BEAM JIT found that calling between JITed and interpreted BEAM code was very difficult, so this JIT takes a different approach. This approach makes it hard to apply in depth optimizations, because they would tend to delay loading. However, it eliminates the interpreter overhead completely.