And it is not so much WHAT he explained but HOW he explained it, what really made it stick. It was his sheer unbridled genuine enthusiasm that put him on the map for me.
One of the best to do it.
And it is not so much WHAT he explained but HOW he explained it, what really made it stick. It was his sheer unbridled genuine enthusiasm that put him on the map for me.
One of the best to do it.
https://sqlquantumleap.com/2018/09/28/native-utf-8-support-i...
The main impediment seems to be that languages like Java, JavaScript, and Python treat UTF-8 as just another encoding, but really it’s the most fundamental encoding.
The language abstractions get in the way and distort the way people think about text
Newer languages like Go and Rust are more sensible, they don’t have global mutable encoding variables
https://en.wikipedia.org/wiki/UTF-16#U+D800_to_U+DFFF_(surro...
UTF-8 is the most ubiquitous encoding on the web, but that doesn't make it more fundamental than any other.
Either that, or I'm completely off base and part of the vast horde who don't get it.
And while UTF-8 tries very hard to be simple whenever possible, it does have some nonobvious constraints that make it significantly more complex than the actual simplest variable-width encoding (a continuation bit and then 7 data bits).
Except when it's 1, because that's an invalid start, and except when it's 0, because that means the character is a single byte.
And it also means you're dealing with three classes of byte now.
Plus UTF-8 has more invalid encodings to deal with than a super-simple format.
If your format supports non-canonical encodings you're in for a bad time no matter what, so a whole lot of that simplicity is fake.
> And it also means you're dealing with three classes of byte now.
If you're working a byte at a time you're doing it wrong, unless you're re-syncing an invalid stream in which case it's as simple as a continuation bit (specifically, it's two continuation bits).
> If you're working a byte at a time you're doing it wrong, unless you're re-syncing an invalid stream
It's very relevant to explaining the encoding and it matters if you're worried that invalid bytes might exist. You can't just ignore the extra complexity.
Also if you're not working a byte at a time, that kind of implies you parsed the characters? In which case non-canonical encodings are a non-problem.
Unless you want to actually do anything with the string beyond decode a codepoint.
Non-canonical encodings make it difficult to do things without decoding, but you have bigger problems to deal with in that situation, and the non-canonical encodings don't make it much worse. Don't get into that situation!
Specifically, even with only canonical encodings, one and two byte characters can appear inside the encoding of two and three byte characters. You can't do anything byte-wise at all, unlike UTF-8. But you already said "If you're working a byte at a time you're doing it wrong" so I hope that's not too big of an issue?
It seems like you're advocating people learning incorrect information and forming their impressions of it on falsehood. Which is probably why people think utf-32 frees you from variable-length encoding.
As part of a comprehensive dive into Unicode it's a minor part, but for teaching an encoding it's a significant difference.
I've lectured computer science at the university level, and I think you could introduce all this information to a CS undergrad pretty coherently and design a lab or small assignment on it no problem. Maybe you could ask them to parse some emojis that require multiple 32-bit codepoints.