Code Pages, Character Encoding, Unicode, UTF-8 and the BOM [video]
hanselman.com
hanselman.com
I don’t know about everyone else, but I had to implement an ISO 8859 encoding to UTF-8 converter in assembly when I was at uni around 8-10 years ago. So this is standard stuff for most developers who graduate from the University of Oslo.
I guess assembly made a little more sense 8 years ago and they only had so much class time but I wouldn’t have minded a “decide utf-8” assignment. It would have been another good practice doing bit manipulation which is kind of difficult to get right in C++.
Take the BOM as an example.
I work on backend Java projects for large banks. Over the years I fought with BOM on numerous times. For some reason 95% of software that says it is UTF-8 compliant is not. My modus operandi for dealing with it is to remove BOM on ingestion and only add it when the string leaves the system when we know the outside party absolutely requires it (though it should not...)
UTF-8 BOM is largely a Microsoft idea, they've got a bunch of code that thinks in UCS-2 (now retrofitted to more or less pretend it knows UTF-16) and so it thinks about byte order when decoding text files, and from there a Byte Order Mark in files that don't have byte ordering seems like a reasonable idea.
If the files actually _mean_ something then a UTF-8 BOM just introduces confusion. Lots of code I'm responsible for processes UTF-8 just fine, but if it handles say files full of key = value pairs and your file begins with a BOM, well, OK then, that first key starts with U+FEFF, weird choice but no reason we should disallow that. And of course that isn't what you wanted and so now Windows users are complaining I'm not "compatible".
It seems arbitrary, but I don't think they could have made a better choice.
I agree that getting into details of text rendering and so on is one of those 3xx or 4xx courses with narrow appeal because maybe one student in a thousand will actually make use of this knowledge, but grokking why UTF-8 is how it is and the basic outlines of Unicode seems very broadly applicable.
For most software development jobs you can get by without knowing this stuff but it's great there are things like this to clearly explain fundamentals that are either assumed knowledge or communicated in (to an outsider) gatekeeping levels of dense terminology.
https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...
Wait, what's so wrong about mentioning UTF-7? Wasn't it just a (proposed but abandoned) way to represent Unicode characters in MIME email?
It would have been cool to be able to incrementally upgrade legacy environments to use UTF via UTF-7. Unaware parts would just have displayed the encoding. String lengths would have sort of worked.
(All of these things would of course have come with horrible drawbacks, so in that alternative universe I might have been cursing that we got UTF-7...)
His single most important fact still rings true though, "It does not make sense to have a string without knowing what encoding it uses."
I'm wondering if it's widely used.
Which is why all my robots.txt files have a comment on the first line.
That doesn't stop a BOM being generated or consumed.
You can occasionally see it in git diffs as U+FEFF, or if you open a text file in a hex editor as EF BB BF
Neither does any other of the hundreds of existing text encodings.
It's debatable how much of a magic number it's supposed to be anyway, considering that few people have insisted on having magic numbers in text files, and that you get the BOM at the beginning by simply naively converting a UCS-2/UTF-16 file codepoint by codepoint (and vice versa, enforce it to be there if you ever happen to do the conversion the other way around because of course you're conversion couldn't include that extra logic in it).
$ curl -sO https://www.gutenberg.org/cache/epub/16681/pg16681.txt
$ file pg16681.txt
pg16681.txt: UTF-8 Unicode (with BOM) text, with CRLF line terminators
$ head -c3 pg16681.txt | xxd
00000000: efbb bf ...