I'm wondering if it's widely used.
I'm wondering if it's widely used.
$ curl -sO https://www.gutenberg.org/cache/epub/16681/pg16681.txt
$ file pg16681.txt
pg16681.txt: UTF-8 Unicode (with BOM) text, with CRLF line terminators
$ head -c3 pg16681.txt | xxd
00000000: efbb bf ...Which is why all my robots.txt files have a comment on the first line.
That doesn't stop a BOM being generated or consumed.
You can occasionally see it in git diffs as U+FEFF, or if you open a text file in a hex editor as EF BB BF
Neither does any other of the hundreds of existing text encodings.
It's debatable how much of a magic number it's supposed to be anyway, considering that few people have insisted on having magic numbers in text files, and that you get the BOM at the beginning by simply naively converting a UCS-2/UTF-16 file codepoint by codepoint (and vice versa, enforce it to be there if you ever happen to do the conversion the other way around because of course you're conversion couldn't include that extra logic in it).