The WTF-8 Encoding
simonsapin.github.io
simonsapin.github.io
Charset="WTF-8" (xn--stpie-k0a81a.com) 25-nov-2024 437 comments https://news.ycombinator.com/item?id=42226953
The WTF-8 encoding (simonsapin.github.io) 27-may-2015 104 comments https://news.ycombinator.com/item?id=9611710
> Any WTF-8 data must be converted to a Unicode encoding at the system’s boundary before being emitted. UTF-8 is recommended. WTF-8 must not be used to represent text in a file format or for transmission over the Internet.
I strongly disagree with that part. When you need to be able to serialize every possible Windows filename, WTF-8 is a great choice. This could be a backup tool, or an NTFS driver for Linux.
I also think rust's serde should always serialize OsString as a bytestring, using WTF-8 on Windows. Instead of the system dependent union of u16/u8 sequences it currently uses.
> There is no and will not be any encoding label [ENCODING] or IANA charset alias [CHARSETS] for WTF-8.
The goal is to ensure WTF-8 remains fully contained, so that ill-formed strings don't end up processed by systems that expect well-formed strings.
If you need to serialize every possible Windows filename, then you must also own the corresponding de-serializer (ie make your solution self-contained), and cannot expect users to work with the serialized contents using tools you do not control.
> WTF-8 (Wobbly Transformation Format − 8-bit) is a superset of UTF-8 that encodes surrogate code points if they are not in a pair. It represents, in a way compatible with UTF-8, text from systems such as JavaScript and Windows that use UTF-16 internally but don’t enforce the well-formedness invariant that surrogates must be paired.
WTF-8 is necessary for Rust’s compatibility with Windows filesystems (it underlines OsString on Windows) as e.g. file names are sequences of UTF-16 code units (and thus may contain unpaired surrogates).
Somehow UTF-16 reserves some of those decoded integer values (instead of solving its whatever problem it had in its encoding itself)
The fact that UTF-8 didn't need to also destroy some output integer values to work proves it's not necessary to do that
Encoding and decoded value should be separate concerns
That's like having a mathematical encoding of integers that's like base 10, but for some reason you decide that integer values 100 to 110 are reserved and may never be used by anyone, not even other legit encodings like regular base 10
This property is not true of UTF-8 - if you get a byte-string with bytes between 0x80 and 0xFF, it might be UTF-8, or it might be one of a bunch of other encodings, you need to do a more involved check to be sure.
Granted, the presence of a value between 0xD800 and 0xDFFF does not guarantee that the text is UTF-16, that's why this "WTF-8" encoding exists. But confusion would be a whole lot more likely if the U+D800-U+DFFF range were not reserved.