Any Encoding, Ever – ztd.text and Unicode for C++
thephd.dev
thephd.dev
That reinterpret cast from char* to a char8_t* is undefined behavior.
std::u8string_view utf8_input(reinterpret_cast<const char8_t*>(argv[1]));
This is not just pedantry either, it was purposely designed this way to allow compilers to better optimize code without worrying about aliasing issues.Link to the paper which explicitly calls out that char8_t be defined as a new type that does not alias with any other type (hence making that reinterpret cast undefined behavior):
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p048...
reinterpret_cast<char*>(T*); // Perfectly fine.
reinterpret_cast<T*>(char*); // Undefined behavior.
Otherwise it would be trivial for any type to alias any other type: reinterpret_cast<T*>(reinterpret_cast<char*>(U*));Defining char8_t as a new type is specifically to avoid unnecessarily granting it these all-aliasing powers, but you can still read it as bytes.
Accessing the data using two different types would be problematic, except that if the other type is char, it is still fine as char is allowed to alias anything.
so tldr; as the programmer you can posit that the underlying data is indeed char8_t and the reinterpret cast is valid. You can also read it as char and would still be safe.
That said, I am in the rowdy “Only UTF-8, ever, and nothing else” gang. Even thinking about UTF-16 gives me hives: life is too short. But Shift-JIS is OK.
wobble transform 8bit seems like an unnecessary hijacking of a well labeled error state
People had occasionally used the label before for mojibake of various kinds, but the term was never popular under that meaning. Simon’s work is now a vastly more popular meaning.
It’s true that sometimes you may need language annotations, which will sometimes also need to be applied to substrings. I don’t think that invalidates my claim that rigid UTF-8 allows you to have sanity, though I will tweak it to state that using UTF-8 is a necessary but not always sufficient condition.
Given that context, I don't see that UTF-8 actually helps much. Your fundamental structure has to look something like a rope of (byte sequence+annotation) entries. With that structure, using different encodings for different segments doesn't make things noticeably worse.
In Unicode 5.2, I think, this range was elevated from “strongly discouraged” to “deprecated”: https://www.unicode.org/versions/Unicode5.2.0/ch16.pdf#page=....
In-band signalling in Unicode in general is fraught. The Unicode 5.2.0 specification linked goes on to show various of the reasons why this sort of tagging is generally problematic and should not be used in normal text. (And this is why they were strongly discouraged from the start.)
Text direction signalling is another troublesome area of Unicode; there are multiple techniques, some strongly discouraged, and it’s a perennial source of interesting security bugs. The only reason direction signalling is supported at all is because it’s needed. Life would be easier with it gone.
Stateful control characters make unicode mostly pointless - the whole point is to be self-synchronizing and have a universal representation for each character. (Granted emoji are busy destroying that already).
Other than that, it seems to present a C++ version of encoding conversion library (like iconv), but instead of handling encodings dynamically it uses C++ types. So not very good for any place with user-specified encoding (like browser or database or text editor), but can be used when the encoding is known at compile time?
I can see a few uses of it, but it seems kinda niche for all that grandiose talk about “liberation of the encoding hell”.
In any event, a single switch statement turns a compile-time encoding into a run-time encoding.
Personally I've enjoyed using the tiny-weeny header only utfcpp[0] for simple manipulation of UTF-8 and unicode-to-unicode conversions, and typically use Boost.Locale[1] when I need to do I/O.
95% of localization problems have almost nothing to do with encoding. Encodings are boring. It's like arguing over JSON vs YAML. To do I/O properly across platforms you may need wrappers for the standard output streams that will handle conversion for you, sure, but... you also need to handle date, time and currency formatting/parsing, message formatting, and to enable translations.
See [2] regarding Windows:
> All of the examples that come with Boost.Locale are designed for UTF-8 and it is the default encoding used by Boost.Locale.
Personally I think doing I/O as UTF-8 on Windows is the right direction, as Microsoft have been enhancing UTF-8 support in Windows for quite a while now.
See[3]:
> Until recently, Windows has emphasized "Unicode" -W variants over -A APIs. However, recent releases have used the ANSI code page and -A APIs as a means to introduce UTF-8 support to apps.
[0] https://github.com/nemtrif/utfcpp
[1] https://www.boost.org/doc/libs/1_76_0/libs/locale/doc/html/i...
[2] https://www.boost.org/doc/libs/1_76_0/libs/locale/doc/html/r...
[3] https://docs.microsoft.com/en-us/windows/apps/design/globali...
That said, with appropriate new api in the standard and with the compiler explicitly requiring explicit encoding specified sources, this could all be solved. The c++ streams are a very useful construct but horribly implemented, so we could just make a new one, and then deprecate the broken part of the standard lib. Its just, this is never a problem you realize when coding on your own, since it only ever hits projects big enough to be multilocal.
Projects big enough to be multi-locale (or actually all projects) should definitely be using format strings (to account for different parameter orders), locale sensitive date, time and currency formats, and a good externalized translation framework... but I still think they should be using UTF-8 throughout, because the encoding you use is completely unrelated to these problems.
C++ in particular is used less and less for human facing software, and where it is libraries like Qt are fairly excellent at handling this stuff.
Qt still leaves you with the problem of string litterals possibly changing when moving the code from one locale to another as the code itself is not of defined encoding, so if reparsed it will either look bad in code, or bad in output. The everpresent tr is also rather annoying.
Apparently the GP found a bug in this code. I'd be interested in seeing the github issue.
I'd assume that most filesystems don't care about encoding (and I think they really shouldn't), but there are exceptions in non-Linux worlds as well.
I agree the low level stuff should never care, they are always just binblobs in kernel, and should be. Though there are certainly still bugs caused by encoding changes of some conf file gets read by it.
The advantage of forcing it on a filesystemlevel was that the gigantic pile of old material with varying encoding at least was consistently off between locales rather than working for the people who used it most, and fail in the rare but critical cases for the people who rarely worked on it. Basically forcing people to fix their shit.
> If your type always has a replacement character, regardless of the situation, it can signal this by writing one of two functions: replacement_code_units() (for any failed encode step), replacement_code_points() (for any failed decode step).
They (and their cousins maybe_replacement_code_units/points) do not accept any argument, and that's wrong. There are stateful encodings like ISO/IEC 2022 and HZ where you might need the current encoder state to select a correct replacement character. For example `?` might not be representable in the current state and needs a switch sequence. You can, fortunately, always switch to the desired character set in 2022 with a fixed sequence so you can do have a fixed replacement sequence there, but you can't do so in HZ. It is technically possible to force-switch to the default state when an error is occurred, but not all users will want this behavior. It should obvious to anyone that stateful encodings are in general PITA.
Outside Windows APIs, UTF-16 is pretty much irrelevant (yes there are a few popular languages around which jumped on the UNICODE bandwagon too early and are now stuck with UTF-16, but those languages usually also offer simple conversion functions to and from other encodings - like UTF-8).
UTF-32 has its place as internal runtime string representation, but not for data exchange.
UTF-8 has won, and for all the right reasons. UTF-16 is a backward compatibility hack for legacy APIs, and UTF-32 is a special-case encoding for easier runtime-manipulation of string data.
e.g.
#pragma comment(linker, "/ENTRY:wmainCRTStartup")
int wmain(int argc, const wchar_t** argv)
{
...
// You can use this to actually replace WinMain with proper main (if you use /ENTRY:mainCRTStartup then it's just the good ole main).https://github.com/soasis/text/blob/main/examples/documentat...
Unfortunately this doesn't look so simple anymore.
The issue on Windows is simply that code pages are a per-application setting (or state) and not a system setting. Programs are free to run with, and change between, any "active code page" they want and spit bytes out in that code page.
This is all fine and dandy when you're talking to a -A Windows API, since Windows knows the ACP of your application, but it's disaster for I/O and making arbitrary programs talk to one another.
Nothing you can do in your code can fix this, it's a contract you have to have with the outside world.
On Linux calling setlocale() with a non-UTF-8 locale might not even even work because your /etc/locale.gen file will probably only have a couple of enabled entries, and doing so would be mostly pointless because your terminal and all your other programs are likely using UTF-8.
The downside to the Linux approach is that if somebody does actually send you a ISO-8859-1 encoded file you will have to do some conversion.
#include <boost/locale.hpp>
#include <fstream>
#include <iostream>
#include <string>
int
main(int argc, char** argv) {
std::ifstream ifs ("test.txt");
std::string str;
boost::locale::generator gen; // No dependency on what's in /etc/locale.gen
auto loc = gen ("en_US.ISO-8859-1");
ifs.imbue (loc); // Should make formatting input functions work
std::getline (ifs, str);
str = boost::locale::conv::to_utf<char> (str, loc);
for (auto wc: str) {
std::cout << sizeof(wc) << ": "
<< static_cast<uint64_t>(static_cast<unsigned char>(wc)) << "\n";
}
}Java does this, right?
Proof: int main (int argv, char* argv[]), i.e. both parameters are called argv.
> you want to work with higher-level primitives (code points, graphames) when iterating text that do not break your text apart;
In reality grapheme clusters are implemented nowhere and are only in the glossary. I wanted to see how this handled the sprawling complexity that can beget.
Someone call me when a library can iterate in that form, which isn't written by committee.
edit* Yikes https://github.com/JuliaStrings/utf8proc/blob/master/utf8pro... Still this is a good template to maybe cut a (safter) state-based iterator out of their logic...
The other thing I ran into recently, is that u"" string literals that worked on GCC / Linux, Clang / Mac, and VC++ 2017 / Windows 2012 did not work for a colleague on VC++ 2019 / Windows 10. An emoji, for example, came out as four char16_t code units (one for each byte of the UTF-8 encoding, I think), instead of two. We ended up using Unicode escapes instead, although the source code is less colorful without a pile of poo.
This ztd.text library looks interesting, although it's a little discouraging that the getting started section of the documentation is empty. Is this a header-only library?
(My reasons for doing so are backwards compatibility support and lowering the attack surface.)
utf16_output.size() * sizeof(char16_t)
Keep on keeping on, c++.Used as an example to demonstrate a self-described best of breed library, using all of the latest and greatest features of modern c++ (c++20 features even), and yet you still have to multiply the size of the container by the size of the data type inside the container in order to have correct behaviour.
It’s error prone, even for experienced and careful programmers, and despite all the advances of modern c++, this sort of thing permeates through almost every non-trivial c++ project.
I didn’t read the library’s documentation so maybe there’s an overloaded output operator that prevents people from needing to do this manually, but the fact that this is considered the normal correct way of doing things, is an indictment of c++ as a whole - and I say this as someone who uses c++ daily and who considers himself a safe and careful c++ programmer, who still occasionally makes mistakes due to things like this.
utf16_output.size()* sizeof utf16_output[0];
This removes the redundant indication of the element type (which can introduce bugs when the type of the array is changed). It gets even better when removing the naming slightly (names indicating types rarely make sense): out.size() * sizeof out[0];
With less magic: plain pointers instead of C++ vector like object, the following is total muscle memory for me: out_count * sizeof *out;
I don't see how it is a problem having to do a simple multiplication when what's needed isn't the size that is normally used ("number of elements"), but the size in bytes of the slice.Adding some kind of syntax of library feature to remove this trivial multiplication would add more non-orthogonality, as one would now use sth like
out.size_in_bytes()
but for custom slices would have to do the calculation manually: n * sizeof out[0];
Note also that in the former case, the need to multiply is removed syntactically, but not at runtime. Personally I don't think there is a worthwhile improvement to have.I'm also wondering what is the correct behaviour? In C++ you very very rarely want the size in bytes of a thing (much less often than in C), especially when implementing a generic container.
In C++, every contained element could be a complex type with copy and move semantics and what not, so you should not just deal with plain memory.
Go strings aren’t actually UTF-8; they aren’t even Unicode: they can contain arbitrary bytes. In practice, most strings will be UTF-8, but you can’t depend on it at all.
Rust strings are strictly UTF-8. I appreciate this greatly.
Python 3 strings are sequences of Unicode code points (as distinct from Unicode scalar values), allowing ill-formed Unicode (unmatched surrogates), and its internal representation is a disaster because they decided indexing by code point was worthwhile (it’s not, given the costs), so since CPython 3.3 strings are encoded as ISO-8859-1 (only able to represent U+0000 to U+00FF), UCS-2 (further able to represent U+0100 to U+FFFF) or UCS-4 (able to represent all Unicode code points).
JavaScript strings allow ill-formed Unicode, and the public representation can be described as either UCS-2 or ill-formed UTF-16 (almost all string access techniques work in UTF-16 code units, though new stuff is now mostly working in Unicode and UTF-8 terms). I think all major engines now use an internal representation of ISO-8859-1 if the string contains only U+0000 to U+00FF, or ill-formed UTF-16 for anything else; but Servo’s work has demonstrated that it’s possible to shift to WTF-8 (UTF-8 plus allowing unmatched surrogates), saving lots of memory on some pages and simplifying and slightly speeding up some things, but at the cost of random access performance by code unit index.
Go strings are supposed to be UTF-8, but it's not enforced.
Javascript - bleah.
No, it never looked like it was UTF-8. It did look like it was Unicode (of an unspecified encoding), but Python allows unpaired surrogates, which makes for ill-formed Unicode. Python lets you have '\udead'. Well-formed encodings of Unicode (such as UTF-8, UTF-16 and UTF-32) cannot represent U+DEAD.
There’s a big difference between “Unicode” and “UTF-8”.
> PyPy, though, finally went UTF-8.
Oh cool! I remember vaguely hearing the idea being mulled over, years ago. Good to see it happened: https://morepypy.blogspot.com/2019/03/pypy-v71-released-now-....
Wonder what they’ve done about surrogates; the only way they can truly have switched to UTF-8 internally is if they’ve broken compatibility, which I doubt they’ve done; I’m guessing that they actually use WTF-8, not UTF-8.
(In the meantime, I’ve changed “Python 3.3” in the grandparent comment to “CPython 3.3” for accuracy.)
https://stackoverflow.com/questions/12053168/how-to-properly-output-a-string-in-a-windows-console-with-go
https://github.com/intellij-rust/intellij-rust/issues/766
https://bugs.python.org/issue44275
etc etc