The Wonderfully Terrible World of C and C++ Text Encoding APIs (With Some Rust)
thephd.dev
thephd.dev
It's super convenient for working with Windows file APIs, but everything else is still a pain.
Handling UTF-8 versus WTF-8, and properly round-tripping these representations, is probably a bigger issue than graphemes and canonical forms for most users.
Really? I'd love to see some references to that.
I know there's some operations to convert a javascript string to UTF8, but everything which measures a string counts using 2-byte UCS2 items. For example, array-style indexing, string.length, str.charCodeAt(), split(), slice(), indexOf(), etc.
Also lots of DOM APIs use those weird string lengths too. For example, input element selectionStart / selectionEnd fields measure the selection range by counting surrogate pairs, not characters.
I remember that SpiderMonkey shifted to multiple encodings around 2019 and v8 had done that earlier.
Source: I was part of the SpiderMonkey team when the shift happened.
Edit; that is; it does not have to be linked in the standard lib. Can be a data file somewhere, or a a shared lib.
Locales have some similar behavior as well.
Everyone's so good at patching vulnerabilities, but did you plan for your app breaking in production for all users in particular timezones because a random service in a chain of a 6, say, re-implemented in Go last quarter, or maybe... changed their Node app to port from using Moment to the standard Intl object 10 months ago, thus depending on the ICU C++ library, and a separate copy of the data that is no longer correct? Try to assess that risk and your head will hurt. These are the sharp edges in a polyglot environment that usually aren't considered when deciding to bring in new languages.
Unicode CLDR has the same issue, but the pain usually isn't acute because it takes much, much longer for new glyphs or other data to appear in common use or start appearing in data feeds or from OS APIs/input. As long as you're at least on some update schedule, it's usually fine, and there's usually fewer things to update.
Time zones, on the other hand...
a) conversion was exactly what was provided but C++11, before they deprecated it
b) plenty of staff that the standard library does is complicated! Some of it also seems less fundamental than handling plain text.
If you are writing some software that requires counting the number of characters in a string you should stop and rethink why you're trying to do what you're doing.
(The function is defined as static in the cpp file where RTL detection is done, with a warning that is not general purpose)
It's more crazy that C++ was so slow to mandate UTF-8 support. You have the situation where your modern C++ environment might have several distinct "string" types none of which is guaranteed to just be UTF-8.
UTF-16 should work in there cases eventually though.
For small Japanese-only text, Shift-JIS really will be smaller, but it's just not usually enough to care, and the price is you can't do Unicode. It occupies a similar space to 8859-1 / Windows 1252 where older software uses it but you should just transition.
So, you're right, I need to amend that to "mostly needed for file formats and protocols".
Users should actively file bug reports with any system or product that insists on UTF-16 or abandon use of such legacy systems if they refuse to change.
how would that work if you're on a microcontroller ? if $GOV_AGENCY says "product XXX is made according to international standards, including ISO/IEC 14882:2026" and international standard ISO/IEC 14882:2026 now says that a complete Unicode implementation has to be provided otherwise you're not compliant, any device with less than the 20-30-ish megabytes of memory needed for the unicode database won't be able to pass certification even if they don't handle text at any point, e.g. it's some dsp filter somewhere in a camera lense
> __STDC_ISO_10646__
> An integer literal of the form yyyymmL (for example, 199712L). If this symbol is defined, then every character in the Unicode required set, when stored in an object of type wchar_t, has the same value as the code point of that character. The Unicode required set consists of all the characters that are defined by ISO/IEC 10646, along with all amendments and technical corrigenda as of the specified year and month.
You can see the wording, "if this symbol is defined".
Note that even if your implementation supports Unicode, and has all sorts of character conversion tables and character property tables, you can easily arrange for those to be statically linked, and only included if actually used. This is normally how standard libraries work in programming environments for resource-constrained systems like microcontrollers.
Another interesting macro is the __STDC_HOSTED__ macro--basically, if __STDC_HOSTED__ is missing, then large chunks of the library will not be available. Still completely standards-compliant.
additionally, if your device has less than a few megabytes of storage, you'll almost certainly be using static linking, which automatically prunes unused code from properly designed libraries.
[1] https://github.com/llvm-mirror/llvm/blob/master/lib/Support/...
No. If you have a common function with 13 arguments:
U_CAPI void ucnv_convertEx(
UConverter *targetCnv, UConverter *sourceCnv, // converters describing the encodings
char **target, const char *targetLimit, // destination
const char **source, const char *sourceLimit, // source data
UChar *pivotStart, UChar **pivotSource, UChar **pivotTarget, const UChar *pivotLimit, // pivot
UBool reset, UBool flush, UErrorCode *pErrorCode); // error code out-parameter
you've done something terribly wrong. Refactor it as a method, do "factory" stuff, something. Just shocking what C people are able to put up with.Object oriented code is very common in many larger C libraries (or things like the linux kernel). The only difference is that you don't get any compiler support to prevent errors.
The standard model is:
typedef struct __MyType MyTypeRef; struct MyTypeMethodTable { // the destructor - honestly at an api level you should probably have retain/release instead void (destroy)(MyTypeRef _this);
// void (someMethod)(MyTypeRef _this, int someArg); };
struct __MyType { MyTypeMethodTable vtable; };
void MyType_destroy(MyTypeRef _this) { _this->vtable->destroy(_this); }
void MyType_someMethod(MyTypeRef _this, int someArg) { _this->vtable->someMethod(_this, someArg); }
Then an actual type is implemented as
struct MyConcreteType { struct __MyType base; // some fields };
void MyConcreteType_destroy(MyTypeRef value) { MyConcreteType realValue = (MyConcreteType )value; // cleanup anything you need to do free(realValue); }
void MyConcreteType_someMethod(MyTypeRef value, int someArg) { printf("I: %d\n", someArg); }
MyTypeMethodTable MyConcreteType_MethodTable { .destroy = MyConcreteType_destroy, .someMethod = MyConcreteType_someMethod };
MyTypeRef CreateMyConcreteType() { MyConcreteType *result = calloc(1, sizeof(MyConcreteType)); result->base.vtable = & MyConcreteType_MethodTable; result->someField = whatever; return &result->base; // Or similar. avoid UB in C can make this weird }
You can see that trivially this is pretty much what notionally OO languages like C++, Java, Haskell, etc do.
An apparently not-uncommon error that happens in COM is:
someObject->whateverTheirMethodTableIsCalled->someMethod(theWrongObject)
Possibly with someObject and theWrongObject the other way around. The end result is sadness either way.
OO programming is super effective for many things, and generally better for a lot of design, especially for libraries and frameworks. But nothing about OO requires compiler/language support - indeed the first discussions of OO code predated OO languages - but it's hopefully obvious to see that if the compiler can manage this, it results in less code, and less opportunity for error, and because the compiler manages those semantics it should technically produce better code (e.g if the compiler knows that "vtable" is actually a vtable pointer, it knows there is no circumstance it can change[1]).
[1] Yes you could have incorrect code through UB, but in that case the compiler is allowed to do the "wrong" thing.
There's no hidden magic that will come back and bite you when it changes in another part of the code base.
The truth is that this isn’t really a “common” function; most programs will call it (or one of its siblings) just a few times for input and output. You’re write that code once and then never worry about it again.
Leaving all that aside, if you think you can do better, feel free to write a wrapper library and prove it.
CFStringRef CFStringCreateWithBytes(CFAllocatorRef alloc, const UInt8 *bytes, CFIndex numBytes, CFStringEncoding encoding, Boolean isExternalRepresentation);
CFDataRef CFStringCreateExternalRepresentation(CFAllocatorRef alloc, CFStringRef theString, CFStringEncoding encoding, UInt8 lossByte);
Nice and easy, perhaps a teeny tiny bit inflexible? :D