C String Creation
maxsi.org
maxsi.org
Also open_memstream is POSIX 2008, now if it would just get into OS X…
asprintf()
aswprintf()
vaswprintf()
fmemopen()
open_memstream()
open_wmemstream() int main(int argc, char *argv[]) {
wchar_t buf[100];
wprintf(L"Hello, world!\ntype something>");
if (fgetws(buf, 100, stdin))
wprintf(L"You typed '%ls'\n", buf);
if (argv[1]) {
char *s = argv[1];
/* Convert char string to wchar_t string */
size_t len = mbsrtowcs(buf, &s, 99, NULL);
if (len != (size_t)-1) {
buf[len] = 0;
wprintf(L"argv[1] is '%ls'\n", buf);
}
return 0;
}
It's a pain, but the advantage is access to iswalpha() and friends.One hack is to assume that bytes of UTF-8 encoded strings above 127 are all letters. It mostly works :-)
Am I misunderstanding you, because I've always thought that's what the mbtowc(3) family of functions was?
if (isalpha(*s)) {
*d++ = *s++;
while (isalnum(*s))
*d++ = *s++;
}
To use UTF-8 / Unicode should require only small changes: if (iswalpha(decode(&s)) {
encode(&d, advance(&s));
while (iswalnum(decode(&s))
encode(&d, advance(&s));
}
For efficiency, don't decode twice- have the decoder return a pointer to the next sequence: if (iswalpha(c = utf8(&s, &n))) {
encode(&d, c);
s = n;
while (iswalnum(c = utf8(&s, &n))) {
encode(&d, c);
s = n;
}
}
Also should be able to match a string in line: if ('A' == utf8(&s, &t) && 'B' == utf8(&t, &s) && 'C' == utf8(&s, &t)) // we have 'ABC'.Secondly, no, you're asking for a massive addition of 2 new versions for every interface that mentions wchar_t. That's a huge addition to standard libraries. That's error prone and bloats things up. Then additionally you're asking for a rewrite of all software using wchar_t. And only until everything is transitioned, which isn't going to happen, the standard libraries will be much larger.
The solution is rather to embrace wchar_t and fix it. All sensible and modern platforms, which is a premise of this article on modern POSIX functions, have a 32-bit wchar_t type. That's excellent. It's only Windows, which due to historical short-sightedness that have 16-bit wchar_t. But writing portable C for native Windows is a losing game, the winning move is not to play. (Do see midipix which is upcoming and will provide a new POSIX environment for Windows with musl and 32-bit wchar_t). In fact, 16-bit wchar_t violates the C standard. That moment you give up broken platforms with 16-bit wchar_t, wchar_t works as intended, and this is a non-problem. Embracing char16_t and char32_t is a worse problem and isn't solving anything.
Those applications don't really care about the actual unicode codepoints besides ASCII. If you start to deal with visual representation of strings, calculating the column for error messages, advanced unicode-aware parsing, font rendering, and so on, then you do want to convert on the fly to wchar_t. mbsrtowcs and such are kinda bad, because they convert the whole string at once, which means an allocation that can fail in the unbounded case. It's usually sufficient to decode one wchar_t at a time with mbrtowc.
This way, char and wchar_t are not replacements for each other, but complement each other by being better abstractions for various purposes. Now, the wide stdio functions is where things start to get a bit useless, because the regular stdio char functions are perfectly fine and those functions don't really appeal well to the strengths of wchar_t.
The power of C -- which distinguishes it from most other languages -- is the ability to allocate almost everything statically.
In fact, the older I get the more I appreciate pre-Algol60 way to allocating stack frames statically.
But when C is appropriate, and these problems arise, which they will in any C codebase of appreciable complexity, these string creation interfaces are waiting for you, and will help you write correct code. It's usually a worthwhile effort to reconstruct higher level abstractions in C, in a good manner, for the same reasons you use them in higher level languages.
Note that statically allocating everything is hardly always possible. See the distinction between bounded and unbounded in the article, the unbounded case is really common.