I'm reminded of a time a colleague needed something like string.split, and working in c++ he filled a std::vector<std::string> with the result. Using a more C way he'd really only have needed a couple of pointers on the stack.
Ah, yes, the good 'ole C way that only works reliably in English.
You can parse utf-8 character at a time. Some characters advance the pointer by 4 at an iteration and some less.
I do NOT mean that char itself corresponds to a glyph or codepoint, you are seriously preaching to the choir making that lecture to me.
And when do you stop? UTF-8 strings can have zero bytes in them so treating them as C strings is potentially error prone depending on the context.
If you really do need it there are some C language libraries that use "pascal-ish" structs to do strings. UNICODE_STRING in Windows comes to mind. Doing strings in C doesn't force you to use C strings, it's just the most common thing to do.
This is not true. A zero-byte in a utf-8 string is the null-terminator and utf-8 strings can be treated exactly like C strings in terms of where the string ends.
What you do need to look out for is malformed utf-8, for example, 1 byte before the null terminator you get a lead byte saying the next character is 4-bytes long.
If you're not checking each byte for null and just skipping based on the length indicated by the lead byte then you're in for a crash.
Where utf-8 strings differ from C strings is slicing. You can't just slice the string at some random point without doing extra validation to make sure you only slice on codepoint boundaries.
No, the parent was correct: UTF-8 encodes NUL (i.e. \0) as a single zero byte (e.g. in contrast, Modified UTF-8[1] uses an overlong for NUL, so there's never any possibility of an internal zero). Of course, an application/library can choose to restrict itself to only handling UTF-8 that doesn't contain internal NULs, but the spec itself allows for zero bytes in a string.
By definition, with a null-terminated string, NUL is the terminator.
If you want to have strings that contain NUL, then by definition you can't use a null-terminated string.
This is true of utf-8 or regular C strings.
If someone passes you a text file that is verified to be valid UTF-8 and contains, say, access permissions, then you better not stop parsing it at the first '\0' character.
None of this is a huge problem, but it's something to be aware of. C string handling is incompatible with UTF-8.
That's separate from string handling.
UTF-8 was originally designed to be compatible with NUL terminated strings and keep NULs out of well formed text.
In fact it was the first point in the 'Criteria for the Transformation Format', mentioned in the initial proposal for utf8.
The UTF-8 spec doesn't make that distinction as far as I know. There's a simple fact: A valid UTF-8 byte sequence can contain nul characters. So you can't naively use C string handling functions on it. And as someone else has correctly pointed out, the same is true for ASCII.
I'm just pointing out a potential pitfall and a source of security issues. Some might assume that after validating UTF-8 text input, you could just dump it in a C string and process it using C's string functions. But that's not the case.
Yes, you can avoid this if you're careful and you understand the intricacies of utf-8 (or some other multi-byte encoding), but it very quickly stops being elegant.
Go parse some zalgo with your 4 per iteration algorithm. I'll be there, waiting and laughing.
C string handling is not elegant, nor does it fit the realities of the world.
Parent said string handling in C was elegant. My point is that it becomes fraught with (even more) issues once you throw non-English language at it.
It is C's decision to handle strings in this way, and the decision of many C programmers to treat all strings as if they are just iterable character pointers.
It's a recipe for bugs.
I've heard (mostly here) that Swift does something different and treats glyphs as the basic unit. I haven't had a chance to look at precisely what that does. Given all the issues I've seen elsewhere I'm skeptical that someone, anyone can pull that off correctly.
UTF-8 at least has one elegance (there's that word again) in the design in that you can do some "dumb" ASCII things and if your code does not know what to do with fancy unicode, you can check the high bit of any given octet and "safely" skip over it and any adjacent nonascii sequence if you don't know what it means. This may or may not be applicable to a task at hand.
This is true, however even something as simple as storing the (byte) length as part of the string reduces the complexity and the likelihood for bugs.
Other languages also prevent accidental buffer overruns so while they still need to deal with all the same Unicode problems you mentioned, the program likely won't crash if the programmer gets things wrong. The same is not necessarily true of C.
IIRC C++ is getting slices too, so it might be able to get better APIs around string manip. But I've seen decent string manip code that avoided allocations.
This is pretty much how it's done in Rust too via slices. For example, the standard way to split a string is to create an iterator and it won't do any allocations.
I try to avoid manipulating strings as much as possible. ;)