To achieve safety, one option is to use routines that always assume that the data they are operating on may be invalid — but such routines are always slow, because they are essentially validating the UTF-8 data continually.
Therefore, what most systems do is establish the validity of UTF-8 data as an invariant by running a sanity check once before supplying the data to unsafe-but-fast internal routines. And that's where validators like the one described in this article come in.
I struggle to imagine such a function. If its signature is "const char * utf8end(const char * s, size_t len)", then the implementation is obviously "return s + len". If the len is missing then, uh, you can't find the end unless there is a terminator character, in which case the implementation is "return s + strlen(s)".
And if you mean finding the beginning of the last codepoint, then again, you simply go to the very last byte of the string and rewind back until you see a non-trailing byte (or you've run off the beginning of the string, or you've seen more than 4 bytes in which case you have an invalid string).
The following function will read invalid memory if there are missing trailing bytes:
int32_t
read_last_code_point(const char *s, size_t len) {
s += len; // place pointer at end of string
while (!HEADER_BYTE(s)) { // read backwards until header byte found
s--;
}
return DECODE_UTF8(s); // BOOM! May read beyond string end
}
To fix it, you would either need to establish that the passed-in string was valid UTF-8, or you would need to add a bounds check to DECODE_UTF8, making it slower.When you write string manipulation code, you come across this sort of issue all the time. In isolation, maybe you'd want to fix this function by adding the bounds check. But since you're going to see this problem everywhere, you'll probably solve it through an initial validation pass instead.
1) If position == end, the iterator is done. 2) Otherwise, read the next byte to determine the number of bytes in the next character. 3) Read all the bytes of the next character, convert them to a regular 32-bit code point, and emit that.
In the presence of invalid UTF-8, that iterator is broken. There needs to be an extra check in step 2.5: "Check that the number of bytes reported for the next character doesn't exceed the number of bytes remaining in the string." Iterating over the characters of a string is a pretty common thing to do, and that extra branch is not free.
#define bsr(u) ((sizeof(int) * 8 - 1) ^ __builtin_clz(u))
#define ThomPikeByte(x) ((x) & (((1 << ThomPikeMsb(x)) - 1) | 3))
#define ThomPikeMsb(x) (((x)&0xff) < 252 ? bsr(~(x)&0xff) : 1)
void ThompsonPikeDecoder(const char *s) {
unsigned c, w = 0, t = 0;
do {
c = *s++ & 0xff;
if (0200 <= c && c < 0300) {
w = w << 6 | c & 077;
} else {
if (t) {
printf("%04x\n", w);
t = 0;
}
if (c < 0200) {
printf("%04x\n", c);
} else {
w = ThomPikeByte(c);
t = 1;
}
}
} while (c);
}
That code generalizes to any 32-bit number. It's the full superset of intended behaviors including things like stream synchronization. It'll even decode numbers that were arbitrarily banned by the IETF e.g. \300\200. Most importantly, since it doesn't require a validation pass beforehand, does that mean it goes faster than the OP's code in praxis? D:1. Your code drops naked continuation bytes entirely.
2. If a valid multibyte character is followed by extraneous continuation bytes it will not decode correctly.
3. Using the wrong number of continuation bytes is a giant mess and can give you arbitrary numbers in multiple ways.
The second one especially is a behavior that's hard to defend.
This seems like a premature optimization, unless you're only iterating over tiny strings.
While I acknowledge it would be unusual to provide a library API for finding the value of the last code point specifically, most string libraries provide something like a "codePointAt" function to return the code point at a specific offset. You'll run into the same problem with DECODE_UTF8 there.
int32_t
code_point_at(const char *s, size_t len, size_t offset) {
const char *end = s + len;
while (s < end && offset > 0) {
// Move forward by 1-4 bytes depending on header byte value.
s += UTF8_SKIP(s);
offset--;
}
if (s >= end) {
return -1;
}
return DECODE_UTF8(s); // BOOM! May read beyond string end
}Sure, and if you do strlen(s) without checking if there is actually a NUL first then:
> finding the end of a string may place you into invalid memory if there are missing trailing bytes.
As was originally said. This doesn't mean it's impossible to make such functions but <rest of original comment>.
>> finding the end of a string may place you into invalid memory if there are missing trailing bytes.
You're really losing me here. Either you have a NUL terminated string or you already know the length of the string or you have no way of determining the length of the string that exists in memory that it is safe for you to access and what you have is garbage.
Just because data is "valid UTF-8" doesn't make it safe to read. If the end of your string isn't marked and you don't know how long the string is already, it's over.
Again, it's doable, or you can validate the string once (including more than "it's valid encoding") and make a ton of safe assumptions later instead of each thing that interacts with the string. UTF or not.
Then your problems have nothing to do with invalid UTF-8, because that issue occurs even in pure ASCII.
A function that counts the number of characters/codepoints or maybe decodes utf8 -> ucs4, may wish to operate on entire codepoints instead of always checking for potential end of buffer mid-decoding. Or perhaps avoid handling of invalid codepoints at all.
That sounds like the sort of reasoning that results in compiler bugs like:
thing_t* nextp=p+1;
// other declarations for this function
if(!p) abort_thing("null pointer"); // optimised out
/* CVE-20XX-#####: user-space code can compromise kernel $THING if it mmap()s $STUFF at address zero */Have you ever met a use case for the number of codepoints in a unicode string?
(One that didn't involve the moral equivalent of "because codepoints are what Twitter counts"?)
I feel like you're hinting at something. What is it?
But wait, why are you converting from utf-8 to utf-32? To iterate through? You don't need to count your codepoints for that, you can just advance through your utf-8 linearly. To index into it? You might as well index into byte offsets—sure, that makes some possible indices invalid, but so does the finite length of your buffer.
And what are you doing with the codepoints anyway? utf-32 turns out to be profoundly useless for pretty much every concrete use case, because its whole appeal, that its code units map 1:1 to code points, presupposes that there's something you'd want to do with the codepoints.
There isn't. Code points are practically orthogonal to all of: user-perceived characters, user-perceived graphical symbols, grapheme clusters, text width. They combine arbitrarily in ways that you cannot infer without knowing the semantics of every possible code point. There is no linguistically sensible text operation you can do at the codepoint level. They are completely useless to you.
The worst part is, there are a lot of operations you can do with codepoints that seem, to cursory inspection, to do something you want. But they're wrong, all of them, necessarily so, because you're working at a level of abstraction that does not have the information needed to do just about anything correctly, but the cases that are broken are ones where none of your developers even know what the correct result should be.
Even measuring the width of rendered text is done to make it fit some layout.
But there are a lot of reasons to validate utf-8. The article mentions security vulnerabilities, and as a langsec devotee, I approve of always parsing inputs to prevent that sort of thing.
But even if you're reasonably confident that an invalid string isn't going to power a weird machine in your codebase (and are you, really?), validating is useful, because at that point you can trust algorithms that work on codepoints to actually work on codepoints, and don't have to constantly insert logic to detect malformation and re-synchronize your read, when all you want to be doing is working with your string.
An even simpler reason: It's easy to write something to work on some utf-8, and be quite sure that, for your application, it's going to be fed valid utf-8 or 'close enough'. And then some wire gets crossed and you're feeding it an mp3 or something.
Of course there are situations where you want to make do with your input and try to interpret it the best you can. Sometimes you need to be lenient in what you accept [0]. That's how I'd expect, say, a web browser to operate.
Default to not accepting invalid input. Just fail fast.
[0]: Be liberal in what you accept, and conservative in what you send - Jon Postel
The best way to do this is to TURN the input into valid UTF8 though. Like replacing any uninterpretable-as-utf8 bytes with the unicode replacement character (U+FFFD).
Which is what most well-behaved software I've seen does with bad UTF-8.
I work with lots of text processing, including processing old legacy files etc. Just refusing to go forward with bad input ("just fail fast") would work a lot less well.
Can you imagine if you tried to open up something in UTF8 in a text editor, and it just refused to open it, telling you nothing but "Bad UTF8" or maybe "Bad UTF8 at byte 19734"? Not super helpful.
But yeah, you want to TURN it into valid UTF8 as early in the pipeline as possible, while logging/flagging/notifying what was done. You definitely don't want to let the bad UTF8 just continue through the pipeline like that unless you really really have to, that way lies madness.
I think text editor is a good example where you need to be lenient. Furthermore, in that use case, any unnecessary modification of the data in any way is a bad idea. This includes replacing invalid code points with something else, unless they're edited by the user. The data might not even be UTF-8 in the first place.
Terrible advice, with a long history proving why it's terrible, starting with IE6.
Even valid UTF-8 can turn out to be fairly unusable because it requires unimplemented features such as some corner case of bidirectional text, unexpected combinations of diacritics, or just plain characters that were added to the Unicode standard after your font was designed, so leaving validation till the point of rendering can be a good strategy, and it also means that your validation will be even more ridiculously fast in the case where the string never gets rendered at all.
I agree that usually you should often allow and preserve unknown code point sequences, because rendering degrades gracefully and Unicode keeps changing. But the algorithm this article is talking about checks that a UTF8 string is ‘valid’ in the context of the raw byte encoding format. These checks are much more simple, and they’re needed in basically any software that interacts with strings.
I expect the I/O facility of a modern programming language to either reject bad UTF-8 or "quarantine" it somehow (map the bad bytes to a safe representation according to some documented rules).