Anyone have any ideas about which open source codebases UTF-8 validators exist in?
Anyone have any ideas about which open source codebases UTF-8 validators exist in?
Rust's std library -- the canonical way to read a text file to a string is to (implicitly) use `std::str::from_utf8`. If I remember correctly, the current implementation doesn't use SIMD specificially but will of course contain vectored instructions if the compiler can select them on the platform you target.
I did a comparison with another SIMD based implementation last year. Maybe it's time to update it: https://github.com/killercup/simd-utf8-check
The whole algorithm basically has to be redesigned from scratch to introduce vectorization: something far beyond the capabilities of the compiler.
That said, I'm happy to help anyone get it integrated into their codebase if they need it. There's generally other concerns in large projects, though, like portability, that I don't care as much about.
For more mundane uses like opening a big text file, you often want to tolerate invalid bytes and show a replacement character. You could use the technique, but you'd have to modify it to work for that usage.
Maybe this could be applied to XML, but that's quite a behemoth of a standard--I really wouldn't be surprised if there was a way to switch encodings mid-stream. I have no idea though...
You can't normally blindly parse XML as UTF-8. The encoding has rules for detecting the character set.
> A lot of the software ... don't validate upfront
They almost always should be validating all input up front. Deferring validation tends to become many different pieces of code all informally parsing fragments of the input that are needed locally. Without upfront validation of the complete unit of input, the resulting fragmented parsers are just a weird machine waiting to be programmed by a malicious attacker.
For a much better explanation, I strongly recommend Meredith and Sergey's 28c3 talk[1] about The Science of Insecurity.
> you often want to tolerate invalid bytes and show a replacement character
While this validator wouldn't be useful in that situation, it is still important to validate the input upfront. When showing replacement characters, the "invalid bytes" that will be supported with a replacement character should be formally defined and added to the validation grammar, because they are no longer "invalid", but instead will be handled as a special case.
Postel's robustness principle shouldn't be used as an excuse to skip validation. Being "liberal in what you accept" should still be well-defined and validated.
[1] https://media.ccc.de/v/28c3-4763-en-the_science_of_insecurit...
How unicode replacement characters should work is already well defined. Just look it up.