Show HN: Charcoal – Faster utf8.Valid using multi-byte processing without SIMD
github.com
github.com
Since that's the point of comparison, what's the Go standard library's strategy? Is it inherently slower than this or does it behave better for different scenarios?
How does this compare to using simd (like simdutf8)?
[1] https://github.com/sugawarayuuta/charcoal/blob/main/docs/ind...
It is basically using 64 bit integer operations to check 8 bytes at a time.
UTF-32 always uses 32 bits per character, so you can easily paralelize it because you can easily tell where a character starts in a bit stream.
UTF-16 can have 16 or 32 bits, so if you have a sequence 2 characters of 16-32 bits, or in UTF-8's case, a sequence of 3 characters that are 8-16-16 bits, they won't fit neatly in a box. You can't just say char_ptr + 32 to land on the 32nd character, making it much harder to paralelize.
In addition, UTF-32 can cause issues on the network due to byte ordering (little endian vs. big endian). There’s really two encodings, UTF-32LE and UTF-32BE.
For all of these reasons, the world has largely standardized on UTF-8, plus for historical reasons UTF-16 in the JVM, CLR, Windows, and JavaScript.
Do you have any plans on proposing to upstream some of your work?
Curious to know if there are any known caveats?
Otherwise, go for it. May seem daunting, but it's not. I've submitted a couple of proposals and patches; pretty happy with the process.
If you have invalid utf-8, go fix whatever code it is that is producing invalid utf-8...
Not all sequences of bytes are valid UTF-8. A UTF-8 decoder should be prepared for:
invalid bytes
an unexpected continuation byte
a non-continuation byte before the end of the character
the string ending before the end of the character (which can happen in simple string truncation)
an overlong encoding
a sequence that decodes to an invalid code point
https://en.wikipedia.org/wiki/UTF-8But that's not what I'm answering to, and I was not taking sides so I don't know why you would reply to my comment with this. I was just clarifying something. I'm answering a comment that might have misread what they answered too, that's it.
OP was saying not to validate as part of your normal pipeline. They weren't saying not to validate ever, such as while poking around or testing, which is how you could still figure out that bad utf-8 exists and where it is coming from.
> And even if you don't validate on the hot path, it doesn't hurt to do it faster, so I don't understand why one would be against a more efficient computation.
It's not about being for or against the speed of calculation. It's that removing unnecessary calculations is even faster. So questioning the need is good.
If 'hot path' was unclear and you think I'm making things up, let me try a different wording. As a hypothetical: If I don't think something needs to be validated in production, I might question a library dedicated to doing it really fast, and say your producers should be fixed instead. Me saying that doesn't mean I think you should be stumbling around blind, never testing those producers.
Presumably you think you should always be able to trust anyone who sends you JSON?
I mean Go already does a bunch of UTF-8 validations in its JSON encoder/decoder [1]. No need to make those faster.
I said hypothetical, and I did not use the word JSON.
If there's a workload that makes sense to optimize because it's used in production, it's UTF-8 validation.
So why are we even arguing whether UTF-8 validation is necessary/useful in the "hot path" or in "production"?
If you're asking about OP's idea, then I agree that validating the UTF-8 inside of JSON is needed sometimes, though it's not required in all situations. But I wasn't commenting on that, I was commenting on the specific objection of "how do you know if you don't check?", because I don't think that specific objection is very strong.
Ah, makes sense. Thanks!
Well, the implication is not wrong. If you validate only during testing, you never validate some inputs, the ones that matter: user input.
I now see what you mean by hot path, I was imagining something like validating asynchronously, later, when not serving a request for instance.
If OP wanted to share with us the idea of moving validation to another step in the process, giving a case where it would be applicable, that would have been fine but why question an optimization in the first place? Why not both? It just feels like they are rejecting stuff for the sake of rejecting stuff and I assume this is the reason they got downvoted.
UTF-8 validation is something that happens constantly. And even if it's was not, this optimization is probably an interesting technical feat, and we are in a Show HN post too, rejecting stuff like this is just not going to fly on HN. You can question things but you need to be insightful and nice.
You make good points but OP didn't make them.
Anyways.
And obviously you don't want to spend more time doing this than necessary.