This looks like a nice approach to making wasi-libc faster. Could you submit these changes upstream?
I've also only really tested wazero. I can't know for sure that this is a straight improvement for other runtimes and architectures.
For instance, the code delays using wasm_i8x16_bitmask as much as possible, because on Aarch64 it can be slower than not using SIMD at all, whereas it's plenty fast on x86-64.
One of the nice things about Go is how much that's a solved issue out of the box, compared to almost everything else; certainly compared to C.
Pinging them in an issue: https://github.com/WebAssembly/wasi-libc/issues/580