I was wondering if you like to contribute better scalar code? I'll be happy to include your (or anybody else) code and then compare different approaches.
Your AVX2 is better in benchmarks than Zach's AVX512VBMI, wow. I have to admit that looked at the code but got lost. :) Will need more time to digest it.
However, in despacing English texts and CSV, Zach's variant is still faster.