Interesting but his scalar code is slow. When you care about performance, better to implement such algorithm so it reads bytes one by one, but move blocks with memmove when switching from write to skip state.
Pathological case (skipping every other character) is slightly slower, but on real data it’s much faster overall.
Unfortunately I don’t have AVX512 hardware so I can’t test.