I've had a similar hypothesis, but it didn't work out for me, only bloating the codebase. Please lmk if you have a different outcome :)
If the lack of registers is really the bottleneck, a variant might to use both zmm registers for some cumulative sums and ymm registers for some others. In this case the speed up might be less spectacular though.
Edit: actually, I just discovered that zmm registers overlap the ymm registers, so the only registers left are the ones from the FPU.