I don't know about this code in particular but if you're mostly doing math you tend to run out of architectural registers before you achieve the full throughput of the instructions you care about :(
If the lack of registers is really the bottleneck, a variant might to use both zmm registers for some cumulative sums and ymm registers for some others. In this case the speed up might be less spectacular though.
Edit: actually, I just discovered that zmm registers overlap the ymm registers, so the only registers left are the ones from the FPU.