That makes sense. I did observe significant speed up using this technic for evaluating a polynomial on several values. I used Hörner algorithm and indeed the number of registers is very small in this case.
If the lack of registers is really the bottleneck, a variant might to use both zmm registers for some cumulative sums and ymm registers for some others. In this case the speed up might be less spectacular though.
Edit: actually, I just discovered that zmm registers overlap the ymm registers, so the only registers left are the ones from the FPU.