In the current final code, each loop has 3 independent FMAs. And in each loop iteration, the 3 FMAs are dependent on the FMAs from the previous iteration.
So one thing that could speed up the computation is to do 6 independent FMAs per loop iteration instead of 3. This can be done as follow. Instead of doing cumulative sums from 0 to n-1, one could split the computation of the cumulative sums from 0 to n/2-1 and from n/2 to n-1. So the main loop goes only from 0 to n/2-1, and at the k-th loop iteration, we do the 2x3 FMAs corresponding the entries of indices k and n/2+k. And at the end we add the cumulative sums.