1) Try to load things sequentially don’t load from offset 32, then 0, then 16.
2) Don’t use the loop instruction
3) this code:
dec n
cmp n, 0
jne 1b
Can be replaced with: sub n, 1
jnz 1b
Which will actually be executed by the processor as a single instruction (macro-op fusion).4) Interleave your expensive instructions with less expensive ones. Try to interleave multiple dependency chains to let the processor see more of the parallelism.
Your two divisions one after another will be limited by available execution units capable of doing the divide.
5) Lastly align your loop labels to 16-byte offsets. The assembler will do this for you with the ALIGN directive.