However, from a brief glance, it does not appear to have created a closed form solution. Instead, it contains a single loop:
.L4:
addl $1, %edx
paddd %xmm1, %xmm0
paddd %xmm2, %xmm1
cmpl %edx, %eax
ja .L4
which seems to be using a SIMD instruction (paddd[1]) that adds does 4 32-bit integer additions in parallel.After this loop, it does some "housekeeping" (read, something I don't understand) before proceeding to an unwound version of the last iterations of the loop:
leal 4(%rdx), %ecx
addl %edx, %eax
cmpl %ecx, %edi
jl .L2
addl %ecx, %eax
leal 8(%rdx), %ecx
cmpl %ecx, %edi
...
jl .L2
addl %ecx, %eax
leal 28(%rdx), %ecx
cmpl %ecx, %edi
jl .L2
addl %ecx, %eax
addl $32, %edx
leal (%rax,%rdx), %ecx
cmpl %edx, %edi
cmovge %ecx, %eax
ret
Where .L2 is just: .L2:
rep ret
I assume that this is just some form of return, but the documentation I could find [2] seems to suggest that rep is a prefix for string operations, which doesn't make sense.[0]https://pastebin.com/raw/Y55gQG7p
[1] http://x86.renejeschke.de/html/file_module_x86_id_226.html