Pretty sure it's optimisations.
Even small things like converting multiplications followed by additions to FMA will reduce the operation count.
Add to that constant folding etc. and a speedup factor of ~3 is not so hard to imagine.
Even small things like converting multiplications followed by additions to FMA will reduce the operation count.
Add to that constant folding etc. and a speedup factor of ~3 is not so hard to imagine.
nvcc -arch=sm_86 prospero.cu -o prospero
cuobjdump -sass prospero | grep -E 'FFMA|FADD|FMUL|FMNMX|MUFU\.RSQ' | wc -l