OK so I coded up the first one and the last one in C and compiled on my mac using whatever gcc links to. I found 1e7 points in the sphere:
first routine: 0.521s
last routine: 0.850s
[edit]: C code here: https://pastebin.com/zH0BehSM
OK so I coded up the first one and the last one in C and compiled on my mac using whatever gcc links to. I found 1e7 points in the sphere:
first routine: 0.521s
last routine: 0.850s
[edit]: C code here: https://pastebin.com/zH0BehSM
After fixing that, this way is faster for me. Some trig, but no branches.
But what if one used optimizations like lookup tables and trigonometric approximations instead? Could they end up costing less cycles than three rands? Maybe not LUTs, maybe other trickery someone knows of?
And what about the cost of those three rands? Could one use a fast xorshift with good enough results here? Does one even need thee random vectors? Maybe two + modulo trickery will do?
(rand()/(double)RAND_MAX) etc.
Instead assign some name to 1.0/(double)RAND_MAX, and multiply.