> That's why I call it big-number calculation: manual carry/overflow handling, larger-than-register size operands, tedious indexing/bug prone if coding from scratch, ... but of course you can still argue that it is not "big."
Its not a general big-integer implementation which can be arbitrarily sized. Since you know all results fit inside of 192 bits, and that all operands are less than 64-bits, you can easily code for this special case.
192-bit x 64-bit is way easier to debug and test than 192-bit x 192-bit. Its probably the bulk of your work, but I don't foresee any major complication here. Its just grade-school arithmetic.
struct cube{
long long val[3];
}
cube multiply(cube x, long long y){
cube toReturn;
toReturn.val[0] = x.val[0] * y;
toReturn.val[1] = x.val[1] * y + __umul64hi(x.val[0] * y);
toReturn.val[2] = x.val[2] * y + __umul64hi(x.val[1] * y);
return toReturn;
}
This is grade-school arithmetic level. This isn't very difficult at all. The only thing you needed to know is that CUDA has an intrinsic to take the top 64-bits of a 128-bit multiplication.
Without any if-statements, the GPU can calculate the above across different threads without any divergence. You can likely optimize this further with MAD instructions, but the above would be a good "first step" towards this problem.
As I said earlier: this wouldn't be an "easy" port, but it would be possible, and it would likely be very efficient. I'm not going to think through all the edge cases, but the bulk of the algorithm seems very easy to do efficiently on the GPU architecture to me.