I'm kinda curious if this is a data-size that might be better on CPU rather than GPU. GPUs of course are better compute, but GPU-registers tend to cap out at ~1024 bytes or so, and a number of this size probably would have to be stored in RAM (plus all the associated data structures in ... whatever algorithm happened here).
If this were small enough, I can kinda imagine CPUs being better due to caching effects (L1 and L2 are incredibly fast on CPU, and GPUs kind of miss those tiers in some sense since GPUs focus on register space).
It could just be that this eliptical-curve prime-testing program is currently a CPU kinda function. But I'm curious if this algorithm would be (or wouldn't be) suitable for GPU compute, with more TFLOPS and more VRAM bandwidth for cheaper.
EDIT: A 20kB number would fit inside of __shared__ memory on GPU, so perhaps various operations (multiplication and/or division??) could be parallelized onto a 256-wide/1024-wide workgroup (~1024 to 4096 bytes per clock tick, or ~5 to ~20 clock ticks to perform most kinds of operations with the 20kB number on one compute unit). It really depends on the kind of algorithm going on here, as well as what needs to be searched for. I did a brief read into ECPP testing, and I honestly don't get anything they're saying in the paper (A. O. L. ATKIN AND F. MORAIN, 1993. "ELLIPTIC CURVES AND PRIMALITY PROVING").