In the context of gmp, people write architecture-specific assembly for the inner loop anyway.
Besides that, you raise good points on sources of complexity. I’m waiting for the benchmarks once such developments have been incorporated. Everything else is guesswork.