Indeed, I'm not a fan of C++28 with virtual functions, exceptions, new/delete and whatever in my GPU code :P
With a parallel CPU implementation I have pretty few debugging issues, and most importantly, I definitely enjoy having excellent performance on Intel, AMD and Nvidia GPUs, on Windows, Mac and Linux.