I think this is very slow for some reason. I'll bookmark this and try on an A100 and a H100 box and then will try with multigpu on larger data.
My hope with this blog was to rile up other CUDA enthusiasts. Making them wanna bring in bigger and better hardware that I don't have access to.
+ Someone in the CUDA MODE community got 6 seconds on a 4090.