For starters it looks like they are talking about integer operations (I only skimmed the paper and it mentions an ALU, not an FPU), whereas my GPU numbers below are single precision floating point numbers. So it is apples vs oranges.
So, a modern 14-16nm GPU like the Tesla P100 or RX 480 does about 5 to 10 trillion ops/sec at 200-300 W, and approximately 30 pJ/op. So GPUs are about 5x less power efficient. The paper authors did a good job. However GPUs are not optimized for absolute power efficiency, but mostly for performance per $.