CPUs use significantly more power to perform the same amount of computation that a GPU does, because they're optimized for different workloads.
GPU input programs can be expensive to switch, because they're expected to change relatively rarely. The vast majority of computations are pure or mostly-pure and are expected to be parallelized as part of the semantics. Memory layouts are generally constrained to make tasks extremely local, with a lot less unpredictable memory access than a CPU needs to deal with (almost no pointer chasing for instance, very little stack access, most access to large arrays by explicit stride). Where there is unpredictable access, the expectation is that there is a ton of batched work of the same job type, so it's okay if memory access is slow since the latency can be hidden by just switching between instances of the job really quickly (much faster than switching between OS threads, which can be totally different programs). Branching is expected to be rare and not required to run efficiently, loops generally assumed to terminate, almost no dynamic allocation, programs are expected to use lower precision operations most of the time, etc. etc.
Being able to assume all these things about the target program allows for a quite different hardware design that's highly optimized for running GPU workloads. The vast majority of GPU silicon is devoted to super wide vector instructions, with large numbers of registers and hardware threads to ensure that they can stay constantly fed. Very little is spent on things like speculation, instruction decoding, branch prediction, massively out of order execution, and all the other goodies we've come to expect from CPUs to make our predominantly single threaded programs faster.
i.e., the reason that GPUs end up being huge power drains isn't because they're energy inefficient (in most cases, anyway)--it's because they can often achieve really high utilization for their target workloads, something that's extremely difficult to achieve on CPUs.