An ECP5 will sit on the order of ~100mW and you can clock those up to dozens of MHz. They can have multiple cores running in parallel (an ECP5 85k will fit dozens, probably well over a hundred RISC-V cores if you do your homework.) Even a laptop sitting at 10W is going to be orders of magnitude more power inefficient than this in terms of raw instructions-per-cycle-per-watt if you're emulating. That is not the best metric, necessarily, but there you go.
And since you mentioned perf-per-dollar -- ignoring soft CPUs, any deeply pipelined algorithm is very likely going to destroy price-comparable CPUs in terms of throughput e.g. you can do 16-to-32 bytes per cycle of AES on a dinky FPGA from 10 years ago for a few dollars, and at 50MHz you're doing 1.6GB/s, and people have been achieving this, or multiple times this, for 15+ years. Things like TDP are not a measure of "overall system design efficiency", it's a measure of thermal capacity, thermal budgets, and nothing more. (BTW, the only general purpose CPU that comes close to this number directly for AES is, like, Ice Lake, since VAESNI can turn out 16 bytes per cycle or whatever IIRC, but now you're well back into "multiple watts" territory on a multi-GHz CPU.) The reason people still use CPUs for these tasks isn't because they don't want better performance: it's because software has better agility and is easier to acquire and modify and distribute. You can have systems that are dozens of times more efficient than commodity ones for a wide variety of tasks, they will just be a pain in the ass to use, program, acquire, and build. You can figure out most of this with basic napkin math.
Stop thinking so much about individual components, and start thinking about global system design -- because the entire system has its own performance criteria that may vary drastically compared to an individual component within it.
> There are very very very few compute tasks where an FPGA solves a problem with better performance per watt than both a CPU and a GPU.
This is like stating "There are very few tasks where a car would do as well as a snowmobile." They aren't comparable for purpose. Hacker News is pop-culture-y so everyone thinks "the only thing that matters is a cool CPU running in a rack with a 7nm TSMC process that can run my Go application on Kubernetes that will disrupt The Market of Smart Toilets" or whatever they do day to day, and extrapolate from there. But I'd guess the vast majority (like, 85% or more) of FPGA field has literally nothing to do with this. A huge amount of work basically revolves around "just" interfacing with analog devices at pico/nanosecond level resolutions...
The quest for best perf-per-watt is one largely driven by datacenters and personal consumer electronics, which have both high volume and high yield, and where the largest challenges revolve around power, cooling, etc. Furthermore these systems run workloads that are largely general purpose "state machines" that use some memory and some CPU and some disk, etc, and need to try and hit a balance among all of these. There is a large amount of resource arbitrage going on. "A rising tide lifts all boats" in this case. But little of that applies in this field; people use older nodes and the same chips for 5-10+ years (or longer) straight because they need to deliver latency-sensitive solutions, customized hardware at low volume, "hardware glue" for various analog systems, highly specialized algorithmic solutions for the lowest total BOM cost, etc. They aren't aiming to replace the systems created by digital Silicon Valley software programmers.
There is a push to move FPGAs into the datacenter (see: Xilinx and their exploding revenue) but it's unclear if they will settle into specific niches or be used as supplementary devices or whatnot.