A GPU is much more efficient than an FPGA for what a GPU does. On the other hand for applications for which the set of primitive operations implemented in hardware by a GPU is not a good fit, an FPGA can be much more power efficient than a GPU.
For applications that involve a massive amount of computations with FP32, FP16 or BF16 numbers, for which GPUs have special hardware execution units, i.e. for training and for inference with non-quantized models, there is no chance for an FPGA to be more efficient.
If the GPU is recent enough to have good support for more heavily quantized data types, e.g. INT8, FP8, NVFP4 etc. an FPGA also does not have chances to be competitive.
An FPGA could be more efficient than a GPU if either it is some special AI-oriented FPGA, which instead of having traditional arithmetic units oriented for DSP applications, has execution units implementing the quantized data types popular in ML/AI, or if it implements inference using some new not yet standardized data type, for which GPUs do not have dedicated support yet.
The more application specific you get, the smaller the total volume of chips. The very nature of application specifity ruins the economics of ASICs.
Every time someone tells me an ASIC is more energy efficient I'm thinking, you just ruined the business case. The vast majority of application specific designs are not economically viable unless you use FPGAs to implement them.
Sure, and that is the definition of niche. FPGAs are niche.
You could have made a good point that an application specific design that already requires an FPGA might want to now also have a local LLM, so putting the LLM right on the FPGA might be the most expedient option in that case.
Other than that, I don't think people are reaching for FPGAs to do LLM training or inference in general because I don't think it can be cost effective vs other options.
Anything else is added and unintentional benefits.
FPGAs win against MCUs in terms of performance and they only lose in terms of static power consumption, not on performance per watt.
Also, there was a company doing LLM inference on FPGAs and their entire selling point was that they were more energy efficient than Nvidia so you can add more FPGAs onto the same rack.
Finally, the ASIC Vs FPGA battle is kind of meaningless because the moment you decide to reprogram your FPGA for any reason, ASICs aren't even in the same market anymore.
Then there is the fact that FPGAs tends to have insane amounts of SRAM bandwidth compared to most chips.