- ASICs are too rigid and require high volumes to be profitable
- GPUs are too power hungry
- CPUs are not good for massively parallel processing
FPGAs are heavily used in industrial/military/aerospace applications
- ASICs are too rigid and require high volumes to be profitable
- GPUs are too power hungry
- CPUs are not good for massively parallel processing
FPGAs are heavily used in industrial/military/aerospace applications
FWIW most modern FPGAs use discrete DSPs anyway so you're not really getting the flexibility at that level.
I don't think the comment about process is really true, from what I can tell, Xilinx is only a few months behind the biggest SoC makers in terms of its process adoption, and is shipping 14nm parts currently. Not sure about Altera, but they are on Intel's process, which is bit ahead of the competitors anyway.
In terms of switching power, you definitely pay a penalty to have the reconfigurability in hardware, but on the other hand you don't have all the unused logic that you would on a GPU. I'd guess the comparative efficiency depends on the specific problem and specific implementation, but I don't have any numbers to back that up.
Also there's a ton of SRAM for the 4-LUT configuration so you're paying leakage costs there as well.
NVidia managed to get it right about year and half ago. Before that their gates leaked power all over the place.
The LUTs on Stratix are 6-to-2, with specialized adders, they aren't at all that 4-LUTs you are describing here.
All in all, there are places where FPGAs can beat ASICs. One example is complex algorithms like, say, ticker correlations. These are done using dedicated memory (thus aren't all that CPU friendly - caches aren't enough) and logic and change often enough to make use of ASIC moot.
Another example is parsing network traffic (deep packet inspection). The algorithms in this field utilize memory in interesting ways (compute lot of different statistics for a packet and then compute KL divergence between reference model and your result to see the actual packet type - histograms created in random manner and then scanned linearly, all in parallel). GPUs and/or CPUs just do not have that functionality.
(I don't know if you can publicly get Stratix 10 devkits yet, but you can get an Arria at least.)
On the DSP side, you're using a ASIC DSP(can't change the width for instance) anyway on most modern FPGAs so you're comparing ASIC to ASIC at that point.
Of course you can get better price/GFLOP with GPUs + quicker time to market
GPUs have very particular cache and computation hierarchy which is not necessarily a best/good fit for all problems that are being thrown at them.
You can fuse more operations into DSP using FPGA and/or you can perform less operations per FLOP. One example is to avoid rounding and packing/unpacking when creating deep pipelines for floating point processes.
For open source hardware I doubt you'll see people shell out the cost of a used car to be able to match modern ~$200 GPUs.
I don't even what to know what Stratix 10 starts at.
This was admittedly a space certified radiation hardened chip. Still alarming when you had to pick it up and carry it somewhere
The Xeon Phi board https://www.cnet.com/products/intel-xeon-phi-coprocessor-712... is $4.2K..$5K
The Stratix V board will not consume more than 60W when used as PCI Express card. The requirements for Xeon Phi is at least 250W.
With a difference of ~200W, there will be difference in ~4.5kW/h per day or ~1600kW/h or $160 in hard cash per year (US average). Very probably more - getting rid of heat produced, etc.
Just being a SoC has another advantage: reduced part count, so you can fit more functionality onto a smaller PCBA.
Another commenter mentioned the Cyclone V and Zynq SoCs. My day job is in the telco industry. The equipment vendors we sell to are always pushing for higher and higher densities in their chassis and line cards, and it drives a lot of our design decisions. Chips like the Cyclone V and Zynq help a lot in achieving those aims. Our PCBAs have shrunk dramatically (with increasing functionality) over the last decade.
The continual pressure to reduce costs and power consumption also leads to choices like minimizing the other available system resources (e.g. clocks, memory bandwidth, RAM, flash), which can end up reducing the effectiveness of a GPU.
As with everything in engineering, there are many different problem domains, each with unique constraints dictating different solutions.