More optimistically something along the lines of hands-off offloading, where if the scheduler sees enough of the same type of calculation (e.g sparse matrix multiplication) it can reconfigure the fpga and offload.
More optimistically something along the lines of hands-off offloading, where if the scheduler sees enough of the same type of calculation (e.g sparse matrix multiplication) it can reconfigure the fpga and offload.
FPGAs excel in parallelizing very simple operations. One example where FPGAs are good is where you might want to look for a specific string in a network packet. Because FPGAs are electronic circuits you can replicate the "match a byte" logic thousands of times and have those comparisons all run in parallel and the results combined with AND & OR gates into a yes/no decision. (I think the HFT crowd do this sort of thing to preclassify network packets before forwarding candidate packets up to software layers to do the full decision making).
The moment you are memory heavy the FPGA loses it's edge because CPUs and GPUs are designed around hiding memory latency really well and they overall have higher memory bandwidths.
I don't think that's a good characterization. FPGA are good at what they end up being programmed for. And in the final analysis, everything a chip does is broken down to simple operation.
The FPGA selling point has always been around perf/W and efficiency with regard to a "set of task". An ASIC will always be faster for a specific task, and CPU will always be faster on average on everything. However, when considering say "compression" or "check-sum" as general class of algorithm, and FPGA with a set of predefined configuration could be better and cheaper.
FPGAs being selected for performance per watt was only a fairly recent phenomenon, when they were deployed on semi large scale as password crackers/cryptocurrency miners.
Their real strength is ultimately real time processing (DSP or networking), with reconfigurability often quite valuable for networking applications. For DSP applications it's usually because a MOQ of custom silicon can't be justified.
For instance, SHA-3 is quite slow on x86 CPUs, which, unlike the recent ARM CPUs, do not have any hardware instructions to accelerate it.
If there had been an included FPGA, it would have been easy to implement a very fast SHA-3, as its hardware cost is very low.