Here's a paper about it: https://www.microsoft.com/en-us/research/wp-content/uploads/...
4.4, Feature Extraction: "The first stage of the scoring acceleration pipeline, Feature Extraction (FE), calculates numeric scores for a variety of “features” based on the query and document combination. There are potentially thousands of unique features calculated for each document, as each feature calculation produces a result for every stream in the request—furthermore, some features produce a result per query term as well. Our FPGA accelerator offers a significant advantage over software because each of the feature extraction engines can run in parallel, working on the same input stream. This is effectively a form of Multiple Instruction Single Data (MISD) computation."
> ... in the coming weeks, they will drive new search algorithms based on deep neural networks—artificial intelligence modeled on the structure of the human brain—executing this AI several orders of magnitude faster than ordinary chips could.
Sure, GPUs deliver impressive raw performance. To be useful, the task must benefit from massively parallel hardware. GPU hardware works fantastic for shading polygons, training neural networks, or raytracing. For compression and encryption algorithms however, GPUs aren’t terribly good.
Another reason is while a GPU delivers impressive bandwidth on parallel-friendly workloads, it’s usually possible to achieve lower latencies with FPGA. An FPGA doesn’t decode any instructions, and its computing modules exchange data directly.
Citation needed.
Maxwell Jetson TX1 is claimed to achieve 1TFlops FP16 at <10W, and soon to be released Pascal based replacement will probably be even more efficient.
The TX1 power consumption including DRAM and other subsystems peaks 20-30W. Typical usage is 10-15W if you're running anything useful.
That 1 TFLOP counts a FMA instruction as 2 flops - while accurate and useful for say dot products - for other workloads the throughput will be half of this number.
As an example of an FPGA performing significantly better than the TX1 is DeepPhi [0].
While not the TX1 vs FPGA result you want, this is very close. For example they aren't using the latest FPGA or GPU, and are not using TensorRT on the GPU and on the FPGA side they are using fatty 16-bit weights on an older FPGA rather than newer stuff you can do with lower precision (which improves the efficiency of the FPGA having more high speed RAM collocated with computation vs GPU which is primarily off-chip).
If you want to learn more about this stuff, I suggest a presentation by one of Bill Dally's students (chief scientist at NVIDIA): http://on-demand.gputechconf.com/gtc/2016/presentation/s6561...
I'm not saying you're wrong, just that to make a convincing claim that FPGAs are more power efficient than GPUs, one needs to do an apples to apples comparison.
And of course, let's not forget about price: Zynq ZC706 board is what, over $6k? And Jetson TK1 was what when released, $300? If you need to deploy a thousand of these chips in your datacenter, to save a million per year on power, you will need several years to break even, and by that time, you will probably need to upgrade.
It just seems that GPUs are a better deal currently, with or without looking at power efficiency.
It's hard to tell from the article though, I'm just guessing.
Edit: Though I think it was not mentioned in the article, Microsoft and Intel/Altera have indeed gone this route in no small part due to the empirical death of Moore's Law (which has gone much discussed on HN over the past few years).
[0] https://en.wikipedia.org/wiki/Reconfigurable_computing#Parti...
Thank you for the info though!