Microsoft Bets Its Future on a Reprogrammable Computer Chip
wired.com
wired.com
Intel openly acknowledges it. From the article:
"Microsoft’s services are so large, and they use so many FPGAs, that they’re shifting the worldwide chip market. The FPGAs come from a company called Altera, and Intel vice president Diane Bryant tells me that Microsoft is why Intel acquired Altera last summer—a deal worth $16.7 billion, the largest acquisition in the history of the largest chipmaker on Earth. By 2020, she says, a third of all servers inside all the major cloud computing companies will include FPGAs."
Altera strikes me as an odd choice though. I'd have thought Intel would buy out Xilinx, the industry leader, instead.
Intel had been fabbing (some of?) Altera's chips for a few years before the acquisition, so from this angle it makes more sense than Xilinx. As to why Altera had this partnership and not Xilinx, who knows. Perhaps being in second place motivates you to shake things up.
Yes, FPGAs are a pain to program now but so were general purpose CPU's in the early days. Given time and innovation that will change. We are headed towards a time where most algorithms will be run off specialized chips rather than a general purpose CPUs. It's only a matter of time.
Without great improvement and opening of the tools, FPGAs aren't going to be a general-purpose accelerator. It'll be a question of building one specific piece of firmware for one specific project and then deploying it semi-permanently with minor revisions across the cloud.
Wake me up when Visual Studio has an FPGA backend.
Vendor could pull a IBM or SGI with little boards with high-end CPU's + FPGA's or S-ASIC's for various coprocessors. Not sure if it's ultimately best idea but surprised I haven't seen anyone try it. Wait, just looked up their press releases and it seems they're doing something through OpenPOWER:
http://www.easic.com/easic-joins-the-openpower-foundation-to...
The thing with using FPGA's in systems (they're great for low-volume, high-priced items where ASICs would be too costly) is they end up just emulating logic which could be more cheaply implemented as actual execution units (as many modern FPGA's things like cache, ROM, ALUs, etc.). That is, it's expensive flexibility that isn't really all that useful. Sure you could reconfigure a "computer" from doing database things to suddenly add more GPU cores to play games, but how useful or power/cost efficient would that be? Sure it's nice to cut down on ASICs and upgrading them after the fact, but it seems like more like category development than offering practical advantages to solve a real problem. Maybe a super-fast HPC-on-a-chip would be possible, but I don't see that we're storage or compute constrained, however we maybe bandwidth and latency constrained in terms of shrinking clusters to a single rack of ridiculously power-hungry reprogrammable chips.
Instead of infinitely customizable, arbitrary logic, you might have a crap-ton of simplified RISC cores with some memory and lots of interconnected bandwidth or something in-between FPGA and MPPA.
https://en.wikipedia.org/wiki/Massively_parallel_processor_a...
It's hard to tell from the article though, I'm just guessing.
Edit: Though I think it was not mentioned in the article, Microsoft and Intel/Altera have indeed gone this route in no small part due to the empirical death of Moore's Law (which has gone much discussed on HN over the past few years).
[0] https://en.wikipedia.org/wiki/Reconfigurable_computing#Parti...
Thank you for the info though!
> ... in the coming weeks, they will drive new search algorithms based on deep neural networks—artificial intelligence modeled on the structure of the human brain—executing this AI several orders of magnitude faster than ordinary chips could.
Sure, GPUs deliver impressive raw performance. To be useful, the task must benefit from massively parallel hardware. GPU hardware works fantastic for shading polygons, training neural networks, or raytracing. For compression and encryption algorithms however, GPUs aren’t terribly good.
Another reason is while a GPU delivers impressive bandwidth on parallel-friendly workloads, it’s usually possible to achieve lower latencies with FPGA. An FPGA doesn’t decode any instructions, and its computing modules exchange data directly.
Citation needed.
Maxwell Jetson TX1 is claimed to achieve 1TFlops FP16 at <10W, and soon to be released Pascal based replacement will probably be even more efficient.
The TX1 power consumption including DRAM and other subsystems peaks 20-30W. Typical usage is 10-15W if you're running anything useful.
That 1 TFLOP counts a FMA instruction as 2 flops - while accurate and useful for say dot products - for other workloads the throughput will be half of this number.
As an example of an FPGA performing significantly better than the TX1 is DeepPhi [0].
While not the TX1 vs FPGA result you want, this is very close. For example they aren't using the latest FPGA or GPU, and are not using TensorRT on the GPU and on the FPGA side they are using fatty 16-bit weights on an older FPGA rather than newer stuff you can do with lower precision (which improves the efficiency of the FPGA having more high speed RAM collocated with computation vs GPU which is primarily off-chip).
If you want to learn more about this stuff, I suggest a presentation by one of Bill Dally's students (chief scientist at NVIDIA): http://on-demand.gputechconf.com/gtc/2016/presentation/s6561...
I'm not saying you're wrong, just that to make a convincing claim that FPGAs are more power efficient than GPUs, one needs to do an apples to apples comparison.
And of course, let's not forget about price: Zynq ZC706 board is what, over $6k? And Jetson TK1 was what when released, $300? If you need to deploy a thousand of these chips in your datacenter, to save a million per year on power, you will need several years to break even, and by that time, you will probably need to upgrade.
It just seems that GPUs are a better deal currently, with or without looking at power efficiency.
Here's a paper about it: https://www.microsoft.com/en-us/research/wp-content/uploads/...
4.4, Feature Extraction: "The first stage of the scoring acceleration pipeline, Feature Extraction (FE), calculates numeric scores for a variety of “features” based on the query and document combination. There are potentially thousands of unique features calculated for each document, as each feature calculation produces a result for every stream in the request—furthermore, some features produce a result per query term as well. Our FPGA accelerator offers a significant advantage over software because each of the feature extraction engines can run in parallel, working on the same input stream. This is effectively a form of Multiple Instruction Single Data (MISD) computation."
Also annoyingly this article does not link to the paper about this, which also explains it better than the article. I recall this one from MS about how they use the FPGAs in Bing; was pretty impressed by it at the time. https://www.microsoft.com/en-us/research/publication/a-recon...