Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Learning?
nextplatform.com
nextplatform.com
GPUs are flexible and scalable when you don't know what the large-scale parameters of the network you want to build look like, and need a lot of them to do training. Let a fleet of cloud-based GPUs do the heavy-lifting of training and learning.
But then once training is over, an FPGA or even an ASIC could implement the trained model and run it at a crazy-fast speed with low-power. A piece of hardware like that would be able to handle things like real-time video processing of a DNN potentially. Very handy for things like self-driving vehicles.
https://rcpmag.com/articles/2016/10/10/microsoft-google-ai-s...
If the deep learning architecture stabilizes for a problem with sufficient market demand, seems like ASICs could be economical.
ADDENDUM: But fundamentally, in spaces like this, the underlying algorithms that can be accelerated are fairly simple. In most cutting edge AI these days the heavy lifting is performed by convolutional neural networks and the specialized silicon that works to speed up one set of convolutional neural network operations will speed up another just as well. Baking the network itself into the hardware shouldn't tend to be any better than loading it into specialized memory pools unless you get really exotic and do your neural network in analog electronics.
I hardly believe that this is economical.
At the very least they talk about omitting weights which are 0 in the synthesis.
Until vendors are willing to release bitstream details enabling open source tools and an vibrant ecosystem, applications will be limited.
That said, Lattice has recently started to push into these new areas, but they haven't been that successful. If they start to see more success, I think we will see open toolchains. Lattice also has the advantage of being able to lean on the open toolchain work done by people like Clifford Wolf [1]
It doesn't beat the commercial tools in area / performance right now, but dear god its so much more pleasant to work with!
Whoever has worked with FPGA knows that they are completely different to program than CPU/GPU.
They are not competing at all. You can't take some computer developers and have them work with FPGA. That'd be like taking a dude who knows XML and put him on optimizing C++ low level algorithms.
For starters, a FPGA doesn't run programs, it describes hardware components.
Hardware engineers can design FPGA-based hardware optimized for ML.
A second set of engineers then uses these boards/FPGA's just as they would GPU's.
They write code in whatever language to use them as ML co-processors.
This second group doesn't have to be composed of hardware engineers.
Today someone using a GPU doesn't have to be a hardware engineer who knows how to design a GPU. Same thing.
Bitcoin is a good example. I'm reasonably sure that FPGAs eclipsed GPU there because some "software person" called the right person in. And it certainly was a competition. Similar for the eventual move to ASIC.
If there's enough incremental benefit, it will happen.
[1]: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.50....
I also believe that today's FPGAs are more robust against these kinds of bugs. Because, even with some HDL code, these defects or cross-talk incidents could result in difficult to debug errors.
And if it is just about bang-for-buck and scale, FPGAs seem to be quite competitive.
The big win for them is a 10x reduction in power usage, since in a datacenter/cloud environment this is more important. Still at the research stage though.
Source: https://en.wikipedia.org/wiki/Pascal_(microarchitecture)
section 2.4 Chips claims the Titan XP uses the GP102 chip, and section 3 Performance gives the speed for computing with 16-bit floats.
Overall, not impressed with Stratix 10. It won't be cost effective, it's not much more power-efficient, and Volta will likely leapfrog it across the board within a year.
Wasn't this thing supposed to sample in late 2014? Back then it would have been a gamechanger at any price. Now, 1080Ti for $700 beats it across the board in throughput/$. NVIDIA's confusing messaging about using consumer versus professional HW is about the only thing that might make it viable for deep learning. Although I note the absence of training perf numbers here, just (apparently) inference.
The Titan X Pascal (and 1080ti) uses the smaller GP102 chip, which as znfi pointed out is practically useless for FP16.
I saw some of their II-V kits costing 1-100k+ but didn't find the Stratix 10 mentioned in the article in their kit list: https://www.buyaltera.com/Search?keywords=stratix+kit&pNum=1
>Intel Stratix 10 FPGA is 10%, 50%, and 5.4x better in performance (TOP/sec) than Titan X Pascal GPU on GEMMs for sparse, Int6, and binarized DNNs
My guess is that while electricity costs would be much higher, it would be better at this time to still just buy ceil(.1, .5, 5.4) Titans instead.
Actually, they are similar in price.
medium to high end consumer FPGA and GPU will go up to around $1k.
Then, there are the entreprisey GPU (Quadro and FireThing as I recall) going for a few k.
It's similar for FPGA. The very high-end (Stratix/Virtex) will charge a few k as well, peaking at 5k or 10k for the top models (nude FPGA chip only).
I recall negotiating some FPGA devkits in the 10-15k€ range, that seems to be the top end. If I remember well, there was an option to get 4 Virtex on the same devboard for 30k or 50k. That's as high as it gets.
My memory ain't perfect but that's about this much.
Looks like the cheapest one is $17k (so 15 Titan Xs? 24 1080Tis?), and I'd expect the price will be around that after full release.
If a FPGA vendor made a FPGA solution (both the hardware and the software libraries to integrate with one or two machine learning frameworks) that did basic matrix/tensor calculations faster/cheaper than GPUs, then they'd be able to take a lot of market off nvidia. Users wouldn't have a need to program the FPGA directly if they can work at the level of matrix operations.
It's certainly possible to make. It's also a very expensive very specialized appliance. All of that to do some matrix manipulations.
But the point is that you don't have that much FPGA-specific code - once someone does the matrix manipulations and the proper integration, everyone else can just run e.g. tensorflow code on it faster and/or cheaper without specific expertise; if a FPGA vendor can do this one-time investment in software tools, then they can compete for a slice of the large pie of hardware revenue that nvidia now has for itself.