Virtex UltraScale+ FPGA Augmented with Co-Packaged – Community Forums
forums.xilinx.com
forums.xilinx.com
Although, 8GB of HBM could do some nice things. Also the UltraScale+ plus are one of thew few FPGAs capable of 1 GHz sequential logic for quite a few designs. So you could probably get a lot of performance for some signal processing or compute.
The clock recipe is examined as part of the "AFI Manifest" when Amazon creates FPGA images you can load, so I doubt you can go out of the way very much, here.
https://github.com/aws/aws-fpga/blob/master/hdk/docs/clock_r...
https://github.com/aws/aws-fpga/blob/83c80efd30862d862ea8f99...
Xilinx has the lower/mid scale FPGA market fairly locked up in terms of licensing and available boards, I think, outside of a few niche Intel offerings (DE10 Nano is quite good).
But you aren't going to get anything close to on-board HBM2 without spending north of $10k USD, or something.
Also: that looks scary! Didn't see a price, is $nnnn enough? Whoa.
Are there any really good papers, projects, or products that show where FPGAs provide a major commercial benefit over a GPU?
* http://jaewoong.org/pubs/fpga17-next-generation-dnns.pdf
* https://www.researchgate.net/publication/284691714_The_battl...
The first paper is actually one I've spent a significant amount of time trying to use, to the point of collaborating with one of the authors. His conclusion was that FPGAs used to be competitive with GPUs for approximated nets, but the Tesla GPUs were such a jump forward in practical network performance that it wasn't worth trying to compete outside specialized realms like binary nets.
The second paper was interesting- I can imagine why the problem they are trying to solve would be a good fit for FPGAs. However, I'm suspicious that they implemented an entirely different algorithm on the FPGA, and didn't measure the performance of that algorithm on GPUs. I'm all for using the best algorithm for the hardware, but I worry they just used an overall better algorithm on FPGAs and conflated the results.
I agree this was a bit suspicious. It may be the case that the different algorithm they used for the FPGA would have done well on a GPU -- or perhaps more likely, that if they spent a similar amount of effort in rethinking the algorithm just for the GPU, ending up with a third GPU-specialised approach, perhaps that would have done dramatically better.
Pragmatically, it seems like they chose the GPU for their application anyway - so they had already decided the GPU was the overall winner without needing to improve it.
A couple applications that come to mind:
- RF, including cellular base station hardware and radar
- ASIC prototyping
- Low production run computer hardware, including some RAID controllers
IMO, those are the main factors that justify FPGA selection - low latency and hard real-time performance. I understand that military and industrial designers make extensive use of FPGAs for these reasons; the throughput isn't necessarily any better than an ordinary processor and the cost is drastically higher, but you have absolute certainty about latency.
Regarding cost, these are expensive, complicated instruments. A bottom of the barrel oscilloscope costs $300 and professional grade units are more like $2-3000. The top grade ones can cost half a million dollars ( https://www.keysight.com/en/pcx-x205212/infiniium-z-series-o... ).
Do you have a recommendation for a specific data processing experiment of theirs I should check out? I really feel like I just missed a paper where they proved some real advantage over other hardware, and once I found that I'd understand. I respect the Azure teams generally and assume they know what they're doing- but can't escape the hunch that what they are doing is network acceleration, and are just releasing to the public cloud because they have these sitting around anyway.
So if your looking for a solution that will perform faster on FPGA you are going to want something that is simple but you need compute often. That way you can duplicate it 1000s of times on the FPGA. An other place an FGPA excels is data that quite long. Compare 32 bit number to 1024 bit number. You could do what ever your doing to the 1024 bit number in one pass with an FPGA. However, the GPU's native int size is probably 32 bits. So to just perform one operation the GPU has to perform at least 32 operations for that one number. So that overhead has to be carried around for every operation a GPU core would have to perform. The HBM added to this FPGA makes it even better in cases like this. That's just the general idea though.
So if what you are doing can take advantage of the FPGA's strengths you could come out with a much faster solution. Also there is power usage. If your going to be building clusters to perform what ever processing you need and the FPGA performs about the same as GPU you will use less electricity.