Stratix V accelerator card from eBay
j-marjanovic.io
j-marjanovic.io
https://j-marjanovic.io/stratix-v-accelerator-card-from-ebay...
https://j-marjanovic.io/stratix-v-accelerator-card-from-ebay...
https://j-marjanovic.io/stratix-v-accelerator-card-from-ebay...
Useless. The drivers are horrible and the state of the tooling is insanely bad. The vendors have no clue, Intel personnel have no clue as they don't care about OpenCL 2.0 but will still gladly charge $10,000 in licenses if you really want to go that route. Verilog/VHDL is wtf, and you still need to redo your drivers.
So much potential.
I talked to a couple companies that use these in production, they reveal that they wrote their own drivers and still have issues.
You pretty much have to be a Microsoft doing sci-fi tech to get the benefits out of these, and it doesn't have to be that way.
OpenCL was a nice try but everyone did their own thing with it and are now going back to their own vendor specific frameworks. No leadership, irreconcilable business decisions, unnecessarily crippling an open source community, software licenses are too expensive probably because they can't mass monetize the FPGAs. Its just really bad, it doesn't seem like it belongs in this century.
The FPGA vendors see themselves as ASIC replacements so from their perspective they are "killing it" across all metrics. We know where this needs to go.
https://www.microsoft.com/en-us/research/uploads/prod/2014/0...
So it may be possible, but the architecture of the FPGA is not documented, so you'd have to completely reverse engineer the bitstream format, and if the protection is enabled, that'd still get you nowhere.
Stratix II: https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.58...
There's been a little work on vulnerabilities in the Stratix V, though AFAICT no break on the encryption: http://www.ecs.umass.edu/ece/tessier/ramesh-fccm18.pdf
And aside from the encryption there's been a lot of work on FPGA bitstream reverse engineering tools, so the output wouldn't be meaningless—it would just take some effort to figure out what it's doing.
My advice: get some board from Digilent, they have examples and active support forum. It will be more useful on the long run.
[1] https://www.intel.com/content/dam/www/programmable/us/en/pdf...
[2] https://www.intel.com/content/www/us/en/programmable/buy/des...
The logic blocks and interconnects are relatively slow compared to a tailor made ASIC, but they're still generally capable of propagating signals approaching GHz clock rates.
As an example, with a really nice CPU you might be able to run billions of instructions per second, but you are limited to what those instructions can do. If you want to sum a bunch of integers, you need to do it sequentially by implementing a loop. Maybe you can sum 1 billion integers per second. With an FPGA, you could instead use the available blocks to implement a thousand adders in a logarithmic pyramid, able to sum 500 numbers at once at say 100Mhz, and the circuitry to speak to a CPU and its memory directly over PCI-E. Now, you can sum 50 billion integers per second, assuming the RAM can even keep up.
I can't imagine there's any automatic speed up of C code with this specific card, unless the card maker implemented a really fancy compiler. More likely, there's a specific algorithm that the card was designed to do in parallel faster than a CPU can. Someone had to implement that algorithm in a language that could be compiled down to the FPGA bitstream. On top of that, it's possible the card wasn't even meant to be tightly coupled to a traditional CPU. It's possible that the FPGA was meant to participate directly on a data center network and respond to requests directly. The whole "program" could be implemented as state machines in digital logic, including the TCP/IP stack. Also, "soft" CPUs can be implemented inside an FPGA, making it possible to run traditional programs too.
You can generally use an FPGA to implement any algorithm that's parallelizable to some degree. As a master's project 11 years ago, my team accelerated stereo disparity calculations with an FPGA. The idea is that you match similar pixels between two spatially separated images in order to generate a depth map. It's an expensive calculation. At the time, a CPU could manage a few frames per second from 480p webcam streams. The FPGA version could handle hundreds of frames per second. We couldn't feed it in video fast enough. Of course, nowdays a modern GPU running a shader program might be faster.
I was just about to say: when to use an FPGA over a gpu? There are a few GPU gems books that describe what kind of code runs best on GPUs, and you can buy 30TFlops for $700 these days.
I know gpu is not so good beyond single precision. Is that the application? Do you recommend a similar FPGA gems book/website?
I see there is: designing with xilinx fpgas, design warriors guide to fpgas
An FPGA gives you more flexibility because you can dedicate hardware exactly where you need it. You don't even need to stick with common data widths. If you only need 11 bits, then there's no need to do 32 bit math, for example.
An ASIC with custom circuitry will generally be the fastest, lowest power way to solve a problem, with the lowest per-unit cost, but it has absurdly high development and initial tooling cost. An FPGA gives you the benefits of custom circuitry with significantly lower development cost (than an ASIC), but it's less power efficient and the unit cost is relatively high for what you're getting. A GPU gives you good power efficiency, good development cost, and good unit cost if your problem maps well to what GPUs can do. A CPU gives you the lowest development cost, but sets a mediocre baseline for power efficiency and unit cost.
Most QSFP DAC cables are passive and do not use power for anything other than their ID EEPROM.