Very likely yes, but FPGAs often have hundreds to thousands of hardware multipliers, as part of the DSP blocks. Here for example newer AMD FPGAs: https://eu.mouser.com/datasheet/2/903/ds890_ultrascale_overv...
Very likely yes, but FPGAs often have hundreds to thousands of hardware multipliers, as part of the DSP blocks. Here for example newer AMD FPGAs: https://eu.mouser.com/datasheet/2/903/ds890_ultrascale_overv...
You're giving completely the wrong impression about dsp slices - it is absolutely not 1 dsp slice per FP operator at any precision that you would want to do floating point arithmetic. It's definitely at least 2 plus a whole bunch of LUTs (~500) for FP16 with 4 stages or something like that. And if you want faster (fewer stages) then you need more slices. On alveo u280, which is an ultrascale part, I have never been able to effectively utilize more than ~4000 dsp slices (out of 9024) for 5,4 mults and that cost basically 99% of clbs in SLR1 and SLR2.
And even then, disconnected FPUs are completely meaningless without a datapath implementing eg matmul and boy oh boy do you have no clue what you're in for there.
Takeaway: it's pointless to compare raw specsheet numbers when everything comes down to datapath.
Ain't that the truth.
However: Xilinx's DSP58 blocks (Versal devices), and older Intel DSPs, do integrate floating point into the DSP tiles - which does narrow the gap between the datasheet's DSP count and the number of floating-point operations achievable per clock.
this is correct (i am currently working on designs that target these parts) but those things are more expensive than comparably sized/equipped nvidia gpus so (in the context of this discussion) it's moot.
---------
FPGAs advantage is that systolic array arrangement. I'm pretty sure GPUs will win in a "hardware multiplier" war, but it might be harder to get all the data into a typical GPU
There's a fair number of programs that will never utilize a GPU with any level of efficiency. That's where FPGAs come in, utilization will be better because of LUTs and Routers and all that magic.
But if people are looking for a TFLOP fight with raw numbers of multipliers? My bet is on the GPU and SIMD-processing in general. Those architectures are just crazy.
--------------
BTW: That top end DSP slice is documented here: https://docs.xilinx.com/v/u/en-US/ug579-ultrascale-dsp
That's a 27-bit x 18-bit multiplier per dsp.
In contrast, a 32-bit floating point operation is 24-bit x 24-bit multiplier plus a bit of extra stuff for the exponent bits. I'm pretty sure that the 4096-shader Vega64 actually had 32-bit multipliers but I'm not 100% sure on that.
But the float 24-bit x 24-bit multiplier would need more than one dsp to replicate, maybe two of them. AMD's newest consumer 7900 xt GPUs are 6144-shader... but those shaders can execute _TWO_ multiplies per clock tick for a total of 12288 hardware multipliers operating at 2500 MHz clock.