And my 6-year-old Consumer GPU (Vega64) has 4096 32-bit multipliers at 1500 MHz clock. (And I bought it a year or two late, so it was only $400). With 8GB HBM2 RAM at 512GBps (yes, Gigabytes per second) throughput. You absolutely are not getting anything close to a consumer GPU in terms of performance or cost on any other platform.
---------
FPGAs advantage is that systolic array arrangement. I'm pretty sure GPUs will win in a "hardware multiplier" war, but it might be harder to get all the data into a typical GPU
There's a fair number of programs that will never utilize a GPU with any level of efficiency. That's where FPGAs come in, utilization will be better because of LUTs and Routers and all that magic.
But if people are looking for a TFLOP fight with raw numbers of multipliers? My bet is on the GPU and SIMD-processing in general. Those architectures are just crazy.
--------------
BTW: That top end DSP slice is documented here: https://docs.xilinx.com/v/u/en-US/ug579-ultrascale-dsp
That's a 27-bit x 18-bit multiplier per dsp.
In contrast, a 32-bit floating point operation is 24-bit x 24-bit multiplier plus a bit of extra stuff for the exponent bits. I'm pretty sure that the 4096-shader Vega64 actually had 32-bit multipliers but I'm not 100% sure on that.
But the float 24-bit x 24-bit multiplier would need more than one dsp to replicate, maybe two of them. AMD's newest consumer 7900 xt GPUs are 6144-shader... but those shaders can execute _TWO_ multiplies per clock tick for a total of 12288 hardware multipliers operating at 2500 MHz clock.