Their “CUDA core” is not warp-wide, it’s a single lane.
If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.
If you’re talking about FP32 rates, they double it because of FMA (floating-point multiply-accumulate). Everyone does that.
No comments yet.