Are there benefits to using FP32 vs FP16? I’ve been dabbling with deep learning but not really sure how much affect higher precision is having. Though more precision is better I suppose.
With FP16 one can theoretically get 2x speed and 2x larger models with the same VRAM capacity. For inferencing with INT8/INT4 it can be even way better (good for embedded stuff). The downside is that sometimes more complex/deep models don't converge (or converge less often than FP32). Sometimes there are framework issues with some advanced FP16 stuff.