I imagine that how well this works also strongly depends on the kinds of rounding. I imagine that stochastic rounding, or the rounding used in Google's bfloat16 are different in this regard in comparison with standard IEEE floating point rounding.
I imagine that how well this works also strongly depends on the kinds of rounding. I imagine that stochastic rounding, or the rounding used in Google's bfloat16 are different in this regard in comparison with standard IEEE floating point rounding.
I believe it is not used because it requires 16bits precision which, nowadays, you only get on GPU. People usually train on GPU but then evaluate on CPU (in production) where the discontinuity would be much smaller (as you would use 32 bits precision).
Furthermore I don't know if, in practice, that type of discontinuity trains as well as a classical activation function (the gradient propagation might be hindered by the limited precision).
AFAIK analog computers are still the standard in radars for example and it sound like neural networks would benefit from similar hardware.