I have some ~15 year old experience with the math behind some of this, but actually none with day-to-day deep learning applications using any of the now-conventional algorithms, so my perspective here is perhaps not that of the most pragmatic user. The status quo may have improved, at least de facto.
1. https://spectrum.ieee.org/floating-point-numbers-posits-proc...
I am familiarizing myself with recurrent neural networks and getting them trained online is a pain - I get NaNs all the time except for very small learning rates that actually prevent my networks to learn anything.
The deeper network is, the more pronounced accumulation of errors in online training is. Add 20-30 fully connected (not highway or residual) layers before softmax and you'll see wonders there, you won't be able to have anything stable.
Turns out our results were better than the papers we compared to, both in time and precision.
I am not that familiar with ml, but can't you just ignore those faulty weights?
I'm not at all caught up with the this side of ML but my first instinct is that faulty weights would lead to interpretability issues. The numbers represented by NaN/Inf vastly outnumber the ones within precision range, so interpreting them is much more of a guess.
what good will it do to compute something if its error is unbound?
the issue of the accumulation of roundoff errors is generally speaking unavoidable when it's linear but fortunately they tend to be small