[0] https://cloud.google.com/blog/products/ai-machine-learning/b...
Not trying to be dismissive just saying that... computational limits are limits on what can be done, in the end.
I stand corrected.
Edit: the term I could not remember is tradeoff.
I have some ~15 year old experience with the math behind some of this, but actually none with day-to-day deep learning applications using any of the now-conventional algorithms, so my perspective here is perhaps not that of the most pragmatic user. The status quo may have improved, at least de facto.
1. https://spectrum.ieee.org/floating-point-numbers-posits-proc...
I am familiarizing myself with recurrent neural networks and getting them trained online is a pain - I get NaNs all the time except for very small learning rates that actually prevent my networks to learn anything.
The deeper network is, the more pronounced accumulation of errors in online training is. Add 20-30 fully connected (not highway or residual) layers before softmax and you'll see wonders there, you won't be able to have anything stable.
Turns out our results were better than the papers we compared to, both in time and precision.
I am not that familiar with ml, but can't you just ignore those faulty weights?
I'm not at all caught up with the this side of ML but my first instinct is that faulty weights would lead to interpretability issues. The numbers represented by NaN/Inf vastly outnumber the ones within precision range, so interpreting them is much more of a guess.
what good will it do to compute something if its error is unbound?
the issue of the accumulation of roundoff errors is generally speaking unavoidable when it's linear but fortunately they tend to be small
Instability leads to divergence from the true answer, and I would expect it to mean super-linear divergence (though I am not an expert in this) which would quickly destroy any meaningful result (=> chaotic behaviour). But I'm not an expert.
(Disclaimer: googler, I have nothing to do with this research.)
> One important strength of AlphaTensor is its flexibility to support complex stochastic and non-differentiable rewards (from the tensor rank to practical efficiency on specific hardware), in addition to finding algorithms for custom operations in a wide variety of spaces (such as finite fields). We believe this will spur applications of AlphaTensor towards designing algorithms that optimize metrics that we did not consider here, such as numerical stability or energy usage.
(apologies if I misunderstood, I wasn't calling you out specifically but a generalized misconception I've noticed in a lot of other discussions so far)
The day when a lot of wrong math adds up to a computer drawing a pretty picture. Who would have thought.
This means regardless of how big your matrix is, or how big or small your numbers are — or even the relation between them —, you algorithm is going to be stable and accurate. If it can't be (stable), the library must let you know. This is so you can keep the numbers scaled such that they are not too large, nor too small in order to keep them stable for the operations you need to execute.