bfloat16 is stable enough for ML training, which is what TPUs exist for
Not trying to be dismissive just saying that... computational limits are limits on what can be done, in the end.
I stand corrected.
Edit: the term I could not remember is tradeoff.