You can store it and move it around, but arithmetic operations are prohibitively expensive without hardware acceleration.
(Note that bfloat16 has a different range than float16, so you can't interpret one as the other)
Oh, I should have clarified - could one start with a bfloat16 on software-side, convert to float16 (so that e.g a 3.4E38 float16 becomes a 65504 float16), then do any "heavy math" in fast hardware float16 instructions, and then convert back at the end?
Nothing necessarily wrong with that code, but it also kinda smells. Why even store it as bfloat16 at all? You risk getting the numerical disadvantages of both float16-representation, and none of the advantages.
It's possible to act, in software, as any type of arbitrary length numeric value. It's much faster to do it in hardware thought.
Oh, I should have clarified - could one start with a bfloat16 on software-side, convert to float16 (so that e.g a 3.4E38 float16 becomes a 65504 float16), then do any "heavy math" in fast hardware float16 instructions, and then convert back at the end?
No, you can't. The exponent in a float16 is too small. You'd rather convert back to a 32-bit float, do your operations, and then throw away the surplus precision and convert back to bfloat16.