Supporting half-precision floats is annoying
futhark-lang.org
futhark-lang.org
You have to select AHP by setting an FP config register bit, so (unlike bfloat16 vs binary16) it's a "for this whole chunk of code I am going to use this format" choice.
The upshot is that it's basically an in-memory storage format, and all the actual data-processing gets done at either single or double precision.
Mind this was for a regulated medical device and therefore any change in the software was burdensome. The change didn't affect the internal calculations since it was only for the communication protocol and UI, and the compiler support allowed the minimal changes on both embedded and host source codes.
"just don't use f16" seems like the course of wisdom here.
I had a 3x 5 bit value, packed structure in something where memory pressure was severe, and it was such a bitch on a SPARC to deal with 16 bit quantities that actually running the data as two passes using more memory wound up being an immensely better approach. The Alphas would diddle yer bits any way you liked at speed, but that was clearly an aberrant ability.
https://software.intel.com/sites/landingpage/IntrinsicsGuide...
The result is that Winograd convolutions can achieve an effective FMA rate of twice the peak rate of the CPU.
The Winograd transform reduces the required number of FMAs by a factor of 5x, but you can only do FMAs at half peak rate (because you are bandwidth limited), so you come out ahead by a factor of 2.5x in theory (2x in practice).
Without fp16, that 2x advantage would be lost.
On the GPU on the other hand 16-bit floats are becoming the standard (the M1 GPU for instance has more 16-bit ALUs than 32-bit). With enough precision for possible resolutions/worlds and you save 2x the memory which makes it a no-brainer really.
The main issue is many compilers have issues on x86 platforms due to Intel's bizarre slowness in defining how to pass fp16 parameters in their official ABI.
Intel support is a bit more dubious, especially if you're stuck on older LLVM versions, but __fp16 works for most purposes if all you want is storage. You just have to pass any fp16 function parameters as pointers/references to work around the aforementioned ABI issue. Architecturally intel has supported fp16 conversions necessary for __fp16 for quite a while now, so no reason for it not to work (whereas full fp16 arithmetic is restricted to some less common variants of avx512 iirc, so the distinction between __fp16 and _Float16 isn't much on most intel systems).
It should be fairly straightforward to write a quick test program and see what your local compilers can cope with.
https://lucene.apache.org/core/3_0_3/fileformats.html#N107EF
Does any hardware offer "configurable" bit allocation like f16[e=4,m=11]?
The post is actually quite relatable if you have experience in silly toolchain issues, i.e. worked on a compiler or something like that. It's all fairly straightforward. Please try to read the article next time, you might surprise yourself.
This might be useful in graphics, for example.
In general, I believe it is good to give people options. You never know what people will find useful.
But if you don't want support it and you feel it will make your language better, just don't support it.
But please, quit bitchin' about it.
Java does not support unsigned integer type and everybody is fine.
And Java devs don't write rants about how unsigned integers are useless, supporting them is annoying and everybody should just forget about them.
But still, Java does not have unsigned integer types for 32 and 64 wide integers, just a way to treat them as such if you need it so much that you want to go out of your way. For example when you need to parse a binary with unsigned integers.
There is no way to declare unsigned 32 or 64 bit integer.