The case of the curiously slow shader
raphlinus.github.io
raphlinus.github.io
My crazy shader story: 8 bit floats.
Back in 2007-ish, I was having a hell of a time porting a shader from my desktop to my laptop. For the most part it seemed to work but the output was all one color. That isn't particularly strange for a bug in the shader world, so I went about bisecting it and narrowed the problem down to a critical statement. The input parameters were correct, as verified by color-mapping them, but whenever I performed a division the output would fix to a single color. I could get division to work by feeding it different parameters, though, so what could be the problem?
My big break happened when I noticed that a tiny zone around the edges was outputting a different color. The calculation was happening, it was just unexpectedly almost always producing a single value due to extreme rounding. With the problem identified, it was easy enough to work around by using a scaling factor. After paging through Intel documentation I finally found the cause: the integrated graphics card supported floating point all right, but with 8 bit floats. Implicit, mind you, and there was no compiler warning. I asked for a float and it gave me a float, what's the problem? ¯\_( ツ )_/¯
I'm guessing someone tasked the Intel Integrated graphics team with supporting floating point but didn't really give them the transistor or TDP budget for it, so some evil genius noted that they never specified the bit width and cheerfully squared that circle.
And what shader language? It looks like GLSL expects `float` to be a normal float32.
https://www.khronos.org/opengl/wiki/Data_Type_(GLSL)#Scalars
The language was GLSL. I'm glad to see that sometime in the last decade and a half they decided to mandate float32. Pulling out the float8s to win a benchmark these days would be... something!
Programmers and insufficient specifications, name a more iconic duo?
After spending 2 days on debugging, I found out it used 12 bit floats. Closed as "won't fix".
I don't think there's a universal standard, but an 8-bit float would probably be broken down as:
1 bit sign 4 bit exponent 3 bit mantissa
This allows it to represent numbers in the range of 480 down to 0.0078 (and their negatives).
Of course, it can only accurately represent 256 distinct numbers out of that range, so accuracy will be bit compromised.
I wonder if they were ever legitimately used for a real-world program?
Do you have any other debugging tips or tooling recommendations for a complete graphics noob?
I have only done one webgl project before but that experience was pretty painful. When writing my first glsl shader I basically committed every performance sin possible. Like in your case, the slowness is hard to notice on a beefy desktop, it was only on mobile devices when it becomes apparent. My debugging process was then to comment out lines, rerun, and squint at the fps. There must be a better workflow than this!
(iirc, my issue was that I was using "discard" in my fragment shader which is apparently especially bad on powervr architectures https://stackoverflow.com/questions/8509051/is-discard-bad-f...)
There's a fair amount of academic work on GPU simulators, but a lot of the activity there seemed to be a few years ago. And in any case, I'm not sure how usable these are for tuning real workloads. If you have good hardware performance counters, those can give you almost as good information as you'd get out of a simulator.
I think there's tremendous potential to do better tools in this space. I'm not quite sure what the business model is, though.